Solar Pro 4 by Upstage finishes real agent work end-to-end. 512K context, 128K output, adjustable reasoning, trilingual EN/KO/JA, evidence-based answers.
Capabilities, design details, and architectural traits
Solar Pro 4 is designed to complete multi-step office assignments - reading documents, running tools, and producing final deliverables - rather than stopping at a single answer. It was trained on finished work through OfficeVerse, Upstage's pipeline that synthesizes office tasks from real public data across 11 industry domains and 12 task types, grading each one pass or fail on the final deliverable.
| Trait | Detail |
|---|---|
| OfficeVerse training | Trained and validated on synthesized office tasks graded pass/fail on the final deliverable, not just intermediate steps |
| Evidence-based answering | Says it cannot verify when evidence is absent, rather than fabricating claims that flow downstream unchecked |
| 512K context, 128K output | Loads several contracts, reports, and data files into a single session without splitting |
| Adjustable reasoning effort | Set high for deep analysis or low for real-time, chatbot-speed interaction |
| Terminal task completion | Completes multi-step jobs in a live shell, not just generating commands |
| Sequential deliverable production | Produces Excel workbooks, Word reports, and PowerPoint decks in sequence within one session, with values carrying unchanged between steps |
| Trilingual EN/KO/JA | Handles English, Korean, and Japanese for both input and output |
The model's defining demonstration is end-to-end office work: given a policy document and six market-data files, it screened ten candidate sites and produced an Excel workbook, a review report, and a slide deck in three prompts. A number computed in the workbook carries into the report unchanged - the point is that values from one step survive into the next deliverable.
Upstage published the Solar Pro 4 Agent Cookbook with seven work agents covering Excel workbook generation, evidence-based Word reports, and PowerPoint decks with charts and speaker notes, each with system prompts, actual outputs, and pass criteria.
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | Upstage | Anthropic | Anthropic |
| Release Date | August 6, 2026 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | - | May 2026 | - |
| Context & Limits | |||
| Context Window | 512K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.30 Best Input Pricing | $5 | $10 |
| Output Pricing | $1.20 Best Output Pricing | $25 | $50 |
| Modalities | |||
| Inputs | textfile | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 41.6 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | 52.7 | 78.0 Best Coding Index | 76.5 |
| Agentic Index | 33.6 | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.