Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. Solar Pro 4
Upstage
Released August 6, 2026

Solar Pro 4

Solar Pro 4 by Upstage finishes real agent work end-to-end. 512K context, 128K output, adjustable reasoning, trilingual EN/KO/JA, evidence-based answers.

Visit UpstageAnnouncement
Inputs
Text
File
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Solar Pro 4 - Built to Carry Real Work to the Finish

Solar Pro 4 is designed to complete multi-step office assignments - reading documents, running tools, and producing final deliverables - rather than stopping at a single answer. It was trained on finished work through OfficeVerse, Upstage's pipeline that synthesizes office tasks from real public data across 11 industry domains and 12 task types, grading each one pass or fail on the final deliverable.

TraitDetail
OfficeVerse trainingTrained and validated on synthesized office tasks graded pass/fail on the final deliverable, not just intermediate steps
Evidence-based answeringSays it cannot verify when evidence is absent, rather than fabricating claims that flow downstream unchecked
512K context, 128K outputLoads several contracts, reports, and data files into a single session without splitting
Adjustable reasoning effortSet high for deep analysis or low for real-time, chatbot-speed interaction
Terminal task completionCompletes multi-step jobs in a live shell, not just generating commands
Sequential deliverable productionProduces Excel workbooks, Word reports, and PowerPoint decks in sequence within one session, with values carrying unchanged between steps
Trilingual EN/KO/JAHandles English, Korean, and Japanese for both input and output

From Data to Deliverables in One Session

The model's defining demonstration is end-to-end office work: given a policy document and six market-data files, it screened ten candidate sites and produced an Excel workbook, a review report, and a slide deck in three prompts. A number computed in the workbook carries into the report unchanged - the point is that values from one step survive into the next deliverable.

Seven Work Agents Released

Upstage published the Solar Pro 4 Agent Cookbook with seven work agents covering Excel workbook generation, evidence-based Word reports, and PowerPoint decks with charts and speaker notes, each with system prompts, actual outputs, and pass criteria.

Benchmark Performance

Independent evaluations · Artificial Analysis

41.6%
Intelligence
52.7%
Coding Index
33.6%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science89.1%
Humanity's Last Exam29.2%
SciCode - Scientific Coding43.6%
Long Context Reasoning70.7%
Specs
Context window
512Ktokens
Input pricing
$0.30per 1M tokens
Output pricing
$1.20per 1M tokens
Cached input
$0.06per 1M tokens

Prices in USD.

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderUpstageAnthropicAnthropic
Release DateAugust 6, 2026July 24, 2026June 9, 2026
Knowledge Cutoff-May 2026-
Context & Limits
Context Window512K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$0.30
Best Input Pricing
$5$10
Output Pricing
$1.20
Best Output Pricing
$25$50
Modalities
Inputs
textfile
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index41.6
63.1
Best Intelligence Index
62.1
Coding Index52.7
78.0
Best Coding Index
76.5
Agentic Index33.6
59.2
Best Agentic Index
56.6
Solar Pro 4
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

89%
Solar Pro 4
GPQA Benchmark
Score: 89%
Solar Pro 4
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

29%
Solar Pro 4
Humanity's Last Exam
Score: 29%
Solar Pro 4
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

71%
Solar Pro 4
Long Context Reasoning
Score: 71%
Solar Pro 4
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models