Only 3B active parameters per token, with leading performance in its class and lower costs for repetitive agent workload
We launched Solar Mini 4, an agent-optimized model pretrained from the ground up by Upstage. It activates just 3B of its 35B total parameters per token while achieving the highest overall score among the 3B-active models in the Artificial Analysis Intelligence Index comparison. Solar Mini 4 is designed to handle repetitive tasks—including information retrieval, structured output generation, and tool use—using less compute and at lower cost.
Key capabilities
- Highest overall performance among the 3B-active models evaluated: Solar Mini 4 scored 24.1 on the Artificial Analysis Intelligence Index v4.3.2, the highest result among the models with 3B active parameters included in the comparison.
- Optimized for agent workloads: Solar Mini 4 scored 22.3% on AutomationBench-AA, which evaluates agents on real SaaS workflows, and47.2 in a separate τ³-Banking evaluation.
- Efficient MoE architecture: The model has 35B total parameters while activating only 3B per token.
- 512K context window: Supports up to 512K tokens of context and up to 128K output tokens.
- Low API pricing: Standard pricing is $0.10 per 1 million input tokens, $0.01 per 1 million cached input tokens, and $0.40 per 1 million output tokens.
Highest overall performance among the 3B-active models evaluated
Solar Mini 4 scored 24.1 on the Artificial Analysis Intelligence Index v4.3.2. This is the highest result among the models with 3B active parameters included in the comparison, ahead of Qwen3.6 35B A3B (18) and Nemotron 3.5 Lightning (13).
Compared with Gemma 4 31B, which has 31B active parameters, Solar Mini 4 activates roughly one-tenth as many parameters while scoring 9.1 points higher on the AA Index. It also scored 1.1 points higher than Nemotron 3 Ultra, which has 55B active parameters. Its overall score is 0.9 points lower than MiMo-V2.5, but it delivers comparable performance with one-fifth as many active parameters.
On individual evaluations, Solar Mini 4 scored 83.3% on AA-LCR, which measures long-context reasoning, and 47.6% on SciCode, which evaluates code-based scientific problem solving. It also scored 25.8% on Humanity’s Last Exam, showing strength across long-context reasoning, coding, and knowledge-intensive tasks.
Built to apply policies and use tools in real workflows
Producing a valid function call is not the same as completing a real task. An agent must apply policies, inspect the current state, select the right tools, and adapt across multiple interactions.
Solar Mini 4 scored 22.3% on AutomationBench-AA, which evaluates multi-step workflows across real SaaS applications. It also scored 47.2 on τ³-Banking, a separate benchmark of policy application and tool use across multi-turn interactions.
Solar Mini 4 supports tool calling and parallel tool calling. By defining available functions in the tools parameter and setting parallel_tool_calls=True, developers can call multiple independent tools in parallel—for example, checking an approval status and remaining leave balance at the same time.
Lowering the cost of repetitive agent workloads
The cost of running an agent depends not only on token pricing, but also on the number of calls required to complete a task and how efficiently each call is processed. These differences become more significant when workflows involving retrieval, classification, tool use, and structured outputs are repeated at scale.
Solar Mini 4 is designed to process these high-volume workloads efficiently. Although it has 35B total parameters, its MoE architecture activates only 3B parameters per token. This reduces the compute required for each request while retaining the capacity of the full model.
In internal testing, Solar Mini 4 sustained more than 70 tokens per second per request while processing 32 concurrent requests on two H100 GPUs.*
Active parameter count should not be confused with deployment memory requirements. Although only 3B parameters are active per token, the full 35B model weights must still be loaded into memory. For on-premises deployment, the quantized model can run on a single H100 80GB GPU.
*Measured internally using 4K input tokens and 1K output tokens, without prefix caching, during a continuous load test lasting more than 600 seconds. The throughput test used two H100 GPUs and represents a different configuration from the minimum deployment requirement.
Cost per completed task: Solar Mini 4 vs. MiMo-V2.5
The cost of running an agent depends not only on token pricing, but also on the number of tokens and turns required to complete a task. To measure this, we ran the same agent tasks with Solar Mini 4 and MiMo-V2.5 and compared the total cost of completion.
We ran three agent tasks—covering code analysis and review as well as calendar coordination across multiple tools—three times on each model. Both models passed all quality checkpoints. At standard prices, the estimated cost of completing 1,000 sets of tasks was $1.25 with Solar Mini 4 and $2.20 with MiMo-V2.5, making Solar Mini 4 approximately 43% less expensive.
For the multi-step tool-calling task—checking a calendar and sending messages to attendees—the estimated cost of 1,000 runs was $0.36 with Solar Mini 4 and $0.61 with MiMo-V2.5. Both models completed the full workflow: checking the schedule and attendees, sending the required messages, and summarizing the result.
In this test, MiMo-V2.5 used more input and output tokens to complete the tasks, resulting in a higher overall cost. In environments that call a model repeatedly, teams should consider not only the price per token but also the tokens and turns required to complete the full task.
*This was an initial internal test consisting of three Korean-language tasks, each run three times per model with temperature=0. Costs for both models were calculated without prompt caching. Solar Mini 4 was called directly through the Upstage API, while MiMo-V2.5 was accessed through a third-party router, so latency was excluded from the comparison. Results may vary by workload and operating environment.
Configure it for the task and connect it to your workflow
Solar Mini 4 does not use reasoning by default. For tasks that require multi-step analysis, developers can set reasoning_effort to match the complexity of the task. Higher reasoning levels may increase output tokens and response time, so we recommend starting with the lowest level that meets the task’s quality requirements.
Solar Mini 4 also supports JSON mode, JSON Schema-based structured outputs, tool calling, and parallel tool calling, making it easy to integrate into repetitive agent workflows.
Information extraction and structuring
Extract key terms from contracts, classify customer inquiries, or convert text into JSON and tables. For workflows that repeatedly process a predefined schema, Solar Mini 4 can produce consistent outputs at a lower cost.
Repetitive workflows across internal systems
Check an approval status, apply internal policies, determine the next action, and summarize the result. Tool calling and parallel tool calling can connect retrieval and execution steps across multiple systems.
Validation in document pipelines
Compare results from Document Parse and Information Extract, identify missing fields, or determine the next step based on predefined rules. When a workflow repeatedly validates already structured information, Solar Mini 4 can be a more cost-efficient alternative to a larger model.
Pricing and availability
Solar Mini 4 is available through the Upstage Console API, Solar Chat, and OpenRouter, with on-premises deployment also supported. It will also be free to use on Hermes Agent for a limited time starting October 5.
Standard API pricing is $0.10 per 1 million input tokens, $0.01 per 1 million cached input tokens, and $0.40 per 1 million output tokens. Upstage Console offer a 70% launch discount through October 10 UTC, with standard pricing taking effect on October 1 UTC.
Promotion dates and terms are subject to change.
Get started with the API
You can call Solar Mini 4 with a few lines of code using the OpenAI SDK.
from openai import OpenAI
client = OpenAI(
    api_key="UPSTAGE_API_KEY",
    base_url="https://api.upstage.ai/v1"
)
response = client.chat.completions.create(
    model="solar-mini4",
    messages=[
        {
            "role": "user",
            "content": "Extract the contract value and effective date from the following document as JSON: ..."
        }
    ],
    response_format={"type": "json_object"}
)
print(response.choices[0].message.content)
Get started with Solar Mini 4 in Upstage Console today. For on-premises deployment, contact us to find the right configuration for your environment.
Building and deploying responsibly
For workflows that affect real-world decisions, we recommend using the model’s output as an input to a separate review process rather than as the final decision. Before deployment, evaluate accuracy and failure modes using data that reflects your production environment.