New Models
2026-08-17
A Korean lab tripled its own model's score in one generation — and charges 30 cents per million tokens for it
Upstage released Solar Pro 4 this month. On Artificial Analysis's Intelligence Index it scores 42, against 14 for Solar Pro 3 — one of the largest single-generation jumps any lab has posted this year. That puts it level with Thinking Machines' Inkling and one point behind MiMo-V2.5-Pro. The gains are concentrated in agentic work, which is what Upstage says it built for: Terminal-Bench v2.1 rose from 43.2 to 57, long-document reasoning on AA-LCR from 62.7 to 71, and on the GDPval-AA v2 agentic benchmark its Elo went from 498 to 1277, against a human baseline of 1000. It costs $0.30 per million input tokens and $1.20 per million output, with cache hits at $0.06, currently discounted 90% until 10 September. Context is 384K tokens with up to 256K of output, text only, served through Upstage's API and OpenRouter. There is a cost: the model uses 17% fewer output tokens than its predecessor but takes 8.6 minutes per task rather than 6.0. It works in English, Korean and Japanese.
Why it mattersThe interesting number is not 42 — plenty of models sit there. It is 14 to 42 in one generation, from a lab most people outside Korea have never evaluated. The recipe for reaching this tier is now well enough understood that a mid-sized national lab can execute it in months, which is why the price floor keeps dropping faster than the frontier moves. Treat the agentic Elo above the human baseline with the caution any single benchmark deserves, and note the latency: it is buying that score with time on the clock, not just better weights.
✓ Verified · 3 sources
▶ Related video: Solar Pro 4 First Look & Test – South Korea's DeepSeek Competitor!
Read in the app — free, in 9 languages
Related stories
DeepSeek's cheap workhorse can now see — and on agent tasks that need eyes it says it is close to Anthropic's best
2026-08-21Given four hours and a GPU to improve the way AI is trained, the best agent scored 0.25 out of 1 — and most never tried
2026-08-21ChatGPT can now read your iMessages and send them — and the setting that lets it skip asking is the one OpenAI warns about
2026-08-21Show the robot once — three to twelve seconds — and it gets the job right 59 times out of 100 with no training at all
2026-08-21The agent invented a second person to vouch for its code. A 24-year-old in Texas refused to believe either of them.
2026-08-21