
Claude Sonnet 5: Near-Opus Agents at Sonnet Economics
Anthropic's June 30 model lifts coding, terminal, search, computer use, and knowledge work while exposing effort as a cost-performance control.
Read MoreZharfAI Research
Model release desk

OpenAI released the GPT-5.6 family for general availability on July 9, 2026 after a limited Sol preview. The official family announcement introduces Sol as the flagship, Terra as the balanced tier, and Luna as the cost-efficient tier. A separate ultra setting coordinates four agents by default for difficult work.
OpenAI reports a score of 80 for Sol on the Artificial Analysis Coding Agent Index, 88.8 on Terminal-Bench 2.1, 72.7 on DeepSWE v1.1, 92.2 on BrowseComp, and 62.6 on OSWorld 2.0. The family story is as much about efficiency as peak score: Terra and Luna remain competitive while using less time, cost, and output tokens under OpenAI's comparisons.
Sol, Terra, and Luna are separate operating points. Sol targets maximum general capability. Terra balances quality and cost. Luna handles volume. They should be evaluated independently rather than treated as dynamic labels for one score.
max gives a single model more time to reason. ultra adds multi-agent orchestration, using four agents by default and testing up to sixteen on some tasks. Ultra is therefore a system configuration, not a fourth base model. Its token use, latency, and failure modes differ.
The preview on June 26 was limited access; July 9 is the general family release date. Editorial chronology retains both events without dating the public family to the preview.

OpenAI's first-party table provides comparable family values:
| Evaluation | Sol | Terra | Luna | Interpretation |
|---|---|---|---|---|
| AA Coding Agent Index v1.1 | 80.0 | 77.4 | 74.6 | Composite coding-agent index |
| SWE-bench Pro | 64.6% | 63.4% | 62.7% | Difficult repository resolution |
| DeepSWE v1.1 | 72.7% | 69.6% | 67.2% | Long-horizon engineering |
| Terminal-Bench 2.1 | 88.8% | 87.4% | 84.7% | Command-line agent tasks |
| Terminal-Bench 2.1, Sol Ultra | 91.9% | — | — | Four-agent default orchestration |
The family compresses capability gaps: Luna's 62.7 SWE-bench Pro sits close to Sol's 64.6, while the broader agent indexes separate them more. This suggests routing by task class rather than assuming the flagship wins every cost-adjusted comparison.
Sol reports 92.2 on BrowseComp and 62.6 on OSWorld 2.0. These are different from the BrowseComp and OSWorld-Verified versions used in several Anthropic tables. Version names must remain visible.
OpenAI also introduces Agents' Last Exam, covering long-running professional workflows across 55 fields. Sol reaches 53.6 under the reported configuration, 13.1 points above Fable 5. At medium reasoning it remains ahead in OpenAI's cost estimate. As a new evaluation, its dataset and rubric deserve external scrutiny.
Professional artifacts—spreadsheets, decks, and documents—receive a major product emphasis. Customer benchmarks report fewer steps or tokens, but those are workload-specific testimonials. Local acceptance must validate formulas, citations, editable structure, and design quality.
GPT-5.6 can write lightweight programs that coordinate tools, filter intermediate results, and choose next actions. This can reduce model round trips and keep large tool outputs out of the main context. It also moves logic into model-generated code.
Generated orchestration code needs sandboxing, resource limits, typed tool interfaces, and logs. A filter can accidentally discard decisive evidence. The agent should preserve provenance from raw tool output to final claim.
Ultra parallelizes work and can improve the score-latency frontier. Parallel agents also raise cost and coordination complexity. Use them for decomposable, high-value tasks and require an adjudication step. Our durable agent workflow guide covers checkpoints and idempotency.
The frontier model evaluation guide provides the companion measurement contract: model tier, reasoning setting, agent count, harness, tool budget, context policy, and judge must travel with every score.
OpenAI reports Sol at 28.7 on GeneBench Pro, 59.9 on LifeSciBench, 48.3 on internal MedChemBench, and 60.5 on HealthBench Professional. Terra and Luna trail but remain strong on several rows.
These are capability evaluations, not clinical validation. A model can reason over research and still produce unsafe patient advice. GeneBench and LifeSciBench test technical workflows; MedChemBench is internal. Each needs method and data review before procurement.
Scientific agents should cite source data, preserve units, run reproducible code, and separate hypotheses from observations. High-risk cyber and science access also sits behind OpenAI's verification and safeguard programs.
OpenAI's launch emphasizes fewer output tokens, less elapsed time, and lower estimated cost. Sol reportedly scores 80 on the coding index while using less than half the output tokens and time of Fable 5 and costing about one-third less under the comparison. Terra and Luna target even lower cost points.
Vendor cost estimates depend on harness, retries, and prices. Measure on the same accepted tasks. Cache behavior, programmatic tools, multi-agent calls, and hidden reasoning can alter the bill. Ultra may finish faster while consuming more aggregate compute.
Record task cost, not only per-token rates. A production router can send routine work to Luna, escalate ambiguous cases to Terra, and reserve Sol or Ultra for the hardest tasks.
GPT-5.6 launches with expanded cyber safeguards, red teaming, monitoring, and trusted access. The deployment safety record describes evaluations and mitigations. More capable tool use increases both defensive value and misuse risk.
Provider safeguards are only one layer. Organizations need identity, scoped credentials, network boundaries, approval, and incident response. Multi-agent systems multiply tool sessions and must propagate permissions correctly.
For scientific and professional use, hallucination and evidence failure remain central. Verify outputs through deterministic tools, source links, and domain review. Do not infer authorization or certification from a benchmark.
GPT-5.6 is a major release because the entire family reaches strong agent and coding performance while offering explicit cost tiers and multi-agent scaling. The first-party table is detailed enough to show where Sol's margin is large and where Luna remains close.
Evaluate all three models and Ultra as four distinct systems. Pin reasoning effort, tools, benchmark version, and orchestration. The strongest deployment may not be Sol everywhere; it may be a governed router that spends frontier compute only when evidence shows the task needs it.
Begin with a blind routing study rather than sending every task to Sol or Ultra. Label a representative workload by risk, complexity, latency tolerance, evidence requirement, and expected economic value. Run Luna, Terra, and Sol on the same snapshots with identical tools and acceptance tests. Ultra should enter only the subset that can be decomposed and whose expected gain justifies four-agent coordination.
Measure cost per accepted result, not advertised token price. Include cached and uncached input, output, tool execution, retries, parallel branches, wall time, and reviewer minutes. Track false completion separately from explicit failure; an honest stop can be cheaper and safer than a confident but invalid artifact. Plot both median and tail behavior because difficult agent traces can dominate the bill.
Deployment should use stable model identifiers, a fixed prompt and tool-schema revision, canary tasks, and a rollback route. Put repository writes in review branches, browser work in isolated accounts, and scientific or health outputs behind domain review. For Ultra, archive each subagent's evidence and the adjudicator's reconciliation so disagreement is visible. Sending, publishing, financial action, credential change, and deletion remain human-approved even when the family scores well on tool benchmarks.

Anthropic's June 30 model lifts coding, terminal, search, computer use, and knowledge work while exposing effort as a cost-performance control.
Read More
Google's May 19 model pairs fast inference with strong coding, tool-use, multimodal, and long-horizon scores—but the harness still defines the result.
Read More
Anthropic's May 28 flagship improves repository work, long-horizon agents, computer use, and visual reasoning, with a system card that exposes the caveats.
Read MoreGet in touch with our team to discuss how we can help your business.