
GPT-5.6: Sol, Terra, Luna, and the Ultra Agent Tier
OpenAI's July 9 family spans flagship, balanced, and efficient models, with state-of-the-art terminal, coding, browsing, and science results.
Read MoreZharfAI Research
Model release desk

Anthropic released Claude Sonnet 5 on June 30, 2026 as its most agentic Sonnet model and the default for Free and Pro plans. The official announcement positions it close to Opus 4.8 at lower cost. The system card reports 63.2 on SWE-bench Pro, 80.4 on Terminal-Bench 2.1, 84.7 on single-agent BrowseComp, 81.2 on OSWorld-Verified, and 1,618 Elo on GDPval-AA v2.
The model's most important control is adjustable effort. Medium effort targets cost efficiency; higher settings spend more tokens and time and can match Opus 4.8 on some tasks. Introductory API pricing is $2 per million input tokens and $10 per million output tokens through August 31, then $3 and $15.
The Sonnet 5 system card says standard results use adaptive thinking at max effort, default sampling, and five-trial averages. Context varies; BrowseComp uses up to ten million tokens with compaction at 200K.
This disclosure makes Sonnet easier to evaluate, but it also prevents casual score mixing. Our frontier model evaluation guide explains how to preserve harness, effort, context, tool, and sampling conditions beside each value.
| Evaluation | Sonnet 5 | Sonnet 4.6 | Meaning |
|---|---|---|---|
| SWE-bench Verified | 85.2 | Not in summary table | Repository issue resolution |
| SWE-bench Pro | 63.2 | 58.1 | Hard software engineering |
| SWE-bench Multilingual | 78.3 | — | Nine-language repository work |
| Terminal-Bench 2.1 | 80.4 | 67.0 | Command-line agent execution |
| BrowseComp | 84.7 single; 86.6 multi | 76.2 single | Difficult agentic search |
| HLE | 43.2 no tools; 57.4 tools | 34.6; 46.8 | Broad hard questions |
| OSWorld-Verified | 81.2 | 78.5 | Computer use |
| GDPval-AA v2 | 1,618 Elo | 1,381 | Professional knowledge work |
These are not the scores of the cheapest effort setting. A product choosing medium effort should build its own curve instead of attaching the max-effort table to a lower-cost endpoint.

Sonnet 5's 63.2 on SWE-bench Pro improves by 5.1 points over 4.6 and remains below Opus 4.8's 69.2 in Anthropic's comparison. On SWE-bench Verified it reaches 85.2, and multilingual performance is 78.3.
The gap can narrow under effort controls, but buyers should compare equal cost and equal harness. Repository success needs patch-quality review, tests, minimal changes, and security. A cheaper model that needs more corrections may not save money.
Terminal-Bench uses mini-SWE-agent in the system card because Anthropic found Terminus-2 suffered 2.7 times more timeouts at xhigh. That methodological choice is exactly why harness labels matter. A raw “80.4” without mini-SWE-agent is incomplete.
Anthropic edited the launch post on June 30 after the original BrowseComp cost-performance chart used a simpler method that underestimated Sonnet 5. The corrected setup matches the system card: a ten-million-token budget, context compaction, and programmatic tool calling.
Corrections are preferable to leaving a bad chart, but the episode shows that benchmark methodology can change on launch day. Editorial records should use the updated score and preserve the changelog. Procurement teams should archive the source version used for a decision.
Multi-agent BrowseComp rises from 84.7 to 86.6, but adds orchestration and cost. Research workflows need citation support, source diversity, and deduplication, not only final answer accuracy.
OSWorld-Verified reaches 81.2, up from 78.5 for Sonnet 4.6. The footnote says Anthropic fixed a zoom-tool bug and raised maximum tokens per turn to 128K. That means part of the measured system improvement comes from the tool and budget as well as the model.
AutomationBench reaches 13.5 versus 5.3 for 4.6, while Gemini 3.5 Flash is 14.5 in the same summary. Scores remain low in absolute terms, reminding us that complex automation is not solved.
Computer agents need isolated accounts, least privilege, action logs, and confirmation for irreversible work. Our human approval design guide explains where to place those gates.
Effort levels create a family of cost-performance points inside one model. Low or medium can serve ordinary coding and research; xhigh can handle difficult debugging or long analysis. The application should route based on task class and risk, not the model's self-confidence alone.
Build curves for success, tokens, latency, and reviewer time at each setting. Include the tokenizer change noted by Anthropic: the same input may map to roughly 1.0–1.35 times as many tokens depending on content. Introductory pricing was designed to ease that transition, but standard pricing starts after August.
Cost per accepted task is the deciding metric. Effort that doubles tokens but eliminates repeated failures can still be economical.
Anthropic reports fewer undesirable behaviors, hallucinations, and sycophancy than Sonnet 4.6 and better resistance to prompt injection. Sonnet 5 shows lower cyber capability than Opus or Mythos and produced no complete exploit in the Firefox evaluation described in the launch.
The model still receives cyber safeguards and shows slightly more partial exploit success than 4.6. It also scores worse on some alignment measures than Opus 4.8 and Mythos. Safety is not monotonic with price or family tier.
Organizations should test their tools, local languages, and data. Provider safeguards do not replace authorization or sandboxing. Safe refusals should be reported separately from capability errors.
Sonnet 5 is a strong default for coding agents, search, document work, and computer use where Opus cost is hard to justify. Its one-million-scale evaluation contexts and effort controls make it flexible for long tasks. Opus or Fable may remain better for the hardest jobs; smaller models may be better for routine extraction.
The release's benchmark record is unusually useful because it exposes trial averages, context, harness choice, a corrected methodology, and predecessor values. Use those details to build a local effort curve. The right conclusion is near-Opus capability over a wider cost range—not that one Sonnet score applies to every setting.
Treat the move from Sonnet 4.x or another provider as a versioned system change. Freeze a representative set of repository issues, browser-research questions, spreadsheet and document tasks, and computer-use flows. Preserve the old model's output, tool transcript, elapsed time, token bill, reviewer corrections, and final disposition. Then run Sonnet 5 with the exact permissions intended for production rather than an unconstrained demo account.
Use separate thresholds for answer quality, task completion, regression rate, latency, cost, and unsafe action attempts. A higher benchmark score does not compensate for a model that asks for broader permissions or produces harder-to-review patches. Conversely, a longer trace can be acceptable when it exposes evidence and reduces reviewer work. Measure the complete accepted workflow.
Roll out by task class. Start with read-only analysis and draft generation, then bounded code changes and reversible browser actions. Keep sending, publishing, deployment, deletion, and credential changes behind explicit confirmation. Record the model ID, effort, context policy, tool versions, and date for every canary so a later service update can be isolated. A fallback model and resumable checkpoint are especially important for the occasional slowness and timeout risks documented by the provider.

OpenAI's July 9 family spans flagship, balanced, and efficient models, with state-of-the-art terminal, coding, browsing, and science results.
Read More
Anthropic's May 28 flagship improves repository work, long-horizon agents, computer use, and visual reasoning, with a system card that exposes the caveats.
Read More
Google's May 19 model pairs fast inference with strong coding, tool-use, multimodal, and long-horizon scores—but the harness still defines the result.
Read MoreGet in touch with our team to discuss how we can help your business.