← All models

Claude Sonnet 4.6: historical model reference

1M tokens · Text / Vision / Code · Prompt cache

Okou no longer runs Claude Sonnet 4.6. This page is kept as a reference for its specs, pricing and benchmarks. For the same kind of work, use Claude Sonnet 5.5.

See Claude Sonnet 5.5

Vendor list price is $3 / $15 per 1M tokens, with cached input dropping to $0.30 / 1M. Reach for Opus only when Sonnet visibly fails on the hardest reasoning, and for Kimi K2.7 Code or GPT-5.4 Mini when unit cost dominates.

What is Claude Sonnet 4.6?

February 2026 (Claude 4.6 generation)

Claude Sonnet 4.6 sits in the middle of Anthropic's Claude 4 family. It is the workhorse model designed to handle the full breadth of typical agent work. Multi-tool routing, code edits, long-running conversations, and structured-output tasks. Without the cost premium of Opus.

What's notable about Claude Sonnet 4.6

Headline architecture and capability features.

Sonnet 4.6 ships with the 1M-token context window at standard pricing, adaptive thinking inherited from Opus 4.6, and prompt caching that bills cached input at one-tenth the input rate. It accepts multimodal input across text, vision, and code.

Specs at a glance

FamilyClaude 4 generation
ModalitiesText, vision, code
LanguagesEnglish-first, multilingual
Prompt cachingSupported (Anthropic)
Context window1M tokens
Max outputUp to 64K tokens

Claude Sonnet 4.6 benchmarks

Sonnet 4.6 sits roughly 3 to 4 percentage points behind Opus 4.6 on Anthropic's headline coding benchmarks while being three to five times cheaper at the vendor level. The typical Opus/Sonnet trade-off.

SWE-bench Verifiedvendor-reported
~77%
Long-context recallinternal observation
Strong across 100K+

Claude Sonnet 4.6 pricing

Historical provider list prices, per 1M tokens. These reference figures are not current Okou charges.

Input$3.00
Output$15.00
Cache read$0.30
Cache write$3.75

How Claude Sonnet 4.6 behaves in practice

Observed behaviour from production agent runs.

Tool routing

Best-in-class tool-routing accuracy at this price. On multi-tool flows across Slack, GitHub, Linear, and Notion, Sonnet 4.6 picks the correct tool with the correct arguments more reliably than any model below ×2.

Long-context coherence

Coherent across 100K+ token transcripts. Drops below Opus 4.7 only on the longest, most adversarial runs.

Speed

Faster than Opus and slower than Kimi K2.7 Code. The right speed/quality balance for production agents.

Best agent tasks for Claude Sonnet 4.6

The Slack agent that knows where things live

Triages incoming questions, follows up on stalled threads, posts status updates, and answers search-style queries ("who's owning the auth refactor?"). Sonnet's tool-routing accuracy means the right tool gets called with the right arguments on the first try, even when the request is ambiguous, so the agent feels reliable instead of flaky.

The PR review agent that doesn't drown in noise

Sonnet handles the bulk of code-aware work — PR review, test scaffolding, refactor suggestions, bug bisection — without leaving stylistic comments that nobody asked for. The 1M-token context window lets it pull in the related files and prior reviews when it matters, and you only escalate to Opus 4.7 for the patches Sonnet visibly struggles with.

The research agent that makes 20 tool calls in a row

GitHub plus Linear plus Notion plus the web, stitched together across twenty-plus tool turns to answer a question like "why did this customer churn last quarter?" Sonnet keeps the goal in view across the whole chain at a fraction of Opus's cost, which is what makes it sustainable for everyday research as opposed to one-off deep dives.

The customer-support assistant with a stable system prompt

Long conversation histories, frequent tool calls into the CRM, the same hefty system prompt and tool schema on every turn. Sonnet's prompt caching turns that fixed prefix into a fraction of the input cost after the first call, which is what keeps per-conversation cost flat as volume grows.

When to skip Claude Sonnet 4.6

Skip Sonnet 4.6 on the hardest reasoning steps where it visibly drops instructions and you should escalate to Opus 4.7, on bulk classification at high volume where GPT-5.4 Mini is the cheaper supported bulk option, and on latency-critical micro-replies where Kimi K2.7 Code is meaningfully faster.

Claude Sonnet 4.6 vs other models

Claude Sonnet 4.6 vs Claude Opus 4.7

Sonnet 4.6 is ×1; Opus 4.7 is ×2. Sonnet handles most agents; Opus is the upgrade when reasoning depth matters more than throughput. Many teams use Opus as the planner and Sonnet as the worker.

Claude Sonnet 4.6 vs GPT-5.4 Mini

GPT-5.4 Mini is the cheaper OpenAI-side bulk option. Use Sonnet when tool-routing reliability matters more; use Mini for high-volume pre-filtering and simple steps that do not need Sonnet-class routing.

Frequently asked questions

What is Claude Sonnet 4.6's context window?

1 million tokens with up to 64K tokens of output per response.

Does Sonnet 4.6 support image input?

Yes. It's multimodal. Text, code, and images.

When should I switch off Sonnet 4.6?

Switch to Opus 4.7 if Sonnet visibly drops the goal on long agent loops or fails on hard code edits. Switch to Kimi K2.7 Code or GPT-5.4 Mini for high-volume simple flows where cost dominates.

Is Sonnet 4.6 the same as Sonnet 4.5?

No. 4.6 is the newer generation in the Claude 4 family with better long-context behaviour and adaptive thinking. The vendor pricing per token is identical.

Availability of Claude Sonnet 4.6 on Okou

Claude Sonnet 4.6 is no longer offered on Okou. See Claude Sonnet 5.5 for a currently supported alternative. This page retains historical model information, not current Okou pricing.