Claude Sonnet 4.6: historical model reference
1M tokens · Text / Vision / Code · Prompt cache
Okou no longer runs Claude Sonnet 4.6. This page is kept as a reference for its specs, pricing and benchmarks. For the same kind of work, use Claude Sonnet 5.5.
See Claude Sonnet 5.5Vendor list price is $3 / $15 per 1M tokens, with cached input dropping to $0.30 / 1M. Reach for Opus only when Sonnet visibly fails on the hardest reasoning, and for Kimi K2.7 Code or GPT-5.4 Mini when unit cost dominates.
What is Claude Sonnet 4.6?
February 2026 (Claude 4.6 generation)
Claude Sonnet 4.6 sits in the middle of Anthropic's Claude 4 family. It is the workhorse model designed to handle the full breadth of typical agent work. Multi-tool routing, code edits, long-running conversations, and structured-output tasks. Without the cost premium of Opus.
What's notable about Claude Sonnet 4.6
Headline architecture and capability features.
Sonnet 4.6 ships with the 1M-token context window at standard pricing, adaptive thinking inherited from Opus 4.6, and prompt caching that bills cached input at one-tenth the input rate. It accepts multimodal input across text, vision, and code.
Specs at a glance
Claude Sonnet 4.6 benchmarks
Sonnet 4.6 sits roughly 3 to 4 percentage points behind Opus 4.6 on Anthropic's headline coding benchmarks while being three to five times cheaper at the vendor level. The typical Opus/Sonnet trade-off.
Claude Sonnet 4.6 pricing
Historical provider list prices, per 1M tokens. These reference figures are not current Okou charges.
How Claude Sonnet 4.6 behaves in practice
Observed behaviour from production agent runs.
Tool routing
Best-in-class tool-routing accuracy at this price. On multi-tool flows across Slack, GitHub, Linear, and Notion, Sonnet 4.6 picks the correct tool with the correct arguments more reliably than any model below ×2.
Long-context coherence
Coherent across 100K+ token transcripts. Drops below Opus 4.7 only on the longest, most adversarial runs.
Speed
Faster than Opus and slower than Kimi K2.7 Code. The right speed/quality balance for production agents.
Best agent tasks for Claude Sonnet 4.6
The Slack agent that knows where things live
Triages incoming questions, follows up on stalled threads, posts status updates, and answers search-style queries ("who's owning the auth refactor?"). Sonnet's tool-routing accuracy means the right tool gets called with the right arguments on the first try, even when the request is ambiguous, so the agent feels reliable instead of flaky.
The PR review agent that doesn't drown in noise
Sonnet handles the bulk of code-aware work — PR review, test scaffolding, refactor suggestions, bug bisection — without leaving stylistic comments that nobody asked for. The 1M-token context window lets it pull in the related files and prior reviews when it matters, and you only escalate to Opus 4.7 for the patches Sonnet visibly struggles with.
The research agent that makes 20 tool calls in a row
GitHub plus Linear plus Notion plus the web, stitched together across twenty-plus tool turns to answer a question like "why did this customer churn last quarter?" Sonnet keeps the goal in view across the whole chain at a fraction of Opus's cost, which is what makes it sustainable for everyday research as opposed to one-off deep dives.
The customer-support assistant with a stable system prompt
Long conversation histories, frequent tool calls into the CRM, the same hefty system prompt and tool schema on every turn. Sonnet's prompt caching turns that fixed prefix into a fraction of the input cost after the first call, which is what keeps per-conversation cost flat as volume grows.
When to skip Claude Sonnet 4.6
Skip Sonnet 4.6 on the hardest reasoning steps where it visibly drops instructions and you should escalate to Opus 4.7, on bulk classification at high volume where GPT-5.4 Mini is the cheaper supported bulk option, and on latency-critical micro-replies where Kimi K2.7 Code is meaningfully faster.
Claude Sonnet 4.6 vs other models
Claude Sonnet 4.6 vs Claude Opus 4.7
Sonnet 4.6 is ×1; Opus 4.7 is ×2. Sonnet handles most agents; Opus is the upgrade when reasoning depth matters more than throughput. Many teams use Opus as the planner and Sonnet as the worker.
Claude Sonnet 4.6 vs GPT-5.4 Mini
GPT-5.4 Mini is the cheaper OpenAI-side bulk option. Use Sonnet when tool-routing reliability matters more; use Mini for high-volume pre-filtering and simple steps that do not need Sonnet-class routing.
Frequently asked questions
What is Claude Sonnet 4.6's context window?
1 million tokens with up to 64K tokens of output per response.
Does Sonnet 4.6 support image input?
Yes. It's multimodal. Text, code, and images.
When should I switch off Sonnet 4.6?
Switch to Opus 4.7 if Sonnet visibly drops the goal on long agent loops or fails on hard code edits. Switch to Kimi K2.7 Code or GPT-5.4 Mini for high-volume simple flows where cost dominates.
Is Sonnet 4.6 the same as Sonnet 4.5?
No. 4.6 is the newer generation in the Claude 4 family with better long-context behaviour and adaptive thinking. The vendor pricing per token is identical.
Availability of Claude Sonnet 4.6 on Okou
Claude Sonnet 4.6 is no longer offered on Okou. See Claude Sonnet 5.5 for a currently supported alternative. This page retains historical model information, not current Okou pricing.

