GPT 5.5: historical model reference
400K tokens · Text / Vision / Code · Prompt cache
Okou no longer runs GPT-5.5. This page is kept as a reference for its specs, pricing and benchmarks. For the same kind of work, use GPT 6 Luna.
See GPT 6 LunaGPT-5.5 is the model you reach for when the work needs both deep reasoning and reliable tool use: orchestrating multi-step agent loops, code edits that have to land first try, and computer-use workflows that span many GUI actions. Vendor benchmarks (SWE-bench Verified, AIME 2025, GPQA Diamond) put concrete numbers on the gains over GPT-5.4.
What is GPT-5.5?
April 2026 (successor to GPT-5.4)
GPT-5.5 is the flagship of OpenAI's GPT-5 generation, released in April 2026 as the recommended upgrade from GPT-5.4. OpenAI frames it as a step-change improvement on agentic tool use and computer-use tasks rather than a refresh on the surface API. The 400K-token context window and reasoning_effort parameter introduced with GPT-5 carry over unchanged, so existing Codex agents drop in without rewrites.
Independent leaderboards (Artificial Analysis, Vellum) corroborate the relative ordering against GPT-5.4 and place GPT-5.5 within a few points of Claude Opus 4.7 on most agentic-coding tasks. Absolute numbers shift weekly and OpenAI itself has flagged training-data contamination on SWE-bench Verified across frontier models. Treat the public scores as directional rather than authoritative; the structured behavioural differences (tool-call accuracy, computer-use reliability, first-attempt patch quality) are the more durable signal.
What's notable about GPT-5.5
Headline architecture and capability features.
GPT-5.5 keeps the 400K-token context window from GPT-5.4, billed at standard input pricing across the entire window. It supports the reasoning_effort parameter at four levels (minimal, low, medium, high), prompt caching where cached input bills at one-tenth the input rate, and the Responses API surface that codex CLI uses by default. Tool-use, structured outputs and computer-use are unchanged from 5.4. Inputs are multimodal across text, vision and code; the model has no native image generation (use the Images API for that).
Specs at a glance
GPT-5.5 benchmarks
Vendor-reported scores from OpenAI's GPT-5.5 release materials, with deltas shown against the public GPT-5.4 numbers. Independent reviews place 5.5 within a few points of Claude Opus 4.7 on agentic-coding tasks. Treat absolute percentages as directional; OpenAI has flagged training-data contamination on SWE-bench Verified across all frontier models.
GPT-5.5 pricing
Historical provider list prices, per 1M tokens. These reference figures are not current Okou charges.
How GPT-5.5 behaves in practice
Observed behaviour from production agent runs.
Tool routing
Lowest rate of mis-routed tool calls in the GPT-5 family. The gap versus 5.4 widens on hard edge cases such as conditional tool selection, deeply nested arguments, and tool calls dispatched after long stretches of reasoning.
First-attempt code edits
Strongest patch quality in the GPT-5 family. The right pick when an agent has to modify code that must keep compiling and passing tests, especially when the patch spans multiple files. Vendor-reported SWE-bench Verified reflects this directly.
Computer use
Materially more reliable than 5.4 on multi-step GUI sequences, which is what the OSWorld delta captures. Reach for it when the agent is driving a browser or desktop app over dozens of steps and the cost of a mid-run derailment is high.
Speed
Slower than 5.4 and noticeably slower than 5.4 Mini. Around 70 tokens/sec at medium effort per Artificial Analysis. Reserve it for the steps that actually need the extra reasoning depth and run lighter tiers in parallel.
Hallucination behaviour
GPT-5.5 carries OpenAI's stricter calibration from the GPT-5 generation and tends to admit uncertainty rather than confabulate, which is the reason production teams keep paying the premium for high-stakes reasoning despite cheaper alternatives like DeepSeek V4 Pro now matching it on benchmarks.
Best agent tasks for GPT-5.5
The orchestrator running a multi-tool plan
Use GPT-5.5 as the planner that breaks a customer's request into ten steps, dispatches each step to a GPT-5.4- or 5.4 Mini-tier sub-agent, and stitches the results back together. Running 5.5 only at the planner layer (and the cheaper tiers everywhere else) costs a fraction of running 5.5 end-to-end, with most of the quality preserved.
The first-try code edits that don't waste a CI run
Ask GPT-5.5 to migrate a 50-file codebase from one ORM to another, refactor a tangled module, or apply a security fix across the repo. The patch applies cleanly on the first attempt more often than any other model in the family, and that's exactly what your CI bill will reflect.
The computer-use agent that has to finish the workflow
When the agent is driving a browser through a multi-step booking flow, a desktop app, or a legacy admin UI, 5.5's stronger OSWorld score translates to fewer mid-run derailments and fewer human takeovers. The premium pays for itself the first time a long session doesn't have to be restarted.
The hard-math or hard-science research step
Drop a competition-grade math problem set or a graduate physics derivation in and 5.5 will work through it without the off-by-one slips you see in 5.4. AIME 2025 and GPQA Diamond pick up exactly this kind of behaviour.
When to skip GPT-5.5
Skip GPT-5.5 on high-volume routine work where GPT-5.4 hits the same quality bar at half the credit cost, on latency-sensitive chat replies where GPT-5.4 Mini is much faster, and on bulk classification or extraction jobs where GPT-5.4 Mini is the cheaper supported bulk option.
GPT-5.5 vs other models
GPT-5.5 vs GPT-5.4
GPT-5.4 is the workhorse default in the GPT-5 family and the right pick for most agents. Promote to GPT-5.5 only when 5.4 visibly fails on hard reasoning, long agentic loops or first-attempt code edits, usually as the orchestrator that delegates downward to 5.4- or 5.4 Mini-tier sub-agents.
GPT-5.5 vs Claude Opus 4.7
Same role in different families: the high-stakes orchestrator and the model you escalate to when the cheaper tier fails. Opus 4.7 has the 1M-token context window and Anthropic's safety profile; GPT-5.5 has stronger computer-use scores and is the natural pick for teams already on the Codex framework. Pick by which framework and ecosystem your existing agents target.
GPT-5.5 vs Gemini 3 Pro
Gemini 3 Pro leads on raw long-context reasoning (2M-token window) and on some multimodal benchmarks. GPT-5.5 leads on agentic coding (SWE-bench Verified, Terminal-Bench) and computer use. Pick GPT-5.5 when the agent edits code or drives a UI; pick Gemini 3 Pro when the workload is heavy document or video understanding.
Frequently asked questions
What is GPT-5.5's context window?
400,000 tokens, with up to 128K tokens of output per response. The full window bills at standard rates.
Can GPT-5.5 handle images?
Yes. GPT-5.5 is multimodal. It accepts image inputs alongside text and code, so screenshot-driven and document-vision agents work natively. For image generation use the OpenAI Images API.
When should I pick GPT-5.5 over GPT-5.4?
When (a) the agent is the planner / orchestrator and decisions cascade, (b) the run is long enough that 5.4 starts mis-routing tool calls, or (c) the output must apply cleanly on the first attempt (code edits, structured payloads, computer-use workflows).
Does GPT-5.5 support prompt caching?
Yes. Cached input bills at $0.50 per 1M tokens — a 10× discount on the cached portion. Worth using whenever your system prompt or tool schema is stable across calls.
Availability of GPT-5.5 on Okou
GPT-5.5 is no longer offered on Okou. See GPT 6 Luna for a currently supported alternative. This page retains historical model information, not current Okou pricing.

