Live Prices
Markets

Anthropic Debuts Claude Sonnet 5.5 With Lower Task Costs and High Token Variance

TheCryptoDesk Editorial · 3m read
Anthropic Debuts Claude Sonnet 5.5 With Lower Task Costs and High Token Variance

AI developer Anthropic released Claude Sonnet 5.5 on September 28, matching the pricing of Sonnet 5 at $2 per million input tokens and $10 per million output tokens. The launch arrives as major technology firms ramp up artificial intelligence deployments, parallel to efforts like Michael Dell and Nvidia launching an open agent safety platform.

Key takeaways from the Sonnet 5.5 release include:

  • Pricing & Speed: Retains $2/$10 per million tokens while operating over 30% faster with up to 30% cost savings per standard task.
  • Benchmark Scores: Achieved 70.6% on Terminal-Bench 4.0 in internal tests, outperforming Opus 5.5 (66.4%) and Sonnet 5 (10.3%).
  • Third-Party Index: Artificial Analysis rated Sonnet 5.5 at 56 on its Intelligence Index, placing it 2 points behind Opus 5.5 at maximum effort.
  • Token Usage: Used ~193k output tokens per Intelligence Index task at max effort—60% higher than Opus 5.5 and 7 times the volume of GPT-6 Astra.

Benchmark Results and Model Comparisons

Anthropic introduced Sonnet 5.5 as the second release in its Claude 5.5 model family, following the September 22 launch of Opus 5.5. On that same day, OpenAI launched GPT-6 Sol and Luna, though OpenAI subsequently shelved its planned GPT-6.1 Astra model after CNBC confirmed on Monday that it failed to meet internal safety standards. Anthropic noted that Sonnet 5.5 does not expand its overall capability frontier, focusing alignment reviews on misuse risks and adding cyber safeguards and anti-distillation classifiers. The smaller Claude Haiku 5.5 is expected to follow in the coming weeks.

While Anthropic described Sonnet 5.5 as optimal for everyday tasks, bug fixes, and document creation, its coding benchmarks were strong. Anthropic reported a 70.6% score on Terminal-Bench 4.0. Independent benchmarking firm Artificial Analysis—which ranked Grok 4.7 fourth overall—recorded 64% for Sonnet 5.5 on Terminal-Bench 4.0, compared to 60% for both Opus 5.5 and GPT-6 Astra. Sonnet 5.5 also performed on par with Opus 5.5 on knowledge-work evaluations such as GDPval-AA.

Discrepancies in Token Volume and Enterprise Feedback

Despite reaching intelligence metrics close to Opus 5.5, Artificial Analysis found that Sonnet 5.5 required substantially higher token usage at maximum effort. The firm reported that the model consumed ~193k Output Tokens per Intelligence Index Task, roughly 60% above Opus 5.5 and Sonnet 5, and approximately 7 times the max-effort output of GPT-6 Astra. Artificial Analysis noted it tested a pre-release build containing a structured-output bug that Anthropic has since resolved.

In contrast, early enterprise feedback showed enhanced token efficiency on domain-specific workloads. Joe Poirier, Senior AI Engineer at Balyasny Asset Management, reported positive results across 2,441 finance tasks covering Q&A, extraction, analysis, and forecasting. Poirier noted that Sonnet 5.5 scored ahead of Sonnet 5 while using about 121k tokens per answer, down from 497k tokens consumed by Sonnet 5.

Why It Matters

The contrasting benchmarks between enterprise evaluations and stress tests highlight how reasoning depth alters model economics. While high-effort multi-step reasoning can drastically increase output token counts and operating costs, structured enterprise tasks display significant efficiency gains over older models. As businesses integrate agentic workflows, navigating these token variations will be critical for managing infrastructure expenses.

Read next