Skip to content
KB.consultingKB.consulting
Stack8 min read

Claude Opus 4.7: Honest Notes After 1,500+ Sessions at Our Studio

After two weeks of intensive use, our honest take on Opus 4.7: a real jump in coding benchmarks, a new tokenizer that raises the bill, and when migrating from 4.6 is worth it.

Cover for the Claude Opus 4.7 article — large numeral typography with a summary of benchmarks and API pricing
Cover for the Claude Opus 4.7 article — large numeral typography with a summary of benchmarks and API pricing

This is our second week using Claude Opus 4.7 intensively at the studio. We're past 1,500 sessions, from restoring a three-year-old Laravel codebase to helping write dozens of marketing pages for client products. I'm writing this as a developer who uses it hands-on every day — not a press-release summary, not a cherry-picked demo.

The short version: it's a meaningful upgrade, but it isn't free. The "isn't free" part is what generic reviews most often miss.

When it shipped and what's in it

Anthropic released Opus 4.7 on April 16, 2026. Its official positioning: the most capable generally available model, at exactly the same API price as Opus 4.6 — $5 / $25 per million tokens for input and output. It's available right away through the Claude API, Amazon Bedrock, Vertex AI, and Microsoft Foundry.

What's interesting isn't the list price, but the combination of three things:

  1. A real jump in coding benchmarks, not a 0.3% claim on an arbitrary number.
  2. A new tokenizer that can raise the cost per task, even though the per-token price hasn't changed.
  3. A change in tone you notice immediately in daily use — more direct, with far less filler.

Benchmarks: what matters for developers

I'll keep it brief. These are the official numbers relevant to our work:

  • SWE-bench Verified: 80.8% → 87.6% (up about 7 points from 4.6)
  • SWE-bench Pro: 53.4% → 64.3% (up about 10 points)
  • Agentic multi-step reasoning: ~14% more accurate, with roughly a third fewer tool errors
  • Context window: still 1 million tokens, with a 128k maximum output
  • Image input: maximum size raised to 2,576 px / 3.75 MP
  • A new thinking level: xhigh, sitting between high and max

SWE-bench is a benchmark that asks the model to resolve real issues from open-source repos. The Pro version uses larger tasks and is designed to be hard to game by a model that simply memorized its training data. The 10-point jump on SWE-bench Pro is the most relevant one for us — it's close to what we ask the model to do every day: understand a large repo, make multi-file changes, keep the tests green.

Vellum and several other independent teams that had published comparison numbers at the time of writing report results consistent with Anthropic's claims. To me that's a healthy signal — usually when official benchmarks look too polished, some independent party has very different data. Not this time.

What we felt immediately in our work

1. Multi-file refactors in one pass, not ping-pong

We have an internal e-commerce project (about 84,000 lines of Laravel + Inertia + Vue 3) that had to be migrated from Spatie Permission v5 to v6, with several conflicting changes in middleware. On Opus 4.6, a task like this usually took 3–5 rounds of fixes — the model gets something wrong once or twice, we feed the error back, and repeat.

On 4.7 the flow is much more direct. For a similar migration, it now lands in 1–2 rounds on average. That doesn't mean the model is never wrong — but when it is, the mistakes are smaller in scope and more honest ("I'm not sure file X is still used, please check first") rather than a convincing-sounding fix that's actually a guess.

This matches Anthropic's claim of "a third fewer tool errors". In long agentic workflows, the difference compounds. Tasks that used to take 30 minutes of back-and-forth now often finish in 8–10 minutes.

2. Long context that's finally actually used

The 1-million-token window has existed since 4.6, but honestly we rarely used it anywhere near the limit. On 4.7, retrieving information from the middle of a long context feels much more consistent. I tried loading the entire tartil.id backend codebase (~310,000 tokens including migrations + tests) and asked for a detailed answer about the validation rules in the memorization module.

The result: the model cited the right file paths and the relevant lines, and flagged two small bugs we hadn't noticed ourselves. On 4.6, a similar request would sometimes "forget" files sitting in the middle. Not every time, but often enough that I was reluctant to use the full window for critical work. That hesitation has dropped a lot.

3. Vision that's finally practical, not just a demo

The maximum image resolution went up to 3.75 megapixels. For us — we review UI screenshots from clients all the time — the difference is real: small text in Figma mockups is finally readable without manual cropping.

One routine use at the studio: paste a screenshot of a Filament panel error and ask the model to point out what's wrong. On 4.6 we often had to crop first so the error text was legible. On 4.7, a full-size screenshot (2,560×1,440) is read in full — even an 11-pixel stack trace.

4. xhigh: a new option between high and max

Anthropic added a new thinking level called xhigh, between high and max. For us it's the sweet spot for work that needs deep reasoning without paying the latency of max.

A concrete example: reviewing the architecture behind a bug in a crash report for one of our products. high sometimes misses a subtle race condition; max gets it right but takes over 40 seconds. xhigh finds the answer in 18–22 seconds with reasoning quality almost on par with max. For interactive debugging, that's a big help.

The parts you have to be honest about

This is the section reviews that just repeat the press release tend to skip.

The new tokenizer: your bill can rise 5–35%

Anthropic replaced the tokenizer in Opus 4.7. The result: the same text can become more tokens — between 1.0x and about 1.35x compared with 4.6. How much depends on the type of content. Heavily affixed Indonesian ("memperhitungkan", "ketidakpastian"), code with long variable names, and technical documents in table format are the categories that grow the most.

On our own workloads, the average token increase was ~12% for English-language coding work and ~22% for conversations mixing in Indonesian. So even though the per-token price is the same, our total monthly bill went up by about 15–18% after moving fully to 4.7.

Practical advice: if you're evaluating 4.7 for production, don't benchmark from a single sample prompt. Measure real token consumption over several days. The savings from more accurate results (fewer retries) may or may not offset the token increase — it depends on your workload.

A change in tone: more direct, sometimes to the point of cold

Opus 4.7 speaks more directly. Fewer "Great question!"s, fewer emoji, fewer hedging words that were really empty. For me personally, that's a feature — I don't need a model to cheer me on, I need a correct answer.

But non-developer teams at our clients, who use Claude to draft emails or help write copy, report a different feel. Some say the model is "less friendly". If you're building an end-user-facing product on top of the API, this change in tone needs to be re-tested with real users. 4.7's default suits power users; it may not suit every audience.

More frugal about spawning subagents

In agentic workflows with tool delegation, 4.7 tends to spawn new subagents less often than 4.6. The intent is good — there were plenty of cases on 4.6 where it spawned too many subagents and burned tokens without making progress.

The side effect: if your work genuinely needs parallelization (say, checking 12 independent files), you may need to be more explicit in the prompt. We updated our internal system prompts for genuinely parallel work: "carry out the following steps in parallel subagents." That wasn't needed before; now it is.

Migrating 4.6 → 4.7: our short guide

For teams already running Opus 4.6 in production, here's the migration order we used:

  1. Days 1–2: run in parallel on staging. Run 4.6 and 4.7 side by side on the same workload. Collect token-consumption logs plus output quality.
  2. Days 3–5: audit the tone. If your product is used directly by end users, run a blind A/B test with some real users. 4.7's default tone is cooler; adjust it through the system prompt.
  3. Week 1: review your subagent prompts. Workflows that used to spawn subagents automatically may need explicit instructions on 4.7.
  4. Week 2: recalculate costs. Don't just read the first week's bill — cache-hit ratios and retry frequency take time to settle.

For teams building coding agents — especially long, multi-file tasks that need planning — the migration is clearly worth it. Accuracy goes up, retries go down, and changes are tighter in scope.

For teams that only use LLMs for single-turn classification or short extraction, the 4.7 jump isn't worth the cost yet. Sonnet 4.6 or Haiku 4.5 remain the more economical choice for that kind of work.

Verdict: when we use 4.7, and when we don't

After two weeks of intensive use, here are our current defaults:

  • Multi-file coding agents, refactoring, debugging: Opus 4.7 with xhigh. The performance-to-cost sweet spot.
  • Long-context analysis (codebase audits, long documents): Opus 4.7 with the full window.
  • Vision for UI review: Opus 4.7. The higher resolution is too practical to go back to 4.6.
  • Fast classification, autocomplete, short summarization: still Sonnet 4.6 or Haiku 4.5.
  • End-user-facing content with a friendly tone: test first; you may need to adjust the system prompt — or stay on 4.6 until user feedback is positive.

Opus 4.7 isn't a revolution. It's a well-targeted iteration. If you do serious software development — coding agents, large refactors, automated code review — 4.7 is useful from day one. If your workload is simple, hold off — the added complexity on your bill isn't worth it for work that doesn't need the most advanced model.

One thing is certain: in my second week on 4.7, I don't want to go back to 4.6 for complex work. That's my bar for calling a model upgrade "real".


Note: the token-consumption and latency figures in this article come from KB Consulting's internal workloads during the first 14 days Opus 4.7 was available. Your results may differ depending on the type of work and language mix. Benchmark sources: Anthropic, Vellum, and independent analysis by Verdent.


Written by Tim KB Consulting

Discuss this topic on WhatsApp

Start now

Got an idea?
Let's validate it first.

30 minutes, free, no commitment. If we click, we move on to a validation session. If not, we'll recommend someone who's a better fit — still at no cost.