You're paying for the subscription every month, and yet every time you sit down to work you hesitate: which model do I actually run this on? Now that Opus 5 makes it three, the benchmark charts do even less to settle it.
So I ran the same prompts through all three and priced out token costs and usage limits alongside the results. This isn't a leaderboard — it's about where each model earns its keep. Quick note on the lineup: Opus 5 and Fable 5 are Claude models, and Sol is the model behind ChatGPT 5.6. Spoiler: the three differ less in raw ability than in temperament.
- Hard, messy work: Fable 5 first, then Opus 5. The gap between them is smaller than you'd think.
- Fast, repetitive work: Sol is still the light, quick option.
- Opus 5 runs at exactly half Fable 5's price and draws from a separate usage pool — so there's one more value play on the table now.
Three models, and the difference is personality, not IQ
I gave each one the same three builds on a shared theme: a game, a short story, and an interactive stats page. Conditions matched as closely as I could make them. I turned off the option that lets a model check and repair its own output, and I ran each model exactly once at its default top setting. The question was simple: how good is the first take?

The surprise was how much Opus 5 and Fable 5 look like twins. Same texture in the finished work, same instincts about how to approach a problem, even a similar sense of style — I honestly couldn't tell which was which from the output alone.
Sol clearly went its own way. Right idea, execution a half-step short of finished — and that same gap showed up across all three builds.
So who actually handles the hard stuff?
On the game prompt, all three got the logic running on the first try. It was a simple structure — forecast the flood, run the drainage pumps, send a rescue signal, all inside six turns — so logic was never going to be the differentiator.
The difference showed up in the finish. On Sol's build, the arrow pointing to the safe route sat slightly off-target, and a line that should have connected to a circle floated detached from it. Not major bugs, but it read like nobody double-checked the render. Opus and Fable nailed the details — animation alignment, how elements snapped together — on the first pass.
My working theory: first-draft quality is decided less by raw model intelligence than by how complete the harness behind it is — the layer that checks its own output and fixes it. Sol repairs these same issues just fine if you follow up with a second prompt. The one-shot is where the Claude models pulled ahead, and that's the whole takeaway here. Turn Sol's verification setting up and most of the gap closes.
Writing in another language? Claude doesn't lose that round anymore

It used to be that if I needed output in Korean, I just reached for GPT. Claude's sentences landed a half-step behind — a little stiff. Non-English output was a GPT job by reflex, at least for me.
This time I handed the same short-story brief to all three. Every version came back natural enough that ranking them felt arbitrary. Awkward grammar, that translated-from-English cadence — basically gone.
Where they split was atmosphere, of all things. The opening scene has the power cut out and the emergency lights kick on, and how densely each model held that mood is where the daylight was. Subjectively, Fable 5's prose flowed best, with Opus 5 a hair behind. One thing is clear either way: if you've been defaulting to GPT for non-English writing, that rule has expired.
Web and interaction work is where they really split
The interactive stats page drew the sharpest line. Numbers change when you mouse over them, values refresh when you click a sensor — the kind of dashboard pattern that's everywhere right now.
Opus and Fable both implemented hover-based interaction naturally, and the dots and lines on the timeline lined up exactly. Sol went with click instead of hover, and its timeline points were misaligned. Outside of mobile, hover is the far more natural pattern on the web today.
That habit of missing on precision UI — where lines and coordinates have to land exactly — is the same pattern I saw in the game. Which makes it a tendency, not a fluke.
"Half price" — I built the table to check
"Half price" is such a worn-out marketing line that my first instinct is to doubt it. So I dropped the published rates straight into a table. These are the rates as published in July 2026.
| Model | Input (per 1M tokens) | Output (per 1M tokens) |
|---|---|---|
| Opus 5 | $5 | $25 |
| Fable 5 | $10 | $50 |
The pricing is unambiguous. Opus 5 costs exactly what the previous Opus did, and exactly half of Fable 5, the tier directly above it. On the hardest Cursor Bench tier it lands within 0.5% of Fable 5's top score while costing half as much per task — of course you pull the calculator back out.
Speed, funny enough, depends on the job. Opus was a little quicker on code and UI work, while Fable finished the writing tasks first. Not something you can flatten into one line.
| Criterion | Opus 5 | Fable 5 | Sol (ChatGPT 5.6) |
|---|---|---|---|
| Polish on hard tasks | Excellent | Excellent | Good, rough at the edges |
| Non-English prose | Natural | Smoothest | Natural |
| Precision UI & alignment | Strong | Strong | Frequently off |
| Relative price | Low (about half) | High | Middle |
| Sweet spot | Value plus complex work | Top-tier quality | Fast, simple jobs |
The real headline: the usage limits are separate
Honestly, the thing that hit hardest wasn't the performance or the price. It was that Opus 5 draws from a usage pool separate from the tier above it.
Before, burning through your top-model credits basically ended your workday. Now you can max those out and Opus 5 keeps running from its own bucket. If you're the type who squeezes every dollar out of a monthly plan, that's more practical than any benchmark number.
Reasoning effort comes in tiers. Here's how I'd feel it out:
- High — the value zone. Most coding and automation work is perfectly happy right here.
- Extra high / Max — save it for the jobs that need that last push: complex 3D, long refactors.
One more thing. The previous top-tier model would trip a safety classifier whenever a prompt got even slightly ambiguous, and the work would stall out mid-run. This release reportedly cut those interventions by 85% compared to before. A workflow that doesn't stop halfway matters more than it sounds like it should.
How far a few lines of prompt get you now

Give it a few lines — "a ranch tycoon where you raise cows, milk them, and sell the milk" — and you get a running demo: 3D cows standing around, a milk gauge filling up, cash stacking as you ship. That used to be a few days of work. Now one paragraph gets you the shape of it.
Same story for a site where elements move in 3D as you scroll, or a web app that turns an uploaded photo into a merch-box mockup. We're in the stretch where the resolution of your idea is the resolution of the output.
Models now generate the images and video too

The fun development lately is that it doesn't stop at code — the model generates the images and video that go inside the build, right there in the same session.
Two ways to do it. In the web interface, you connect image and video generation tools through an MCP connector. If you'd rather have the files land straight on your own machine, you paste a CLI command into your terminal and hook it up with a single login.
So what should you actually run?

- If tokens are tight — run everything on Opus 5 and switch on the higher verification setting only for the complex jobs.
- If you're GPT-only right now — it's worth adding Claude this billing cycle or next. With the non-English gap gone, the day-to-day feel is genuinely different.
- Ranked by difficulty — hard work: Fable → Opus → Sol. Fast and simple: Sol → Opus → Fable. If Sol's fast mode never clicked for you, dropping Opus into that slot works fine.
If you only do one thing today, take a task you actually run often and put the same prompt through all three. One result in your own hands beats a hundred lines of benchmarks. The best model isn't the one at the top of the chart — it's the one that matches the rhythm of your work.