Kiyoro on X: OPUS 5.5 IS MAKING YOUR AGENT DUMBER RIGHT NOW AND IT WILL NEVER TELL YOU. JEV C
@0xKiyoro@0xKiyorolisted 47m agoreviewed by JevOPUS 5.5 IS MAKING YOUR AGENT DUMBER RIGHT NOW AND IT WILL NEVER TELL YOU. JEV CAUGHT IT. The default effort on Opus 5.5 dropped from high to medium. If your agent never set effort explicitly, it is now reasoning at a lower setting than the one you tested it on. No error, no warning, every health check still green. Jev sets effort on every single message, so the requests that went through it never changed. The ones that relied on the default did. That gap is how it showed up. The rest of the migration is loud. Each of these returns a 400 on the first call: > thinking: disabled, now rejected. Drop the field and set effort > tool_choice any and tool, removed. Use auto + strict > computer_20251124, retired. Use computer_toolset_20260801 > editing above a thinking block, rejected. Append only, or drop_block You will find all four in five minutes. The effort change you will not find at all. The cache has its own quiet trap. Hop Opus 5.5 to Sonnet 5 and back and a session that cost 3.32 costs 4.36, +31%, because Sonnet cannot read Opus's reasoning. Change effort at the top of a request or switch fast to standard mid-session and the cache is gone again. Jev picks speed once on turn one and sets effort per message, so it survives. The upgrade is still worth it. $4 / $20 per 1M instead of $5 / $25, cache reads at $0.20 instead of $0.50, and 66.4% on Terminal-Bench 4.0 against 57.9% for GPT-6 Astra. Write effort into every request before you touch the model ID. The default stopped meaning what you think. ❤️ 7 likes on X
- Author
- @0xKiyoro
- Use case
- Benchmarks & Evals
- Added
- 2026-09-25
- Cost
- $4
All figures come from the author. Check the source before you quote them.
More in Benchmarks & Evals
- Morgan on X: It has been a really interesting experience to build an eval suite for System On▲ 0x.com
- silentguy on X: Grok Bot does the job, Jev decides where the job goes next▲ 0x.com
- spect on X: The founder of Jev just dropped a 1-hour masterclass on how Jev actually works▲ 0x.com
- Dain on X: A beautiful pattern can still be noise.▲ 0x.com
- Utkarsh Maheshwari on X: Is the Jev hype real,▲ 0x.com
- OpenMed on X: Four typed questions across four authored fictional notes, labels written before▲ 0x.com
Jev guides for this use case
- Using Jev as a judge for evalsGrading with a decision model instead of a prose-writing judge, and why calibrated confidence is the real prize.
- How Jev sorts a build into one of 21 use casesThe 21 criteria Jev classifies against, published in full, plus what the reviewer sees and how ambiguity is handled.
- Jev statistics: latency, cost, and this directory's own numbersPublished benchmarks with their caveat attached, plus live directory figures that update automatically.
Bid history
No bids yet — the first one takes this project straight to the spotlight.