
Dorian Smiley on X: Testing Jev’s accuracy tonight for next best action prediction. The results:
@dsmiley411@dsmiley411listed 2h agoreviewed by JevTesting Jev’s accuracy tonight for next best action prediction. The results: Canonical accuracy: 98.6% Generalization accuracy: 41.5% We ask Jev to predict the next state in a program from the current partial program. The suite has 25 cases and we ran it 20 times. Seven cases are represented in the in context examples. The other 18 are held out cases that require Jev to generalize from those examples. The failures are not uniformly random. Jev generalizes perfectly on some unseen compositions and fails almost deterministically on others. That suggests there may be specific structural boundaries to what it can infer from context. Maybe some of this is prompt design. Maybe it is a capability boundary. We’re testing that now. But next best action prediction is important. A huge amount of software today contains really brittle decision logic: onboarding, payments, claims, revenue cycle, approvals, exception handling, etc. If Jev is a bet on software consuming intelligence, this is exactly the kind of high frequency, high value logic it needs to improve. ❤️ 4 likes on X
- Author
- @dsmiley411
- Use case
- Benchmarks & Evals
- Added
- 2026-09-25
All figures come from the author. Check the source before you quote them.
More in Benchmarks & Evals
- Morgan on X: It has been a really interesting experience to build an eval suite for System On▲ 0x.com
- silentguy on X: Grok Bot does the job, Jev decides where the job goes next▲ 0x.com
- spect on X: The founder of Jev just dropped a 1-hour masterclass on how Jev actually works▲ 0x.com
- Dain on X: A beautiful pattern can still be noise.▲ 0x.com
- Utkarsh Maheshwari on X: Is the Jev hype real,▲ 0x.com
- OpenMed on X: Four typed questions across four authored fictional notes, labels written before▲ 0x.com
Jev guides for this use case
- Using Jev as a judge for evalsGrading with a decision model instead of a prose-writing judge, and why calibrated confidence is the real prize.
- How Jev sorts a build into one of 21 use casesThe 21 criteria Jev classifies against, published in full, plus what the reviewer sees and how ambiguity is handled.
- Jev statistics: latency, cost, and this directory's own numbersPublished benchmarks with their caveat attached, plus live directory figures that update automatically.
Bid history
No bids yet — the first one takes this project straight to the spotlight.