BuiltOnJev
Dorian Smiley on X: Testing Jev’s accuracy tonight for next best action prediction. The results:
x.com

Dorian Smiley on X: Testing Jev’s accuracy tonight for next best action prediction. The results:

@dsmiley411@dsmiley411@dsmiley411listed 2h agoreviewed by Jev

Testing Jev’s accuracy tonight for next best action prediction. The results: Canonical accuracy: 98.6% Generalization accuracy: 41.5% We ask Jev to predict the next state in a program from the current partial program. The suite has 25 cases and we ran it 20 times. Seven cases are represented in the in context examples. The other 18 are held out cases that require Jev to generalize from those examples. The failures are not uniformly random. Jev generalizes perfectly on some unseen compositions and fails almost deterministically on others. That suggests there may be specific structural boundaries to what it can infer from context. Maybe some of this is prompt design. Maybe it is a capability boundary. We’re testing that now. But next best action prediction is important. A huge amount of software today contains really brittle decision logic: onboarding, payments, claims, revenue cycle, approvals, exception handling, etc. If Jev is a bet on software consuming intelligence, this is exactly the kind of high frequency, high value logic it needs to improve. ❤️ 4 likes on X

Author
@dsmiley411
Use case
Benchmarks & Evals
Added
2026-09-25

All figures come from the author. Check the source before you quote them.

Bid history

No bids yet — the first one takes this project straight to the spotlight.