
x.com
Ram Vinjamuri on X: (1/5) trying out jev this week with all the hype. kept it apples to apples vs th
@RamV2003@RamV2003listed 46m agoreviewed by Jev(1/5) trying out jev this week with all the hype. kept it apples to apples vs the fast models it competes with: llama 3.1 8b and deepseek v4 flash same labelled inputs. jev: 235ms median, 82% accurate. llama: 497ms, 68%. deepseek flash was accurate but ~2s genuinely cool. then some odd emergent behaviour: no position or label invariance ❤️ 7 likes on X
- Author
- @RamV2003
- Use case
- Benchmarks & Evals
- Added
- 2026-09-25
- Time
- 235ms
All figures come from the author. Check the source before you quote them.
More in Benchmarks & Evals
- Morgan on X: It has been a really interesting experience to build an eval suite for System On▲ 0x.com
- silentguy on X: Grok Bot does the job, Jev decides where the job goes next▲ 0x.com
- spect on X: The founder of Jev just dropped a 1-hour masterclass on how Jev actually works▲ 0x.com
- Dain on X: A beautiful pattern can still be noise.▲ 0x.com
- Utkarsh Maheshwari on X: Is the Jev hype real,▲ 0x.com
- OpenMed on X: Four typed questions across four authored fictional notes, labels written before▲ 0x.com
Jev guides for this use case
- Using Jev as a judge for evalsGrading with a decision model instead of a prose-writing judge, and why calibrated confidence is the real prize.
- How Jev sorts a build into one of 21 use casesThe 21 criteria Jev classifies against, published in full, plus what the reviewer sees and how ambiguity is handled.
- Jev statistics: latency, cost, and this directory's own numbersPublished benchmarks with their caveat attached, plus live directory figures that update automatically.
Bid history
No bids yet — the first one takes this project straight to the spotlight.