Ranjan Kumar on X: ๐๐ก๐๐ซ๐ ๐๐๐ฏ ๐๐๐ฅ๐จ๐ง๐ ๐ฌ ๐ข๐ง ๐๐ง ๐๐ ๐๐ง๐ญ ๐๐๐ซ๐ง๐๐ฌ๐ฌ (๐๐จ๐ญ ๐
@ranjankumar@ranjankumarlisted 45m agoreviewed by Jev๐๐ก๐๐ซ๐ ๐๐๐ฏ ๐๐๐ฅ๐จ๐ง๐ ๐ฌ ๐ข๐ง ๐๐ง ๐๐ ๐๐ง๐ญ ๐๐๐ซ๐ง๐๐ฌ๐ฌ (๐๐จ๐ญ ๐ ๐๐จ๐๐๐ฅ ๐๐ฐ๐๐ฉ) Your agent harness needs a decision model. Where you place it depends on one brutal fact: Jev's ordering is trustworthy. Its confidence numbers are not. Most placement guides skip this distinction entirely. They tell you to pick a threshold and execute above it. That works if you actually have a probability. You might not. ๐๐ก๐ ๐๐จ๐ซ๐ ๐ฉ๐ซ๐จ๐๐ฅ๐๐ฆ: A model can rank beautifully and lie about magnitude simultaneously. When you gate execution on if confidence >= 0.85: approve_transfer, you are trusting a number that was never calibrated to your prevalence, your cost ratio, or your queue. Move the threshold up and you trade recall you never measured for precision you cannot state. You have no idea how far along an unmarked axis you moved. ๐๐ก๐ซ๐๐ ๐๐๐๐ข๐ฌ๐ข๐จ๐ง๐ฌ, ๐ญ๐ก๐ซ๐๐ ๐๐ข๐๐๐๐ซ๐๐ง๐ญ ๐ง๐๐๐๐ฌ: Routing and ranking consume ordering only - which answer is best matters, the number beside it does not. Gated execution consumes magnitude - the threshold is your boundary and it must land on a calibrated scale. Relative logic like if top1 - top2 < 0.1: escalate consumes differences - and rescaling the score axis will flatten margins unevenly, breaking your code. The sepsis alert systems of 2020 learned this hard way. Michigan switched off alerts when COVID shifted patient prevalence beneath a fixed threshold. No weights changed. The denominator moved, the promise broke silently, and nurses drowned in false alarms. ๐๐ก๐ ๐๐ข๐ฑ ๐ข๐ฌ ๐ฌ๐๐ช๐ฎ๐๐ง๐๐, ๐ง๐จ๐ญ ๐ญ๐ฎ๐ง๐ข๐ง๐ : First question: is this number admissible as a probability at all? Run a calibration study on a few hundred labelled cases from your own queue. Samuel Sacco's measurements from 18 September and Adil Muhammad Pervez's 8,000 judgments both reach the same conclusion: fit your own map. The weights are shared across every account by design - there is no per-customer adaptation - so a deployer-side calibration map is your only mechanism. ๐๐๐๐จ๐ง๐ ๐ช๐ฎ๐๐ฌ๐ญ๐ข๐จ๐ง: given that the number is calibrated, where does your cost ratio put the line? That part of the familiar advice still works. Run them backwards and you are tuning a dial with no markings. Read the full analysis on calibration, harness placement, and where Jev actually fits: https://ranjankumar.in/jev-system-one-model-agent-harness-placement Follow for more practitioner insights on agentic AI systems and production AI engineering. #AgentiveAI #AIEngineering #SystemOne #Calibration #MLOps #DecisionModels #HarnessEngineering #Jev โค๏ธ 0 likes on X
- Author
- @ranjankumar
- Use case
- Agents & Browsers
- Added
- 2026-09-25
All figures come from the author. Check the source before you quote them.
More in Agents & Browsers
- nicolay on X: Jev is the semantic if in your agent loop.โฒ 0x.com
- Ricker on X: Jev Engineering turns a static agent workflow into a graph that can rewrite itseโฒ 0x.com
- Cartwright on X: Built Cortex, a local MCP server: Claude plans, a fast layer does the clicking.โฒ 0x.com23 s
- Coach Shweta Bajaj on X: What stands out to me in the Jev + Grok Bot setup is the division of work:โฒ 0x.com
- YZ on X: ไฝฟ็จ #Jev ๅไบไธไธชๅกซๅๅค้พไฟกๆฏ็ๆผ็คบใโฒ 0x.com
- dealer.eth on X: I LEAKED THE JEV STACK I RUN MY AGENTS ON, IT CAUGHT 11 OF THEM IN ONE NIGHT BEFโฒ 0x.com311 ms
Jev guides for this use case
- Jev in the agent loop: routing, guardrails, contextRouting, tool-call guardrails and context compaction โ the three places a decision model earns its keep inside an agent.
- Jev vs LLM: a decision model is not a smaller text modelWhy a decision model and a text model aren't substitutes, where the boundary sits, and the cost math.
- What people actually build with Jev, by use caseAll 21 use cases with live counts from real submissions, plus what each one is actually for.
Bid history
No bids yet โ the first one takes this project straight to the spotlight.