BuiltOnJev
x.com

Ranjan Kumar on X: ๐–๐ก๐ž๐ซ๐ž ๐‰๐ž๐ฏ ๐๐ž๐ฅ๐จ๐ง๐ ๐ฌ ๐ข๐ง ๐š๐ง ๐€๐ ๐ž๐ง๐ญ ๐‡๐š๐ซ๐ง๐ž๐ฌ๐ฌ (๐๐จ๐ญ ๐š

@ranjankumar@ranjankumar@ranjankumarlisted 45m agoreviewed by Jev

๐–๐ก๐ž๐ซ๐ž ๐‰๐ž๐ฏ ๐๐ž๐ฅ๐จ๐ง๐ ๐ฌ ๐ข๐ง ๐š๐ง ๐€๐ ๐ž๐ง๐ญ ๐‡๐š๐ซ๐ง๐ž๐ฌ๐ฌ (๐๐จ๐ญ ๐š ๐Œ๐จ๐๐ž๐ฅ ๐’๐ฐ๐š๐ฉ) Your agent harness needs a decision model. Where you place it depends on one brutal fact: Jev's ordering is trustworthy. Its confidence numbers are not. Most placement guides skip this distinction entirely. They tell you to pick a threshold and execute above it. That works if you actually have a probability. You might not. ๐“๐ก๐ž ๐œ๐จ๐ซ๐ž ๐ฉ๐ซ๐จ๐›๐ฅ๐ž๐ฆ: A model can rank beautifully and lie about magnitude simultaneously. When you gate execution on if confidence >= 0.85: approve_transfer, you are trusting a number that was never calibrated to your prevalence, your cost ratio, or your queue. Move the threshold up and you trade recall you never measured for precision you cannot state. You have no idea how far along an unmarked axis you moved. ๐“๐ก๐ซ๐ž๐ž ๐๐ž๐œ๐ข๐ฌ๐ข๐จ๐ง๐ฌ, ๐ญ๐ก๐ซ๐ž๐ž ๐๐ข๐Ÿ๐Ÿ๐ž๐ซ๐ž๐ง๐ญ ๐ง๐ž๐ž๐๐ฌ: Routing and ranking consume ordering only - which answer is best matters, the number beside it does not. Gated execution consumes magnitude - the threshold is your boundary and it must land on a calibrated scale. Relative logic like if top1 - top2 < 0.1: escalate consumes differences - and rescaling the score axis will flatten margins unevenly, breaking your code. The sepsis alert systems of 2020 learned this hard way. Michigan switched off alerts when COVID shifted patient prevalence beneath a fixed threshold. No weights changed. The denominator moved, the promise broke silently, and nurses drowned in false alarms. ๐“๐ก๐ž ๐Ÿ๐ข๐ฑ ๐ข๐ฌ ๐ฌ๐ž๐ช๐ฎ๐ž๐ง๐œ๐ž, ๐ง๐จ๐ญ ๐ญ๐ฎ๐ง๐ข๐ง๐ : First question: is this number admissible as a probability at all? Run a calibration study on a few hundred labelled cases from your own queue. Samuel Sacco's measurements from 18 September and Adil Muhammad Pervez's 8,000 judgments both reach the same conclusion: fit your own map. The weights are shared across every account by design - there is no per-customer adaptation - so a deployer-side calibration map is your only mechanism. ๐’๐ž๐œ๐จ๐ง๐ ๐ช๐ฎ๐ž๐ฌ๐ญ๐ข๐จ๐ง: given that the number is calibrated, where does your cost ratio put the line? That part of the familiar advice still works. Run them backwards and you are tuning a dial with no markings. Read the full analysis on calibration, harness placement, and where Jev actually fits: https://ranjankumar.in/jev-system-one-model-agent-harness-placement Follow for more practitioner insights on agentic AI systems and production AI engineering. #AgentiveAI #AIEngineering #SystemOne #Calibration #MLOps #DecisionModels #HarnessEngineering #Jev โค๏ธ 0 likes on X

Author
@ranjankumar
Use case
Agents & Browsers
Added
2026-09-25

All figures come from the author. Check the source before you quote them.

Bid history

No bids yet โ€” the first one takes this project straight to the spotlight.