How People Are Actually Using Jev (AI Daily Brief)
How People Are Actually Using Jev (AI Daily Brief)
Nathaniel Whittemore's main episode (≈10 days after launch) cataloguing how people are actually using Jev — Typesafe's "System 1" model — beyond the viral visual/game demos. Core claim: Jev is not another LLM but a complementary primitive — fast, near-free snap judgments at volume — and it belongs inside an AI stack next to frontier LLMs, not instead of them. See Judgment Models (System 1 AI) for the concept.
"Because this is at core a new primitive in its ability to apply simple judgment at scale, at speed, and for effectively no money, I think it's going to take some time for us to really figure out just how deeply we can weave this into all sorts of different use cases."
What Jev is
- A judgment model, not a writer. Named after Kahneman's Thinking, Fast and Slow System 1 — fast, instinctive pattern-matching. It won't write code, draft contracts, or make multi-factor nuanced calls.
- The fit test: repeatedly read something → make a small judgment → take a predictable next step. E.g. a file lands in Downloads: a rule says "PDF"; Jev says "invoice, for project X, someone needs to see it."
- Three question types:
- Choice (pick one) — up to 255 options; returns the pick plus a probability for every option. "Which team handles this ticket: billing, tech support, sales, other?"
- Score — a 2–10-level scale, each level described in words; returns position + certainty. "Calm → very angry: how frustrated is this customer?"
- Bernoulli (yes/no) — probability 0–1 the statement is true. "Is this person explicitly asking for a refund?"
- Economics (Typesafe claims): 20–200× faster and 40–400× cheaper than comparable LLM processes; ~$4.20 per million input tokens (transcript reads "4.2"), output tokens free; 70–500 ms per call. Many questions per item run in parallel — the 10th question costs tokens, not time. Typesafe: 13 questions in one call = 12.2× cheaper, 10× faster than one-at-a-time, identical answers. Evidence must fit in <32K tokens.
- Market: reportedly in talks to raise up to $1B at ≥$10B valuation (The Information), up from a $40M seed at $200M.
Six use-case categories
| # | Category | Question it answers | Representative examples |
|---|---|---|---|
| 1 | Analyze what you already have | "What's in this pile?" | 724 live ads × 12 questions in 40s for $0.09; then 723 ads × 30 buyer archetypes = 21,690 stop/scroll calls for $0.22 (framed as hypothesis generator, not data); ~3,300 past X posts × 8 questions vs engagement |
| 2 | Search by meaning | "Which of these match what I mean?" | Zillow listings by architectural style/renovation/freeway proximity (Justine Moore, a16z); clip a 90-min video by described theme in <2s; SEO internal-link map across 586 pages in 45s for ~$0.21 vs Claude Opus 5 getting through 21 pages for $1.43; real-time "AI slop" filter on X; 384 news stories × 15 brands in 25s |
| 3 | Triage what comes in | "What is this and where does it go?" | Live-prioritised inbox (100 emails in 453 ms, ~$0.001, matched the author's own ratings); Downloads-folder invoice filer; live chat moderation; Dub link-shortener malicious-URL flagging trained on 10K past bad domains — "we solved it in 2 hours"; Box incident triage (impact/severity/escalation route); CRM lead priority/readiness/spam |
| 4 | Check work against rules | "Does this meet the bar?" | LangChain grading agent traces — Harrison Chase: "great for evals, especially online evals"; Every planted-error test: Jev 6/7 in 0.35s vs Claude Fable 5.1 7/7 in 8.83s at ~580× the cost; style guide → yes/no questions run on every paragraph |
| 5 | Speed up AI agents | "Which model / effort / skill / tool fits?" | Jev switching GPT-6 reasoning effort mid-task inside Codex (~50% lower cost, faster runs); "Jev skill suggestion" for Claude Code injects only the matching skill → 88% fewer tokens; a harness that learns the job and migrates steps from LLM calls to code — compliance alerts from ~$2.95 → ~$0.25 per alert by alert 10,000 |
| 6 | Respond instantly | "What does this person want right now?" | Smart copy-paste (paste a resume → form fields fill themselves, only confident matches pasted) |
Whittemore's enterprise pick: if you explore only one area, make it triage and routing — especially where many people touch the same leads/customers/contacts. "Most workflows eventually hit the same question: what should happen next?"
Scale changes kind. Checking every sentence for AI-isms is categorically different from one generic LLM review of the whole document — "a difference in kind rather than a difference in scale."
Limits and cautions
- Typesafe's own "bad at" list: multi-step questions (accuracy drops per hop), counting/math/dates (extract facts, compute elsewhere), consistency, reading intent.
- High-stakes domains — hiring, money, security. A score with no reasoning attached is dangerous for ranking candidates. Whittemore's pattern: Jev ranks, then an LLM reviews borderline cases — a multi-model system, not a single judge.
- Four-criteria task-fit test:
- Answers can be written down in advance (categories, yes/no, scale) — not a sentence or calculation.
- Volume — a pile or a stream.
- Low stakes — a wrong answer is cheap or easy to catch.
- Evidence fits as text in <32K tokens.
- Question-writing rules: one judgment per question ("is this a good lead?" is really fit + size + intent…); describe every scale level in words; ask more questions than you think you need (they're cheap); hand-label ~50 items and compare before trusting it in production.
Why this matters for the user
For an enterprise IT org, the load-bearing uses are service-desk ticket routing, incident severity triage, compliance-alert handling, and online evals of production agents — the last one Whittemore flags as likely "core infrastructure" as agent fleets grow. Jev-class models also sharpen the Model Routing story: routing decisions themselves become near-free micro-judgments. And the "LLM-to-code migration" harness is a concrete instance of the cost curve bending with volume.
note Transcription caveats Names are auto-transcribed and unverified: "Matthew Burman" (ad analysis — possibly Matthew Berman, not confirmed), "Ian Nuttall", "Boura", "Burhan", "Vchen", "Daniel Son", "AJ Aspher", "Steven Tay (Dub)", "Marcel Pio (Beyond Code)". Dollar amounts like "4.2", "13", "21" are read as dollars/cents from context. The Jev launch itself isn't independently sourced in this vault yet.
Connections
- Concept · Judgment Models (System 1 AI)
- Entities · Jev · Typesafe · Nathaniel Whittemore · The AI Daily Brief
- Routing & cost · Model Routing · Skills (Claude Code) · Harness (LLM Agents)
- Evals · LLM as Judge · Binary Eval Assertions
- Adjacent thesis · Narrow Agents — narrow, repeated, high-volume judgments are exactly where Jev lives