00 TypeSafe / System One

Jev

A field guide to the model

Presented by Han Cheol Moon

27 September 2026

00 Jev / A field guide to the model

Decisions your code
can act on.

Jev · TypeSafe's flagship model, and the first System One model

A typed question goes in. A typed answer with probabilities comes out.

ONE ANSWER, VERBATIM SHAPE
"department": {
  "type": "choice",
  "choice": "returns",
  "confidence": 1.0,
  "probabilities": {
    "returns":  1.0,
    "shipping": 0.0,
    "billing":  0.0
  }
}

Branch on it. Sort by it. Route with it.

Response shape from the Choice documentation
01 The idea02 The primitives03 Building with it04 The case against05 Hands on
How to read this deck

No chat, no prose, no parsing — that is the pitch in three negatives. Everything in that response object is machine-readable, so the application branches on it directly instead of scraping a sentence.

Documented claims are kept separate from illustration throughout: the stamp at the foot of each slide says which is which, and the key is on the final slide.

A companion deck, RLCD — The Mathematics of Calibrated Decisions, covers the probability theory behind the numbers shown here.

Chapter 01 of 05

01

The idea

  • 01The problem Jev answers
  • 02What Jev is
  • 03System One, the category
  • 04Trained for calibrated decisions
01 Why another kind of model

The problem Jev answers

Most AI products generate text for people. Most automation needs decisions for software.

The usual workaround

Coerce a text generator into structure, then parse the result back.

Prompt for a format · parse · retry · still no honest uncertainty.

The System One answer

Skip text generation. Return typed values and probability distributions.

No parsing step exists, so no parsing step can fail.

TypeSafe's name for the goal: Machine Native Intelligence.

The premise behind the product

The docs describe the workaround in their own words: teams end up “coercing a text-generation system into outputting structured decisions, then parsing the results back into something your code can depend on.” The alternative: “System One models make fast, structured decisions for software,” returning “typed values and probability distributions that your code can branch on, sort by, and route with.”

Machine Native Intelligence is defined as “AI with software-like properties such as structure, reliability, observability, testability, speed, consistency, and low cost.”

TypeSafe's stated premise is that large-scale automation will consist mostly of machine-to-machine interactions, not human-facing chat — which shifts the design priority from readable responses to inspectable ones. The company explicitly does not aim to build an all-purpose model; text generation, explanation, and conversation are out of scope by design.

02 The model

What Jev is

Typed questions, asked against a state you supply. Typed answers with probabilities come back.

state

your content

→

questions

choice · score · noul

→

jev

one call

→

typed answers

values + probabilities

→

your code

thresholds · routing

Not a text generator

It answers your questions. It never writes a reply.

Typed by construction

Every answer is constrained to the options you supplied.

Calibrated by training

Probabilities optimized against outcomes.

The documented wording, and where the names sit

“Jev is TypeSafe's flagship model and the first System One model.” “jev-1.13 is not trained to generate text.” “Decisions and probabilities conform to the structured software types and JSON schema your code expects.” System One models are “trained for calibrated decisions: their probabilities are optimized against outcomes to reflect uncertainty.”

TypeSafe is the company. System One is the model category — named after the fast-intuition system in dual-process psychology. Jev is the flagship model in it; jev-1.13.0 is current and jev-latest is the SDK default alias.

Requests go to a single endpoint, POST /v1/systemone, with the model field selecting the version. Client SDKs exist for Python and JavaScript. The state is your content — a message, a record, an application snapshot; answers come back keyed by question id.

03 The category

System One — thinking fast, for software

Dual-process psychology splits fast intuition from slow deliberation. TypeSafe claims the fast half for machines.

System One · Jev

“I just know.” Fast, automatic, intuitive: spotting an angry face, reading simple words, answering 2+2.

Snap judgments, at machine volume.

Milliseconds. Fractions of a cent. Millions of times a day.

System Two · reasoning models & people

“Let me think about this.” Slow, deliberate, effortful: working out 27 × 43, comparing mortgages, checking an argument.

Slow, deliberate, generative.

Plans, explains, handles the novel — and takes the escalations.

System 1 answers first“A bat and a ball cost $1.10. The bat costs $1 more than the ball. How much is the ball?” System 1 shouts 10¢. System 2 checks: that makes $1.20. The answer is 5¢.

The division of laborSystem One takes the million small judgments. System Two takes the rare hard residue. Confidence is the handoff signal.

Take the psychology as branding, not as a spec

The documented claims: System One models “make fast, structured decisions for software,” “don't write replies or generate explanations,” and are “trained for calibrated decisions.” The escalation target is spelled out too — “escalate uncertain cases to a person or a more expensive reasoning model.” That is why the routing slides later matter more than any single accuracy number.

The System 1 / System 2 vocabulary was popularized by Kahneman's Thinking, Fast and Slow. It communicates intent but guarantees nothing; the properties that matter are the documented ones — typed outputs, a calibration target, text-only input, independent parallel questions.

Kahneman did not mean two literal systems or physical parts of the brain; they are conceptual labels for automatic versus deliberate processes. The key dynamic: System 1 usually acts first, and System 2 often just accepts its answer. That makes everyday life efficient, and it is also where the biases come from.

The analogy imports a useful warning: human System 1 is fast and biased. The jaggedness slide is exactly this model's bias list. Fast intuition needs slow verification around it, in machines as in people.

Jev is also “no longer the only claimant to the label” — the open-source Laya (chapter 04) calls itself a “System-1 decision model” too.

04 The training story, briefly

Trained for calibrated decisions

RLCD — reinforcement learning for calibrated decisions — as a third post-training family.

RLHF

Rewards answers people prefer.

Risks sycophancy and confident hallucination.

RLVR

Rewards answers that pass a check.

The code compiles; the sum is right.

RLCD

Targets calibrated probabilities.

0.8 predictions should hold about 80% of the time.

The calibration target

correct ×8 not ×2

Among comparable 80% answers, about 8 in 10 correct.

The recipe is not published. The mathematics is in the companion deck: RLCD.

Why calibration is the property software needs

The docs' critique of RLHF: optimizing preference can encourage “sycophancy and confident-sounding hallucinations,” and narrows output diversity. RLVR optimizes what a verifier can confirm. RLCD's stated aim is that predictions marked 0.2 succeed about 20% of the time and those marked 0.8 about 80%.

A category alone cannot tell your application whether to act automatically or ask a person. A calibrated probability can be multiplied by the cost of being wrong and compared with the cost of checking — that arithmetic is the whole automation decision, and it only works if the probability honestly reflects how often outcomes occur.

“Human preference and machine trustworthiness are different optimization targets” is the docs' one-line version. These goals overlap; the claim is about emphasis. Reviewed sources state the goal, not the reward, loss, or update rule — the recipe itself is undisclosed. One biographical note the docs offer: TypeSafe cofounder Diogo Almeida co-invented RLHF.

Chapter 02 of 05

02

The three primitives

  • 01State and questions
  • 02Choice · Score · Noul
  • 03Confidence
  • 04Routing on confidence
05 Anatomy of a call · state

State — the content you evaluate

One state per request. Three shapes are accepted.

String

"My card was charged twice."

Simple cases.

Object — recommended

{"message": …, "order_id": "A-104"}

Named parts, clear relationships.

Array

["Hi", "My number is TS1337.", …]

Sequential messages or records.

Text only — no image, audio, or video input.

What the docs say about state

“State is the content you ask a System One model to evaluate… The state contains the content and supporting facts. Questions define the judgments the model should make about that material.”

On shape: use an object “for most requests so each part of the state has a descriptive name and its relationships remain clear.” The build guide adds that you should reference exact values by path, like support.tickets[0].message, and keep the state lean — accuracy falls as unrelated content grows, and you pay per input token.

05 Anatomy of a call · questions

Questions — batch, don't loop

Every question about the same state travels in one request, evaluated independently and in parallel.

Independent by design

One answer never becomes hidden context for another.

Size limits

64k tokens per request · 32k for state + the longest single question.

Three question types: Choice · Score · Noul — one slide each, next.

The quickstart, in eleven lines of Python

The documented rule: “Send every question that uses the same state in one request” — “one primitive's result does not become hidden context that changes another primitive's result.”

from typesafe_sdk import Choice, Noul, Score, TypeSafeClient

client = TypeSafeClient()
response = client.system_one(
    state=ticket,                       # one state…
    questions={                          # …many judgments
        "department":  Choice(instructions="Which team should handle this",
                              criteria={"billing": …, "technical": …, "sales": …}),
        "frustration": Score(instructions="How frustrated the customer appears",
                             criteria=["Calm, just stating facts",
                                       "Frustrated but civil",
                                       "Very angry, strong language"]),
        "is_urgent":   Noul(instructions="The message conveys urgency"),
    },
)
response.answers["department"].choice   # "technical"
response.answers["frustration"].score   # 1.0
response.answers["is_urgent"].noul      # 1.0

Three question types, one API call, three typed answers.

06 Primitive 1 of 3

Choice — which of these options?

One option from a defined set, when the options have no order: routing, classification, language detection.

You ask

"state": "My running shoes arrived in the
          wrong size. Can I swap them for
          a size 10?",
"questions": {
  "department": {
    "type": "choice",
    "instructions": "Which team should handle this?",
    "criteria": {
      "returns":  "Exchanges, wrong or damaged items",
      "shipping": "Delivery status, delays, lost packages",
      "billing":  "Charges, invoices, payment problems"
    } } }

Jev answers

"department": {
  "type": "choice",
  "choice": "returns",
  "confidence": 1.0,
  "probabilities": {
    "returns":  1.0,
    "shipping": 0.0,
    "billing":  0.0
  }
}

choice is just the top probability. Acting on it is your code's call.

  • Describe options concretely — one differentiating line each.
  • Include an “other” option so edge cases have somewhere to go.
  • Up to 255 options — give the full list, not a shortlist.
Vague descriptions, and when one line isn't enough

The definition: “A Choice is a System One question type for selecting one option from a defined set. The answer includes the selected option, a probability for each option, and confidence.” Vague option descriptions are the documented cause of misclassification.

When options are similar enough to risk confusion, the docs recommend structured option objects with fields like what (the definition), not_for (what belongs elsewhere), and examples (typical cases). Examples strongly influence behavior — only use ones that match your real inputs.

Remember the companion deck's caution: the largest probability wins the selection, but it is not necessarily the cheapest action. A 41% winner in a three-way split may still warrant a human check.

07 Primitive 2 of 3

Score — which level?

2–10 ordered levels, low to high. A probability comes back for each.

You ask

"state": "Export button crashes Safari
          settings page. Works in Chrome.",
"questions": {
  "bug_severity": {
    "type": "score",
    "instructions": "How severe is the reported issue?",
    "criteria": [
      "Cosmetic; no impact to functionality",
      "Broken or degraded feature, but workaround exists",
      "Blocking issue; no workaround exists"
    ] } }

Jev answers

"bug_severity": {
  "type": "score",
  "score": 1.43,
  "confidence": 0.35,
  "probabilities": {
    "0": 0.00,
    "1": 0.57,
    "2": 0.43
  },
  "legend": { "0": "Cosmetic; …", … }
}

Read it in this orderprobabilities — where: 57% workaround, 43% blocking. confidence — how sure: 0.35, not very. score — a one-number summary: 1.43, between levels 1 and 2, closer to 1.

Designing levels that work

The definition: “A Score is a System One question type for rating content against ordered, descriptive levels.”

Describe concrete situations, not abstract degrees. “Broken or degraded feature, but workaround exists” beats “moderately severe” — the model evaluates each level independently, so comparative language (“worse than the previous level”) means nothing to it.

One dimension per Score. Severity and urgency are two questions; a multifaceted judgment hides several answers behind one number.

Normalize before combining. To weight several Scores of different lengths together, divide each by its top level number first, then apply your domain weights.

07 Primitive 2 of 3 · reading it

What the score actually is

“Each level number multiplied by its probability, added up.”

The score
score = ∑i i · pi = 0×0.00 + 1×0.57 + 2×0.43 = 1.43

What 1.43 is notThere is no level 1.43. It is an expected index, not a measurement.

Reading this answer

1.43 sits between level 1 (workaround) and level 2 (blocking), a bit closer to 1.

Confidence 0.35 says the model is torn between the two. Send it to a person, and the question to settle is “is there a workaround?”

Same score, opposite meanings

p 0 / 1 / 2scoremeaning
0.00 / 1.00 / 0.001.00surely level 1 → act
0.50 / 0.00 / 0.501.00trivial or blocking → escalate

Never act on the score alone. Check probabilities and confidence first.

Two distributions, one score

A confident “middling” answer — nearly all mass on level 1 — and a bimodal “either trivial or catastrophic” answer can report the same score. The first is safe to act on; the second is precisely the case to escalate. The score alone cannot tell them apart, which is why the docs say “always examine probabilities and confidence.”

The next slide lets you build both distributions by hand and watch the number stay put.

08 The Score, alive

Play with a Score distribution

Move probability mass between three severity levels; watch the score slide along the 0–2 axis.

The docs' example: P(0) = 0, P(2) = 0.43 → score 1.43.
Now reach 1.00 from the extremes. Same number, very different situation.
What this widget does and does not show

The bars are set by your sliders, not by Jev — this page computes the weighted average locally to make the definition tangible. In a real response the distribution comes from the model, and a separate confidence field summarizes how concentrated it is.

The middle level's probability is whatever the two sliders leave over; if their sum exceeds 1, the widget rescales proportionally, so the three always sum to one — as in a real Score answer.

The second experiment is the point of the slide: push mass to the extremes until the score reads 1.00 with almost nothing on level 1. A bimodal “either trivial or catastrophic” distribution and a confident “middling” one can report the same score — which is why the score alone should never drive a high-stakes action.

09 Primitive 3 of 3

Noul — is this true?

The probability that the answer is yes. One number, and that number is the whole answer.

Near 1

0.99

A strong yes.

Near 0.5

0.50

Uncertain — not medium. Not an intensity dial.

Near 0

0.02

A strong no. Phrase so high means yes.

Noul is the documented name — not a typo for Bool. Reference ↗

One question per Noul

The definition: “A Noul question asks the model to evaluate a yes/no question and return the probability that the answer is yes.” There is no separate confidence field — the number is the confidence. Near 0.5 means “yes and no get similar probability.” The docs' example is_human_escalation returns "noul": 0.99. The API identifier is noul; the docs don't explain the etymology.

Use a Noul when the answer is yes or no — does the message request a refund, does the résumé mention distributed systems, does the comment contain personal data. A question with multiple conditions (“urgent and from a paying customer?”) should become separate Nouls combined with boolean logic in your code.

From the jaggedness page: structural invariants across separate questions are not guaranteed — a Noul and its negation may not give probabilities summing to one. Ask one question, phrased so high means yes.

09 Primitive 3 of 3 · in practice

Sharpening a Noul

Optional criteria draw the boundary. Your threshold decides what to do about it.

Criteria sharpen the boundary

"is_repeat_contact": {
  "type": "noul",
  "instructions": "Has the customer contacted
                   support about this before?",
  "criteria": {
    "true":  "Mentions a prior attempt, ticket,
              or that they have asked before",
    "false": "No sign of any previous contact"
  } }

Where to set your threshold

Raise it when a false yes is expensive — paging someone, issuing a refund.

Lower it when a missed yes is expensive — failing to flag a safety issue.

The threshold is a ratio of your costs

The documented guidance, verbatim: “Raise it when acting on a false yes is expensive, such as paging someone or issuing a refund. Lower it when missing a true yes is expensive, such as failing to flag a safety issue.”

That is the cost logic from the companion deck in miniature: the right cutoff is a ratio of the two error costs, not a universal number. The RLCD deck's fraud example works it out with actual dollar amounts.

10 The fourth field

Confidence — the shape of the answer

One number, 0 to 1. Concentrated → high. Spread → low.

From the documented Score example

top probability (level 1)0.57
runner-up (level 2)0.43
reported confidence0.35

Not the winner's probability

Two levels nearly tied is a spread-out shape, and the confidence says so — 0.35, not 0.57.

Where the number comes from

Every Choice and Score answer carries a confidence: “a single number from 0 to 1” you can threshold on. It summarizes the distribution's shape — “concentrated on one outcome means a confident answer, spread out means an uncertain one.” Nouls have none; the probability is the answer.

The companion deck's lecture example makes the same point with a Choice: probabilities of 0.60 / 0.38 / 0.02 shipped with a confidence of 0.39 — not 0.60, and not derivable from the winner alone.

10 The fourth field · using it

Two non-negotiables, three bands

confidence tells you how concentrated the answer is — not whether it is right. Use it to decide what to do next.

Don't reverse-engineer a formula

The only documented rule: concentrated → high, spread → low. Don't build logic on the exact math. Need a specific measure, like the gap between the top two options? Compute it from the full probabilities.

Don't read it as ℙ(correct)

The model can be confidently wrong: probabilities piled on the wrong option still give high confidence. Concentration is not correctness.

High

Act automatically.

Medium

Confirm, or queue for review.

Low

Do not act. Route to a human, ask for clarification, or fall back to another system.

Thresholds depend on the actionSet a separate gate for each action, based on the cost of being wrong: a refund or delete needs a higher bar than a read-only lookup. Start conservative, then plot confidence against accuracy on your own outcomes and adjust.

Thresholds scale with consequences

On low confidence the guidance is explicit: “do not act. Route to a human, request clarification, or fall back to a different system.”

“A confidence threshold is not one number. Different actions within the same system should be gated at different levels depending on the consequences of getting it wrong.” A destructive operation deserves a higher gate than a read-only one, even inside the same workflow.

“Start with conservative thresholds, test with your own data, and adjust as you observe results” — and from the build guide: plot confidence against accuracy on your own outcomes. That validation is what turns a heuristic into an engineering decision. The next slide makes the bands draggable.

11 The three bands, alive

Route on confidence

Two gates split the confidence axis into three zones: escalate, confirm, act. The model supplies the confidence; you choose where the gates sit.

Same answer, different actions

Take one answer with confidence 0.80.
Add an internal tag — gate at 0.60 → act
Issue a refund — gate at 0.95 → confirm first
(illustrative numbers)

Try it① Move the answer's confidence and watch where the route changes.
② Move the two gates and watch the same answer land in a different zone.
The gates are values in your code, not in the model.

Try the two failure modes

Escalation is documented: uncertain cases go “to a person or a more expensive reasoning model.”

Gates too generous: drag “act above” down to 0.5. More answers auto-execute, including near-ties — fine for tagging, reckless for refunds.

Gates too timid: drag “escalate below” up toward the act gate. The confirm band vanishes and nearly everything lands on humans, which is the old manual process with extra steps.

Set each gate from the cost of that action going wrong, then validate by plotting confidence against accuracy on your own data. The defaults here (0.40 / 0.80) are arbitrary starting points for the demo.

Chapter 03 of 05

03

Building with it

  • 01The model card
  • 02Jagged edges
  • 03What people build
  • 04The six build habits
12 Practical facts

The model card, in numbers

PropertyValueNote
Current modeljev-1.13.0jev-latest and jev-preview alias it today
Context length64k tokens per request32k for state + longest question
InputText onlystring, object, or array — no image, audio, video
Price$42 / Btok · $0.042 / Mtokcharged per input token; output free
Rate limits250,000 tok/s · 1,200 req/minadjusts dynamically under demand
LanguagesEnglish primaryothers reduced — test before production
CustomizationRequest parameters onlyno per-customer fine-tuned weights

A snapshot of docs.typesafe.ai/models on 27 September 2026 — not a contract.

What the pricing shape implies

Because answers are typed values rather than generated text, output is tiny and TypeSafe simply doesn't charge for it — the bill is a function of how much state you send. That makes the build guide's advice to keep state lean a cost lever as well as an accuracy lever.

At $0.042 per million input tokens, a 1,000-token state with a handful of questions costs on the order of hundredths of a cent — the “run it a million times in the background” use cases are priced deliberately.

Rate limits are “adjusting dynamically” under demand and may change without notice; pricing and limits are the vendor's to change.

13 Documented limitations · 1 of 2

Jagged edges — reading and arithmetic

The vendor's own failure list. The pattern in the fixes: let code do what code does best.

EdgeDocumented behaviorDo instead
Literal readingAnswers the written question, not the intentState the exact condition
CountingDoes not count reliablyCount in code
Dates & durationsOrdering and gaps are unreliableCompare in code; pass the verdict in
Numeric encodingsHex and RGB worse than wordsUse semantic descriptions
Why publishing this list matters

The page opens plainly: “Jev isn't perfect. Here are some jagged edges we are aware of with jev-1.13.” A vendor documenting its model's failure modes is exactly the observability-and-testability posture the manifesto claims — and it gives you a concrete test plan: each row is a category of cases for your own evaluation set before shipping.

Ask Jev the semantic part and keep the arithmetic in code: “does this message mention more than one order?” is a bad question; “which orders does it mention?” plus len() is a good one.

13 Documented limitations · 2 of 2

Jagged edges — context and composition

EdgeDocumented behaviorDo instead
IndirectionMulti-hop lookups degrade accuracyResolve references first
Irrelevant contextAccuracy falls as the state growsKeep state lean
Adversarial contentMisleading framings can sway answersTreat state as untrusted input
Cross-question invariantsRelations between answers aren't guaranteedAsk one question; derive the rest

These are jev-1.13 notes, not permanent truths.

Two more edges from the same page

Contradictory instructions: when instructions and criteria disagree, answers degrade — keep them aligned.

No text generation: by design rather than defect. When you need prose, pair Jev with an LLM — Jev decides, the LLM writes.

Recheck the jaggedness page when the model version changes; the list is versioned to jev-1.13.

14 Where it fits

What people build with it

The ambition, in the docs' words: “run it a million times in the background without a human co-pilot.”

LLM guardrails

Semantic checks on every LLM input, output, and tool call.

LLM orchestration

Route prompts between models; gate tool calls.

Universal verification

Check extractions, reasoning traces, tool calls.

RAG & map-reduce

Semantic search, scoring, and ranking over big datasets.

Real-time decisions

Faster than human perception — games, embedded UI.

Operations automation

Triage, claims, moderation, lead scoring, fraud signals.

Vendor-described applications, not independently benchmarked results.

The common thread

Every entry has the same silhouette: a high-volume stream of small judgments, each cheap to make and cheap to check, feeding deterministic software that owns the consequences. None asks the model to converse, plan, or explain.

The inverse is a useful negative test: if your task needs one large, novel, multi-step judgment — a legal brief, an architecture decision — that is reasoning-model territory, and Jev's role shrinks to verifying pieces of it.

The docs' own phrasings are worth quoting: guardrails “at a fraction of the cost”; “build a custom router that chooses which LLM receives each prompt”; “replace or supplement embeddings in RAG pipelines”; moderation against “company-specific, nuanced criteria.” The cookbook has worked recipes for most of them — re-ranking, classifying RAG passages, double-checking citations, function calling, and guardrails.

15 The method

How to build with it

Six habits, one principle: code owns control flow; the model makes narrow judgments.

1 · Code first

No agent loop where an if suffices.

2 · Lean state

Irrelevant context costs accuracy and money.

3 · Structured state

Nested JSON, descriptive names, exact paths.

4 · Atomize questions

One judgment per question. The guide's most important idea.

5 · Parallelize

Many small questions, one request.

6 · Compose in code

Deterministic rules or weighted sums.

The escape hatch is part of the design: escalate uncertain cases; validate thresholds against your outcomes.
Atomization, worked

The guide's wording: “Keep deterministic work in code. It is reliable and cheap.” Atomization is “probably the most important concept in this guide. Broad questions hide several judgments behind one answer.” Compose with “deterministic rules or weighted sums” — or feed the probabilities into a classical ML model. Reference exact values by path, like support.tickets[0].message.

Instead of one broad question — “should we refund this customer?” — the docs' refund example builds a state holding the customer's message, the transactions, and the policy, then asks independent questions together: does the message request a refund? (Noul) · is the order within the policy window? (Noul) · how frustrated is the customer? (Score) · which team owns this? (Choice).

The refund decision is then an auditable boolean expression in your code. When it misfires you can see which judgment was wrong — the observability a broad question destroys.

Chapter 04 of 05

04

The case against

  • 01Just a fine-tuned classifier?
  • 02How much is hype
  • 03One independent test
  • 04Laya, the open-source rival
16 The case against · 1 of 3

“Isn't this just a fine-tuned classifier?”

The most common objection — and a fair one. Typed classification has a decade of published practice behind it.

The prior art is real

BERT set the template in 2018; bert-base-uncased still sees 47M+ downloads a month.

Community verdict: “the industry rediscovering classification models.”

The published assessment

KDnuggets: “Classification is not new. Intent detection is not new.”

Same piece: “the architecture and product around it may be new.”

Why it stays open

The transparency gapArchitecture, parameter count, training data, and the RLCD reward are undisclosed. So “just a fine-tune” can be neither confirmed nor refuted.

The literature the objection leans on — and the fair restatement

Zero-shot classifiers such as facebook/bart-large-mnli and GLiNER-style encoders “already do something similar.” Reddit and GeekNews verdicts include “essentially BERT with more data” and “I implemented the same functionality with BERT-family models years ago.” KDnuggets also places Jev outside “the same category as GPT, Claude, Gemini, or other frontier LLMs.”

Each ingredient has a citable history: fine-tuning pretrained encoders for classification (Devlin et al., 2018), zero-shot classification via NLI (Yin et al., 2019), and post-hoc calibration including temperature scaling (Guo et al., 2017).

What is documented: one shared set of weights for all accounts, no per-customer fine-tuning, and the task chosen at runtime by the question text — an operational difference from one-classifier-per-task, whatever the architecture turns out to be. The fair restatement of the claim is narrower than the marketing: not a new kind of AI, but engineering on a known recipe.

17 The case against · 2 of 3

How much of it is hype?

Much of the launch coverage came from influencers. Strip the adjectives and measure the claims.

“Zero hallucinations” ≠ zero errors

It means zero out-of-schema outputs. Jev can still choose Billing when the answer was Technical.

The 0% is a definition, not a measurement.

The vendor's 68% is unanchored

Scored against frontier-model reference answers, not ground truth.

On real accuracy, the published verdict: “we do not really know yet.”

Where the two figures come from

The flagship 0% hallucination figure “was plotted at zero because schema conformance is guaranteed by design” — the model cannot emit a value outside the options you defined, so the metric is true by construction and says nothing about whether the chosen option is right.

TypeSafe's accuracy benchmark scores Jev against frontier-model reference answers. That measures agreement with another model, not correctness, so it cannot support a general accuracy claim.

Neither observation says the product is bad — they say these two numbers are not evidence of quality. The next slide is someone actually measuring.

17 The case against · 2 of 3 · measured

One independent test

Phishing detection, 2,000 synthetic emails.

ApproachAccuracy
Jev · one Noul question62.6%zero-shot
Claude Haiku 4.5 · one question81.3%zero-shot
A two-line regex91.8%dataset nearly separable by construction
Haiku · five-question composite93.2%
Jev · five questions + fitted weights95.0%p = 0.063 vs Haiku — not significant

What survives measurementCost and speed do — 12–27× cheaper, 2.9–5× faster than Haiku. Out-of-box accuracy and phrasing robustness don't.

The caveats, and a fair defense from the same source

On phrasing robustness: the same question asked as a Noul returned 0.22 — asked as a yes/no Choice, “no” at probability 0.99. Measured calibration is task-dependent: ECE 0.154 on phishing (worse than Haiku's 0.097), 0.0712 on tool-call classification.

The decomposition workflow — five narrow questions, weights fitted on 1,000 labels — is where Jev won, and it pushes the logic onto the code side where it can be audited. That is the product's thesis working as advertised; it just required labeled data and fitting, like any classifier.

The regex's 91.8% cuts both ways: this synthetic benchmark was easy, which weakens both Jev's win and the criticism drawn from its single-question loss. GeekNews' top comment offers the charitable reading: “the skill of clearly communicating what something does and where it's useful is equally part of innovation.” Others noted the practical value plainly: solving a classification problem “in minutes instead of weeks.”

18 The case against · 3 of 3

Laya — the open-source rival

The strongest version of the criticism ships weights: same three primitives, Apache 2.0, recipe published.

jev-1.13.0laya
Primitiveschoice · score · noulthe same three
Weightsclosed — hosted APIopen — 421M, ModernBERT-large
RLCD recipeundisclosedpublished
Latency~239 ms median~33 ms local GPU
LanguagesEnglish primary100+
Price$0.042 / Mtokfree — your hardware
Specifications and sources

Laya (Convai Innovations) serves the three primitives with calibrated probabilities in one ~33 ms forward pass, 7.2 ms/question batched, and publishes its recipe: proper-scoring-rule reward with Gaussian logit noise. Its question schema is near-identical to Jev's.

From the Hugging Face model card (checked 27 September 2026): three checkpoints (English root 421M on ModernBERT-large, a 322M mmBERT multilingual variant, a typed-decisions fine-tune), 103–332 questions/second on a Tesla T4, and 45 of 51 tested languages above chance. Jev's ~239 ms median latency is from the independent test on the previous slide.

18 The case against · 3 of 3 · the caveats

Laya cuts both ways

“Exactly the same thing” needs a caveat

Zero-shot, Laya's base checkpoints score below the majority-class baseline. The headline number is a fine-tuned checkpoint, and it ships overconfident.

Handle the head-to-head with tongs. The circulating comparison is Laya benchmarking Laya. What needs no tongs: the weights are open and the interface matches.

Both things are trueA 421M fine-tuned encoder can serve this API — and out-of-box calibrated quality is the hard part. Either way, Laya is the baseline to benchmark Jev against.

The numbers behind the caveat

By Laya's own model card, its base checkpoints score 0.362 zero-shot against a 0.461 majority-class baseline — near random. The headline 0.766 is a fine-tuned checkpoint; ECE runs 0.213–0.466 before temperature refitting, and its noul “can follow its option labels instead of the state.”

The circulating head-to-head (accuracy 0.766 vs 0.727; ECE 0.081 vs 0.246) is Laya's vendor benchmarking its fine-tuned model on its own dataset — treat it exactly like TypeSafe's 68%: an advert until independently reproduced.

One further allegation in the GeekNews comments — that a sales-conversion demo model had outcome labels leak into training — is from a single anonymous commenter and is unverified; it is mentioned only so you know it exists.

If the audience takes one thing from this chapter: run Laya as your baseline before paying for anything.

Chapter 05 of 05

05

Hands on, and what to take away

  • 01The same question in Python, twice
  • 02Jev on one slide
  • 03Sources and scope
19 Hands on

The same question in Python, twice

Near-identical schemas. What changes is who runs the weights.

Jev — hosted API

$ pip install typesafe-sdk

from typesafe_sdk import Choice, Noul, TypeSafeClient

client = TypeSafeClient()   # needs TYPESAFE_API_KEY
state  = "Billed twice for March — refund today or we cancel."

r = client.system_one(state=state, questions={
    "department": Choice(
        instructions="Which department should handle this?",
        criteria={"billing":   "invoices, payments, refunds",
                  "technical": "bugs, outages, system errors",
                  "other":     "everything else"}),
    "churn_risk": Noul(
        instructions="Does the user threaten to cancel?"),
})

r.answers["department"].choice       # "billing"
r.answers["department"].confidence   # 0–1 — gate on it
r.answers["churn_risk"].noul         # probability of yes

Laya — open weights, runs locally

$ pip install laya

from laya import Router

router = Router()   # downloads checkpoints on first use
state  = "Billed twice for March — refund today or we cancel."

result = router.predict(state, {
    "department": {
        "type": "choice",
        "instructions": "Which department should handle this?",
        "criteria": {"billing":   "invoices, payments, refunds",
                     "technical": "bugs, outages, system errors",
                     "other":     "everything else"}},
    "churn_risk": {
        "type": "noul",
        "instructions": "Does the user threaten to cancel?"},
})

result["answers"]["department"]["choice"]   # "billing"
result["answers"]["churn_risk"]["noul"]     # e.g. 0.892

Why this mattersAn interface this close means you can A/B both on a thousand labeled examples of your traffic in an afternoon.

Provenance of the code, and each vendor's fine print

The Jev snippet is the documented quickstart pattern (TypeSafeClient().system_one with Choice/Score/Noul objects) with this deck's billing example substituted for the docs' Stripe ticket. The Laya snippet follows its model card's Router.predict usage; the 0.892 noul is the card's own example output. Jev-side values will vary run to run — they are labels here, not reproduced results.

Both SDKs are pip-installable today. Laya also exposes laya.load("convaiinnovations/laya") for pinning a single checkpoint instead of the auto-routing Router.

The fine print each vendor states: Laya wants a fine-tune and a temperature refit before production; Jev wants confidence thresholds validated against your own outcomes.

20 Take this away

Jev on one slide

#IdeaIn one sentenceStatus
1System OneTyped decisions with probabilities — never text.documented
2One callOne state, many independent questions, in parallel.documented
3Choice · Score · NoulWhich option — which level — is this true.documented
4CalibrationRLCD targets honest probabilities; recipe unpublished.goal only
5ConfidenceA shape summary — not ℙ(correct).documented
6ThresholdsYour numbers, set per action from costs.your job
7Jagged edgesCount, compare, and compose in code.documented
8VerificationValid doesn't mean true — log outcomes.your job
9The case againstPrior art is real; Laya is the natural baseline.contested

Define the question precisely, keep control flow in code, and test the probabilities against your own outcomes.

The two decks together

This deck answers what Jev is and how to hold it: the product shape, the primitives, the documented limits, the build method. The companion deck, RLCD — The Mathematics of Calibrated Decisions, answers why the probabilities can be trusted and what to do with them: calibration, proper scoring rules, and expected-cost decisions.

Presented together, run this one first — the mathematics lands better once the audience has seen a real answer object with a probability distribution in it.

Sources: TypeSafe documentation (introduction, System One, primitives, confidence, models, jaggedness, use-case map, build guide), reviewed 27 September 2026, plus the third-party material listed on the next slide.

21 Sources & scope

Sources and scope

Criticism & alternatives

documented illustrative undisclosed third-party
Scope notes

Provenance key in full: documented — reproduced from TypeSafe documentation; illustrative — local widgets and teaching devices, not model output; undisclosed — stated goals whose mechanisms the sources do not specify; third-party — external tests, commentary, and the Laya model card, quoted for balance, not endorsed.

Example values (the shoe ticket's 1.0, the bug's 1.43 / 0.35, the 0.99 noul) are reproduced from the documentation pages above and may change as the docs evolve.

Companion deck: RLCD — The Mathematics of Calibrated Decisions, which reviews the lecture video in this folder.

← → slides · N notes · L language · F fullscreen