A field guide to the model
Presented by Han Cheol Moon
27 September 2026
Jev · TypeSafe's flagship model, and the first System One model
A typed question goes in. A typed answer with probabilities comes out.
"department": {
"type": "choice",
"choice": "returns",
"confidence": 1.0,
"probabilities": {
"returns": 1.0,
"shipping": 0.0,
"billing": 0.0
}
}
Branch on it. Sort by it. Route with it.
Response shape from the Choice documentationNo chat, no prose, no parsing — that is the pitch in three negatives. Everything in that response object is machine-readable, so the application branches on it directly instead of scraping a sentence.
Documented claims are kept separate from illustration throughout: the stamp at the foot of each slide says which is which, and the key is on the final slide.
A companion deck, RLCD — The Mathematics of Calibrated Decisions, covers the probability theory behind the numbers shown here.
Five chapters: what Jev is, how to ask it things, how to build on it, what the critics say, and the code.
Why a model that refuses to write text
slide 04Choice, Score, Noul — and confidence
slide 09The model card, the limits, the method
slide 21Prior art, hype, and an open-source rival
slide 27Real code, and what to take away
slide 33Chapter 01 of 05
Most AI products generate text for people. Most automation needs decisions for software.
Coerce a text generator into structure, then parse the result back.
Prompt for a format · parse · retry · still no honest uncertainty.
Skip text generation. Return typed values and probability distributions.
No parsing step exists, so no parsing step can fail.
TypeSafe's name for the goal: Machine Native Intelligence.
The docs describe the workaround in their own words: teams end up “coercing a text-generation system into outputting structured decisions, then parsing the results back into something your code can depend on.” The alternative: “System One models make fast, structured decisions for software,” returning “typed values and probability distributions that your code can branch on, sort by, and route with.”
Machine Native Intelligence is defined as “AI with software-like properties such as structure, reliability, observability, testability, speed, consistency, and low cost.”
TypeSafe's stated premise is that large-scale automation will consist mostly of machine-to-machine interactions, not human-facing chat — which shifts the design priority from readable responses to inspectable ones. The company explicitly does not aim to build an all-purpose model; text generation, explanation, and conversation are out of scope by design.
Typed questions, asked against a state you supply. Typed answers with probabilities come back.
your content
choice · score · noul
one call
values + probabilities
thresholds · routing
It answers your questions. It never writes a reply.
Every answer is constrained to the options you supplied.
Probabilities optimized against outcomes.
“Jev is TypeSafe's flagship model and the first System One model.” “jev-1.13 is not trained to generate text.” “Decisions and probabilities conform to the structured software types and JSON schema your code expects.” System One models are “trained for calibrated decisions: their probabilities are optimized against outcomes to reflect uncertainty.”
TypeSafe is the company. System One is the model category — named after the fast-intuition system in dual-process psychology. Jev is the flagship model in it; jev-1.13.0 is current and jev-latest is the SDK default alias.
Requests go to a single endpoint, POST /v1/systemone, with the model field selecting the version. Client SDKs exist for Python and JavaScript. The state is your content — a message, a record, an application snapshot; answers come back keyed by question id.
Dual-process psychology splits fast intuition from slow deliberation. TypeSafe claims the fast half for machines.
“I just know.” Fast, automatic, intuitive: spotting an angry face, reading simple words, answering 2+2.
Snap judgments, at machine volume.
Milliseconds. Fractions of a cent. Millions of times a day.
“Let me think about this.” Slow, deliberate, effortful: working out 27 × 43, comparing mortgages, checking an argument.
Slow, deliberate, generative.
Plans, explains, handles the novel — and takes the escalations.
System 1 answers first“A bat and a ball cost $1.10. The bat costs $1 more than the ball. How much is the ball?” System 1 shouts 10¢. System 2 checks: that makes $1.20. The answer is 5¢.
The division of laborSystem One takes the million small judgments. System Two takes the rare hard residue. Confidence is the handoff signal.
The documented claims: System One models “make fast, structured decisions for software,” “don't write replies or generate explanations,” and are “trained for calibrated decisions.” The escalation target is spelled out too — “escalate uncertain cases to a person or a more expensive reasoning model.” That is why the routing slides later matter more than any single accuracy number.
The System 1 / System 2 vocabulary was popularized by Kahneman's Thinking, Fast and Slow. It communicates intent but guarantees nothing; the properties that matter are the documented ones — typed outputs, a calibration target, text-only input, independent parallel questions.
Kahneman did not mean two literal systems or physical parts of the brain; they are conceptual labels for automatic versus deliberate processes. The key dynamic: System 1 usually acts first, and System 2 often just accepts its answer. That makes everyday life efficient, and it is also where the biases come from.
The analogy imports a useful warning: human System 1 is fast and biased. The jaggedness slide is exactly this model's bias list. Fast intuition needs slow verification around it, in machines as in people.
Jev is also “no longer the only claimant to the label” — the open-source Laya (chapter 04) calls itself a “System-1 decision model” too.
RLCD — reinforcement learning for calibrated decisions — as a third post-training family.
Rewards answers people prefer.
Risks sycophancy and confident hallucination.
Rewards answers that pass a check.
The code compiles; the sum is right.
Targets calibrated probabilities.
0.8 predictions should hold about 80% of the time.
Among comparable 80% answers, about 8 in 10 correct.
The recipe is not published. The mathematics is in the companion deck: RLCD.
The docs' critique of RLHF: optimizing preference can encourage “sycophancy and confident-sounding hallucinations,” and narrows output diversity. RLVR optimizes what a verifier can confirm. RLCD's stated aim is that predictions marked 0.2 succeed about 20% of the time and those marked 0.8 about 80%.
A category alone cannot tell your application whether to act automatically or ask a person. A calibrated probability can be multiplied by the cost of being wrong and compared with the cost of checking — that arithmetic is the whole automation decision, and it only works if the probability honestly reflects how often outcomes occur.
“Human preference and machine trustworthiness are different optimization targets” is the docs' one-line version. These goals overlap; the claim is about emphasis. Reviewed sources state the goal, not the reward, loss, or update rule — the recipe itself is undisclosed. One biographical note the docs offer: TypeSafe cofounder Diogo Almeida co-invented RLHF.
Chapter 02 of 05
One state per request. Three shapes are accepted.
"My card was charged twice."
Simple cases.
{"message": …, "order_id": "A-104"}
Named parts, clear relationships.
["Hi", "My number is TS1337.", …]
Sequential messages or records.
Text only — no image, audio, or video input.
“State is the content you ask a System One model to evaluate… The state contains the content and supporting facts. Questions define the judgments the model should make about that material.”
On shape: use an object “for most requests so each part of the state has a descriptive name and its relationships remain clear.” The build guide adds that you should reference exact values by path, like support.tickets[0].message, and keep the state lean — accuracy falls as unrelated content grows, and you pay per input token.
Every question about the same state travels in one request, evaluated independently and in parallel.
One answer never becomes hidden context for another.
64k tokens per request · 32k for state + the longest single question.
Three question types: Choice · Score · Noul — one slide each, next.
The documented rule: “Send every question that uses the same state in one request” — “one primitive's result does not become hidden context that changes another primitive's result.”
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
client = TypeSafeClient()
response = client.system_one(
state=ticket, # one state…
questions={ # …many judgments
"department": Choice(instructions="Which team should handle this",
criteria={"billing": …, "technical": …, "sales": …}),
"frustration": Score(instructions="How frustrated the customer appears",
criteria=["Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language"]),
"is_urgent": Noul(instructions="The message conveys urgency"),
},
)
response.answers["department"].choice # "technical"
response.answers["frustration"].score # 1.0
response.answers["is_urgent"].noul # 1.0
Three question types, one API call, three typed answers.
One option from a defined set, when the options have no order: routing, classification, language detection.
"state": "My running shoes arrived in the
wrong size. Can I swap them for
a size 10?",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"returns": "Exchanges, wrong or damaged items",
"shipping": "Delivery status, delays, lost packages",
"billing": "Charges, invoices, payment problems"
} } }"department": {
"type": "choice",
"choice": "returns",
"confidence": 1.0,
"probabilities": {
"returns": 1.0,
"shipping": 0.0,
"billing": 0.0
}
}
choice is just the top probability. Acting on it is your code's call.
The definition: “A Choice is a System One question type for selecting one option from a defined set. The answer includes the selected option, a probability for each option, and confidence.” Vague option descriptions are the documented cause of misclassification.
When options are similar enough to risk confusion, the docs recommend structured option objects with fields like what (the definition), not_for (what belongs elsewhere), and examples (typical cases). Examples strongly influence behavior — only use ones that match your real inputs.
Remember the companion deck's caution: the largest probability wins the selection, but it is not necessarily the cheapest action. A 41% winner in a three-way split may still warrant a human check.
2–10 ordered levels, low to high. A probability comes back for each.
"state": "Export button crashes Safari
settings page. Works in Chrome.",
"questions": {
"bug_severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists"
] } }"bug_severity": {
"type": "score",
"score": 1.43,
"confidence": 0.35,
"probabilities": {
"0": 0.00,
"1": 0.57,
"2": 0.43
},
"legend": { "0": "Cosmetic; …", … }
}Read it in this orderprobabilities — where: 57% workaround, 43% blocking. confidence — how sure: 0.35, not very. score — a one-number summary: 1.43, between levels 1 and 2, closer to 1.
The definition: “A Score is a System One question type for rating content against ordered, descriptive levels.”
Describe concrete situations, not abstract degrees. “Broken or degraded feature, but workaround exists” beats “moderately severe” — the model evaluates each level independently, so comparative language (“worse than the previous level”) means nothing to it.
One dimension per Score. Severity and urgency are two questions; a multifaceted judgment hides several answers behind one number.
Normalize before combining. To weight several Scores of different lengths together, divide each by its top level number first, then apply your domain weights.
“Each level number multiplied by its probability, added up.”
What 1.43 is notThere is no level 1.43. It is an expected index, not a measurement.
1.43 sits between level 1 (workaround) and level 2 (blocking), a bit closer to 1.
Confidence 0.35 says the model is torn between the two. Send it to a person, and the question to settle is “is there a workaround?”
| p 0 / 1 / 2 | score | meaning |
|---|---|---|
| 0.00 / 1.00 / 0.00 | 1.00 | surely level 1 → act |
| 0.50 / 0.00 / 0.50 | 1.00 | trivial or blocking → escalate |
Never act on the score alone. Check probabilities and confidence first.
A confident “middling” answer — nearly all mass on level 1 — and a bimodal “either trivial or catastrophic” answer can report the same score. The first is safe to act on; the second is precisely the case to escalate. The score alone cannot tell them apart, which is why the docs say “always examine probabilities and confidence.”
The next slide lets you build both distributions by hand and watch the number stay put.
Move probability mass between three severity levels; watch the score slide along the 0–2 axis.
The bars are set by your sliders, not by Jev — this page computes the weighted average locally to make the definition tangible. In a real response the distribution comes from the model, and a separate confidence field summarizes how concentrated it is.
The middle level's probability is whatever the two sliders leave over; if their sum exceeds 1, the widget rescales proportionally, so the three always sum to one — as in a real Score answer.
The second experiment is the point of the slide: push mass to the extremes until the score reads 1.00 with almost nothing on level 1. A bimodal “either trivial or catastrophic” distribution and a confident “middling” one can report the same score — which is why the score alone should never drive a high-stakes action.
The probability that the answer is yes. One number, and that number is the whole answer.
0.99
A strong yes.
0.50
Uncertain — not medium. Not an intensity dial.
0.02
A strong no. Phrase so high means yes.
Noul is the documented name — not a typo for Bool. Reference ↗
The definition: “A Noul question asks the model to evaluate a yes/no question and return the probability that the answer is yes.” There is no separate confidence field — the number is the confidence. Near 0.5 means “yes and no get similar probability.” The docs' example is_human_escalation returns "noul": 0.99. The API identifier is noul; the docs don't explain the etymology.
Use a Noul when the answer is yes or no — does the message request a refund, does the résumé mention distributed systems, does the comment contain personal data. A question with multiple conditions (“urgent and from a paying customer?”) should become separate Nouls combined with boolean logic in your code.
From the jaggedness page: structural invariants across separate questions are not guaranteed — a Noul and its negation may not give probabilities summing to one. Ask one question, phrased so high means yes.
Optional criteria draw the boundary. Your threshold decides what to do about it.
"is_repeat_contact": {
"type": "noul",
"instructions": "Has the customer contacted
support about this before?",
"criteria": {
"true": "Mentions a prior attempt, ticket,
or that they have asked before",
"false": "No sign of any previous contact"
} }Raise it when a false yes is expensive — paging someone, issuing a refund.
Lower it when a missed yes is expensive — failing to flag a safety issue.
The documented guidance, verbatim: “Raise it when acting on a false yes is expensive, such as paging someone or issuing a refund. Lower it when missing a true yes is expensive, such as failing to flag a safety issue.”
That is the cost logic from the companion deck in miniature: the right cutoff is a ratio of the two error costs, not a universal number. The RLCD deck's fraud example works it out with actual dollar amounts.
One number, 0 to 1. Concentrated → high. Spread → low.
| top probability (level 1) | 0.57 |
| runner-up (level 2) | 0.43 |
| reported confidence | 0.35 |
Two levels nearly tied is a spread-out shape, and the confidence says so — 0.35, not 0.57.
Every Choice and Score answer carries a confidence: “a single number from 0 to 1” you can threshold on. It summarizes the distribution's shape — “concentrated on one outcome means a confident answer, spread out means an uncertain one.” Nouls have none; the probability is the answer.
The companion deck's lecture example makes the same point with a Choice: probabilities of 0.60 / 0.38 / 0.02 shipped with a confidence of 0.39 — not 0.60, and not derivable from the winner alone.
confidence tells you how concentrated the answer is — not whether it is right. Use it to decide what to do next.
The only documented rule: concentrated → high, spread → low. Don't build logic on the exact math. Need a specific measure, like the gap between the top two options? Compute it from the full probabilities.
The model can be confidently wrong: probabilities piled on the wrong option still give high confidence. Concentration is not correctness.
Act automatically.
Confirm, or queue for review.
Do not act. Route to a human, ask for clarification, or fall back to another system.
Thresholds depend on the actionSet a separate gate for each action, based on the cost of being wrong: a refund or delete needs a higher bar than a read-only lookup. Start conservative, then plot confidence against accuracy on your own outcomes and adjust.
On low confidence the guidance is explicit: “do not act. Route to a human, request clarification, or fall back to a different system.”
“A confidence threshold is not one number. Different actions within the same system should be gated at different levels depending on the consequences of getting it wrong.” A destructive operation deserves a higher gate than a read-only one, even inside the same workflow.
“Start with conservative thresholds, test with your own data, and adjust as you observe results” — and from the build guide: plot confidence against accuracy on your own outcomes. That validation is what turns a heuristic into an engineering decision. The next slide makes the bands draggable.
Two gates split the confidence axis into three zones: escalate, confirm, act. The model supplies the confidence; you choose where the gates sit.
Take one answer with confidence 0.80.
Add an internal tag — gate at 0.60 → act
Issue a refund — gate at 0.95 → confirm first
(illustrative numbers)
Try it① Move the answer's confidence and watch where the route changes.
② Move the two gates and watch the same answer land in a different zone.
The gates are values in your code, not in the model.
Escalation is documented: uncertain cases go “to a person or a more expensive reasoning model.”
Gates too generous: drag “act above” down to 0.5. More answers auto-execute, including near-ties — fine for tagging, reckless for refunds.
Gates too timid: drag “escalate below” up toward the act gate. The confirm band vanishes and nearly everything lands on humans, which is the old manual process with extra steps.
Set each gate from the cost of that action going wrong, then validate by plotting confidence against accuracy on your own data. The defaults here (0.40 / 0.80) are arbitrary starting points for the demo.
Chapter 03 of 05
| Property | Value | Note |
|---|---|---|
| Current model | jev-1.13.0 | jev-latest and jev-preview alias it today |
| Context length | 64k tokens per request | 32k for state + longest question |
| Input | Text only | string, object, or array — no image, audio, video |
| Price | $42 / Btok · $0.042 / Mtok | charged per input token; output free |
| Rate limits | 250,000 tok/s · 1,200 req/min | adjusts dynamically under demand |
| Languages | English primary | others reduced — test before production |
| Customization | Request parameters only | no per-customer fine-tuned weights |
A snapshot of docs.typesafe.ai/models on 27 September 2026 — not a contract.
Because answers are typed values rather than generated text, output is tiny and TypeSafe simply doesn't charge for it — the bill is a function of how much state you send. That makes the build guide's advice to keep state lean a cost lever as well as an accuracy lever.
At $0.042 per million input tokens, a 1,000-token state with a handful of questions costs on the order of hundredths of a cent — the “run it a million times in the background” use cases are priced deliberately.
Rate limits are “adjusting dynamically” under demand and may change without notice; pricing and limits are the vendor's to change.
The vendor's own failure list. The pattern in the fixes: let code do what code does best.
| Edge | Documented behavior | Do instead |
|---|---|---|
| Literal reading | Answers the written question, not the intent | State the exact condition |
| Counting | Does not count reliably | Count in code |
| Dates & durations | Ordering and gaps are unreliable | Compare in code; pass the verdict in |
| Numeric encodings | Hex and RGB worse than words | Use semantic descriptions |
The page opens plainly: “Jev isn't perfect. Here are some jagged edges we are aware of with jev-1.13.” A vendor documenting its model's failure modes is exactly the observability-and-testability posture the manifesto claims — and it gives you a concrete test plan: each row is a category of cases for your own evaluation set before shipping.
Ask Jev the semantic part and keep the arithmetic in code: “does this message mention more than one order?” is a bad question; “which orders does it mention?” plus len() is a good one.
| Edge | Documented behavior | Do instead |
|---|---|---|
| Indirection | Multi-hop lookups degrade accuracy | Resolve references first |
| Irrelevant context | Accuracy falls as the state grows | Keep state lean |
| Adversarial content | Misleading framings can sway answers | Treat state as untrusted input |
| Cross-question invariants | Relations between answers aren't guaranteed | Ask one question; derive the rest |
These are jev-1.13 notes, not permanent truths.
Contradictory instructions: when instructions and criteria disagree, answers degrade — keep them aligned.
No text generation: by design rather than defect. When you need prose, pair Jev with an LLM — Jev decides, the LLM writes.
Recheck the jaggedness page when the model version changes; the list is versioned to jev-1.13.
The ambition, in the docs' words: “run it a million times in the background without a human co-pilot.”
Semantic checks on every LLM input, output, and tool call.
Route prompts between models; gate tool calls.
Check extractions, reasoning traces, tool calls.
Semantic search, scoring, and ranking over big datasets.
Faster than human perception — games, embedded UI.
Triage, claims, moderation, lead scoring, fraud signals.
Vendor-described applications, not independently benchmarked results.
Every entry has the same silhouette: a high-volume stream of small judgments, each cheap to make and cheap to check, feeding deterministic software that owns the consequences. None asks the model to converse, plan, or explain.
The inverse is a useful negative test: if your task needs one large, novel, multi-step judgment — a legal brief, an architecture decision — that is reasoning-model territory, and Jev's role shrinks to verifying pieces of it.
The docs' own phrasings are worth quoting: guardrails “at a fraction of the cost”; “build a custom router that chooses which LLM receives each prompt”; “replace or supplement embeddings in RAG pipelines”; moderation against “company-specific, nuanced criteria.” The cookbook has worked recipes for most of them — re-ranking, classifying RAG passages, double-checking citations, function calling, and guardrails.
Six habits, one principle: code owns control flow; the model makes narrow judgments.
No agent loop where an if suffices.
Irrelevant context costs accuracy and money.
Nested JSON, descriptive names, exact paths.
One judgment per question. The guide's most important idea.
Many small questions, one request.
Deterministic rules or weighted sums.
The guide's wording: “Keep deterministic work in code. It is reliable and cheap.” Atomization is “probably the most important concept in this guide. Broad questions hide several judgments behind one answer.” Compose with “deterministic rules or weighted sums” — or feed the probabilities into a classical ML model. Reference exact values by path, like support.tickets[0].message.
Instead of one broad question — “should we refund this customer?” — the docs' refund example builds a state holding the customer's message, the transactions, and the policy, then asks independent questions together: does the message request a refund? (Noul) · is the order within the policy window? (Noul) · how frustrated is the customer? (Score) · which team owns this? (Choice).
The refund decision is then an auditable boolean expression in your code. When it misfires you can see which judgment was wrong — the observability a broad question destroys.
Chapter 04 of 05
The most common objection — and a fair one. Typed classification has a decade of published practice behind it.
BERT set the template in 2018; bert-base-uncased still sees 47M+ downloads a month.
Community verdict: “the industry rediscovering classification models.”
KDnuggets: “Classification is not new. Intent detection is not new.”
Same piece: “the architecture and product around it may be new.”
The transparency gapArchitecture, parameter count, training data, and the RLCD reward are undisclosed. So “just a fine-tune” can be neither confirmed nor refuted.
Zero-shot classifiers such as facebook/bart-large-mnli and GLiNER-style encoders “already do something similar.” Reddit and GeekNews verdicts include “essentially BERT with more data” and “I implemented the same functionality with BERT-family models years ago.” KDnuggets also places Jev outside “the same category as GPT, Claude, Gemini, or other frontier LLMs.”
Each ingredient has a citable history: fine-tuning pretrained encoders for classification (Devlin et al., 2018), zero-shot classification via NLI (Yin et al., 2019), and post-hoc calibration including temperature scaling (Guo et al., 2017).
What is documented: one shared set of weights for all accounts, no per-customer fine-tuning, and the task chosen at runtime by the question text — an operational difference from one-classifier-per-task, whatever the architecture turns out to be. The fair restatement of the claim is narrower than the marketing: not a new kind of AI, but engineering on a known recipe.
Much of the launch coverage came from influencers. Strip the adjectives and measure the claims.
It means zero out-of-schema outputs. Jev can still choose Billing when the answer was Technical.
The 0% is a definition, not a measurement.
Scored against frontier-model reference answers, not ground truth.
On real accuracy, the published verdict: “we do not really know yet.”
The flagship 0% hallucination figure “was plotted at zero because schema conformance is guaranteed by design” — the model cannot emit a value outside the options you defined, so the metric is true by construction and says nothing about whether the chosen option is right.
TypeSafe's accuracy benchmark scores Jev against frontier-model reference answers. That measures agreement with another model, not correctness, so it cannot support a general accuracy claim.
Neither observation says the product is bad — they say these two numbers are not evidence of quality. The next slide is someone actually measuring.
Phishing detection, 2,000 synthetic emails.
| Approach | Accuracy | |
|---|---|---|
| Jev · one Noul question | 62.6% | zero-shot |
| Claude Haiku 4.5 · one question | 81.3% | zero-shot |
| A two-line regex | 91.8% | dataset nearly separable by construction |
| Haiku · five-question composite | 93.2% | |
| Jev · five questions + fitted weights | 95.0% | p = 0.063 vs Haiku — not significant |
What survives measurementCost and speed do — 12–27× cheaper, 2.9–5× faster than Haiku. Out-of-box accuracy and phrasing robustness don't.
On phrasing robustness: the same question asked as a Noul returned 0.22 — asked as a yes/no Choice, “no” at probability 0.99. Measured calibration is task-dependent: ECE 0.154 on phishing (worse than Haiku's 0.097), 0.0712 on tool-call classification.
The decomposition workflow — five narrow questions, weights fitted on 1,000 labels — is where Jev won, and it pushes the logic onto the code side where it can be audited. That is the product's thesis working as advertised; it just required labeled data and fitting, like any classifier.
The regex's 91.8% cuts both ways: this synthetic benchmark was easy, which weakens both Jev's win and the criticism drawn from its single-question loss. GeekNews' top comment offers the charitable reading: “the skill of clearly communicating what something does and where it's useful is equally part of innovation.” Others noted the practical value plainly: solving a classification problem “in minutes instead of weeks.”
The strongest version of the criticism ships weights: same three primitives, Apache 2.0, recipe published.
| jev-1.13.0 | laya | |
|---|---|---|
| Primitives | choice · score · noul | the same three |
| Weights | closed — hosted API | open — 421M, ModernBERT-large |
| RLCD recipe | undisclosed | published |
| Latency | ~239 ms median | ~33 ms local GPU |
| Languages | English primary | 100+ |
| Price | $0.042 / Mtok | free — your hardware |
Laya (Convai Innovations) serves the three primitives with calibrated probabilities in one ~33 ms forward pass, 7.2 ms/question batched, and publishes its recipe: proper-scoring-rule reward with Gaussian logit noise. Its question schema is near-identical to Jev's.
From the Hugging Face model card (checked 27 September 2026): three checkpoints (English root 421M on ModernBERT-large, a 322M mmBERT multilingual variant, a typed-decisions fine-tune), 103–332 questions/second on a Tesla T4, and 45 of 51 tested languages above chance. Jev's ~239 ms median latency is from the independent test on the previous slide.
Zero-shot, Laya's base checkpoints score below the majority-class baseline. The headline number is a fine-tuned checkpoint, and it ships overconfident.
Both things are trueA 421M fine-tuned encoder can serve this API — and out-of-box calibrated quality is the hard part. Either way, Laya is the baseline to benchmark Jev against.
By Laya's own model card, its base checkpoints score 0.362 zero-shot against a 0.461 majority-class baseline — near random. The headline 0.766 is a fine-tuned checkpoint; ECE runs 0.213–0.466 before temperature refitting, and its noul “can follow its option labels instead of the state.”
The circulating head-to-head (accuracy 0.766 vs 0.727; ECE 0.081 vs 0.246) is Laya's vendor benchmarking its fine-tuned model on its own dataset — treat it exactly like TypeSafe's 68%: an advert until independently reproduced.
One further allegation in the GeekNews comments — that a sales-conversion demo model had outcome labels leak into training — is from a single anonymous commenter and is unverified; it is mentioned only so you know it exists.
If the audience takes one thing from this chapter: run Laya as your baseline before paying for anything.
Chapter 05 of 05
Near-identical schemas. What changes is who runs the weights.
$ pip install typesafe-sdk
from typesafe_sdk import Choice, Noul, TypeSafeClient
client = TypeSafeClient() # needs TYPESAFE_API_KEY
state = "Billed twice for March — refund today or we cancel."
r = client.system_one(state=state, questions={
"department": Choice(
instructions="Which department should handle this?",
criteria={"billing": "invoices, payments, refunds",
"technical": "bugs, outages, system errors",
"other": "everything else"}),
"churn_risk": Noul(
instructions="Does the user threaten to cancel?"),
})
r.answers["department"].choice # "billing"
r.answers["department"].confidence # 0–1 — gate on it
r.answers["churn_risk"].noul # probability of yes$ pip install laya
from laya import Router
router = Router() # downloads checkpoints on first use
state = "Billed twice for March — refund today or we cancel."
result = router.predict(state, {
"department": {
"type": "choice",
"instructions": "Which department should handle this?",
"criteria": {"billing": "invoices, payments, refunds",
"technical": "bugs, outages, system errors",
"other": "everything else"}},
"churn_risk": {
"type": "noul",
"instructions": "Does the user threaten to cancel?"},
})
result["answers"]["department"]["choice"] # "billing"
result["answers"]["churn_risk"]["noul"] # e.g. 0.892Why this mattersAn interface this close means you can A/B both on a thousand labeled examples of your traffic in an afternoon.
The Jev snippet is the documented quickstart pattern (TypeSafeClient().system_one with Choice/Score/Noul objects) with this deck's billing example substituted for the docs' Stripe ticket. The Laya snippet follows its model card's Router.predict usage; the 0.892 noul is the card's own example output. Jev-side values will vary run to run — they are labels here, not reproduced results.
Both SDKs are pip-installable today. Laya also exposes laya.load("convaiinnovations/laya") for pinning a single checkpoint instead of the auto-routing Router.
The fine print each vendor states: Laya wants a fine-tune and a temperature refit before production; Jev wants confidence thresholds validated against your own outcomes.
| # | Idea | In one sentence | Status |
|---|---|---|---|
| 1 | System One | Typed decisions with probabilities — never text. | documented |
| 2 | One call | One state, many independent questions, in parallel. | documented |
| 3 | Choice · Score · Noul | Which option — which level — is this true. | documented |
| 4 | Calibration | RLCD targets honest probabilities; recipe unpublished. | goal only |
| 5 | Confidence | A shape summary — not ℙ(correct). | documented |
| 6 | Thresholds | Your numbers, set per action from costs. | your job |
| 7 | Jagged edges | Count, compare, and compose in code. | documented |
| 8 | Verification | Valid doesn't mean true — log outcomes. | your job |
| 9 | The case against | Prior art is real; Laya is the natural baseline. | contested |
Define the question precisely, keep control flow in code, and test the probabilities against your own outcomes.
This deck answers what Jev is and how to hold it: the product shape, the primitives, the documented limits, the build method. The companion deck, RLCD — The Mathematics of Calibrated Decisions, answers why the probabilities can be trusted and what to do with them: calibration, proper scoring rules, and expected-cost decisions.
Presented together, run this one first — the mathematics lands better once the audience has seen a real answer object with a probability distribution in it.
Sources: TypeSafe documentation (introduction, System One, primitives, confidence, models, jaggedness, use-case map, build guide), reviewed 27 September 2026, plus the third-party material listed on the next slide.
Provenance key in full: documented — reproduced from TypeSafe documentation; illustrative — local widgets and teaching devices, not model output; undisclosed — stated goals whose mechanisms the sources do not specify; third-party — external tests, commentary, and the Laya model card, quoted for balance, not endorsed.
Example values (the shoe ticket's 1.0, the bug's 1.43 / 0.35, the 0.99 noul) are reproduced from the documentation pages above and may change as the docs evolve.
Companion deck: RLCD — The Mathematics of Calibrated Decisions, which reviews the lecture video in this folder.