00 Jev / A mathematical field guide

From probability
to decision.

RLCD · Reinforcement Learning for Calibrated Decisions

What an 80% forecast means, why honest probabilities can win, and how uncertainty becomes an action.

THE CALIBRATION TARGET
80%

Among comparable 80% forecasts,
about 8 in 10 events should occur.

Illustration · individual batches vary
01 Check the forecast02 Understand the incentive03 Price the decision

Eight mathematical ideas. Jev’s documented goal is distinguished from the teaching examples; the reviewed sources do not specify its training recipe.

01 Why we need probabilities

Why probabilities matter

A category tells your application what the model predicts. A probability tells it how uncertain that prediction is.

For two different support tickets, the model predicts the same category: returns.

Ticket A · predicted category: returns

98%

The model estimates a 98% chance that returns is the correct queue.

Ticket B · predicted category: returns

41%

Returns is the most likely category, but only has a 41% chance of being correct.

Your application can use these probabilities and the cost of mistakes to choose whether to route automatically or request a human check. The model’s category prediction is not the action.

Where RLCD comes inThose numbers are useful only if they reflect real outcomes. RLCD targets calibration: across many predictions assigned 98%, about 98% should be correct.

This example explains why probabilities matter. The next slides explain what makes them calibrated.

Why can 41% still be the selected category?

This example uses one model on two different tickets. On Ticket B, the probabilities could be returns 41%, billing 35%, and shipping 24%. Returns is the largest individual probability, so the model selects it. That does not make returns more likely than all the alternatives combined.

The model predicts the category. Your application chooses what to do with that prediction. It may route the ticket automatically or request a human check, depending on the probability and the costs of those actions.

Purpose of this slide: motivate the need for probabilities. RLCD’s stated goal is to make the probabilities calibrated; this routing example does not describe its training algorithm.

02 Where RLCD sits

What training rewards

Reinforcement learning optimises whatever you reward. So the interesting question about any RL method is simply: what is it rewarding?

Human feedback

Rewards answers people prefer.

Can reflect usefulness, correctness, style, and other human preferences.

Verifiable reward

Rewards answers that pass a check.

Optimises for what a test can confirm — the code compiles, the sum is right.

RLCD

Targets calibrated probabilities.

Optimises for forecasts whose stated rates actually occur.

Teaching notes
These are distinct goals, not rival superpowers. They overlap, and a system can pursue more than one. An answer can be preferred, correct, and accompanied by an honest probability — or any two out of three.
Explanation & analogy
Analogy Three different exams for a doctor. One grades bedside manner. One grades the diagnosis. RLCD grades a third thing entirely — “when you say you're 70 % sure, are you right about 70 % of the time?” You can pass any of these and fail the others.
How a reward changes model behavior

During reinforcement learning, a training procedure adjusts a model to increase its expected reward. What the reward measures influences what the model learns to do.

A human-preference reward evaluates which response people favor. A correctness reward evaluates whether a prediction passes a check. A probability-sensitive reward can distinguish 60% from 99%, even when both forecasts select the same category.

These goals can overlap. The next slide compares two models to show what a category-only reward misses.

03 Training: what should the reward measure?

Reward the probability

Two models classify the same ticket as returns. They attach different probabilities to that prediction.

Model A

60%

Model A predicts returns, with a 60% estimated chance of being correct.

Model B

99%

Model B predicts returns, with a 99% estimated chance of being correct.

If training rewards only the correct categoryIf returns is correct, both models receive +1. If it is incorrect, both receive 0. This reward cannot distinguish a realistic probability from an exaggerated one.

A possible alternative: score each model’s probability against the observed outcome. Brier loss, L = (p − y)², is one way to do this; lower loss means a better forecast score.

This illustrates the incentive behind calibrated probabilities. It does not establish which reward Jev uses.

Worked example: Model A vs Model B

Here p is a model’s probability that returns is the correct queue. The observed label supplies y = 1 when returns is correct, and y = 0 otherwise. Neither model chooses the observed label.

Observed outcomeModel A · p = 0.60Model B · p = 0.99
Returns is correct (y = 1)(0.60 − 1)² = 0.16(0.99 − 1)² = 0.0001
Returns is incorrect (y = 0)(0.60 − 0)² = 0.36(0.99 − 0)² = 0.9801

Model B receives a lower loss when correct, but a much higher loss when wrong. One ticket does not tell us which probability was better calibrated. We must consider both possible outcomes and how often each occurs.

Suppose returns is truly correct in 60% of comparable cases, and the models keep reporting these probabilities. Model A’s expected loss is 0.60 × 0.16 + 0.40 × 0.36 = 0.24. Model B’s is 0.60 × 0.0001 + 0.40 × 0.9801 = 0.3921. Model A performs better on average because its probability matches the true rate.

Training could use reward = −loss: −0.24 is higher, and therefore better, than −0.3921. The later derivation shows why matching the true rate minimizes expected Brier loss.

04 Calibration, before any notation

What 80% means

One model predicts an 80% probability of returns for each of ten tickets. In this illustration, returns is the correct queue for eight tickets and incorrect for two.

Ten forecasts of 80 %

event happened  ×8 did not  ×2

Calibration is a property of a group, never of one case. It asks whether the agreement persists across similar forecasts: among all the 80 % predictions, does the event occur roughly 80 % of the time?

Not every batch of ten contains exactly eight. Samples fluctuate — and calibration cannot tell you which individual case will be the exception. That information does not exist in the number.
Explanation & analogy
Analogy A forecast, not a promise. “80 % chance of rain” is never falsified by one dry Tuesday. A hundred comparable days of which only forty were wet would be strong evidence of miscalibration. Judge the number the way you'd judge a weather service: over a season, not over an afternoon.
What is being counted?

Each dot represents a different ticket, not a different model. A filled dot means the observed category was returns. An empty dot means the observed category was something else.

The model’s 80% prediction does not identify which tickets will be mistakes. Calibration asks whether the frequency matches the prediction across many comparable cases. Ten tickets make the idea visible, but are not enough to establish calibration reliably.

05 The calibration equation

Calibration in symbols

Y

Y = 1 when returns is the correct queue, Y = 0 otherwise. The outcome, after the fact.

p̂  “p-hat”

The model’s predicted probability that returns is the correct queue. A random variable — it changes case to case.

p

One particular forecast value we want to talk about, e.g. the number 0.8.

Eq. 1
ℙ( Y = 1 | p̂ = p ) = p

In wordsLook at every case where the model happened to say p. Among exactly those cases, the event's true rate should be p. Forecasts of 80 % should approach 80 % frequency across comparable cases.

Worked detail

Concrete exampleSuppose the model makes 100 predictions, and for each one it says “there is a 70 % chance this is correct.” If the model is well calibrated, roughly 70 of those 100 predictions should actually be correct. That is Equation 1 with p = 0.7 — and “roughly” matters: sampling uncertainty and dependence between cases affect the strength of the evidence.

Teaching notes
Explanation & analogy
Analogy Read the vertical bar | as “restricted to”. The equation says: filter your history down to the rows where the model wrote 0.8 in the forecast column, then count how often the outcome column says yes. If the answer is 80 %, this forecast value is calibrated.
Why it must be a conditional. Averaging over all predictions would let a model hide: wild overconfidence on one group could cancel wild underconfidence on another. The condition p̂ = p forces the promise to hold separately at every value the model is allowed to say.
Read the equation from left to right

ℙ means probability. Y = 1 means the event occurred. The vertical bar | means “given that.” p̂ = p restricts attention to cases where the model reported the value p.

Read the equation as: “Given that the model reported p, the actual chance of the event is p.” For p = 0.8, the target event rate is 80%. This is a population condition, not a requirement that every finite batch match exactly.

For continuously valued forecasts, the precise statement is 𝔼[Y | p̂] = p̂ almost surely. In a finite dataset, we estimate agreement using nearby probabilities, as on the next slide.

06 Checking it

Check forecasts against outcomes

Group nearby held-out forecasts into bins. For each nonempty bin, compare the mean prediction with the observed event frequency.

Bk

Bucket number k — the set of cases whose forecasts fell in that range.

|Bk|

How many cases landed in the bucket. Just a count.

∑  then  ÷

Sum the values, then divide by the number of cases. A bar over a symbol means “average”.

Eq. 2
p̄k = 1|Bk| ∑i∈Bk p̂i   vs   ȳk = 1|Bk| ∑i∈Bk yi

In wordsWhat the model said, on average, in this bin — against how often the event actually happened in that same bin. Plot one point per bin: said on the horizontal axis, happened on the vertical. Perfect population calibration corresponds to the diagonal; finite samples fluctuate around it.

Worked detail

Concrete exampleA bin holds five forecasts: 0.78, 0.80, 0.82, 0.81, 0.79. Their average — what the model said — is p̄k = 0.80. The outcomes were 1, 1, 0, 1, 0 — the event happened 3 times out of 5, so ȳk = 0.60. The model said 80 % but the event occurred 60 % of the time: plot the point (0.80, 0.60), below the diagonal — overstated in this sample. Five cases alone do not establish miscalibration.

Teaching notes
Read the geometry. → is the promise; ↑ is reality. Below the diagonal = overstated probability — the event probability exceeds the observed rate. Above it = understated. At low probabilities, overstated event probability need not mean overconfidence in the selected class.
Always check the counts. A bin with six cases tells you almost nothing: 4 out of 6 gives 0.67, and that could easily be luck. The same 0.67 from a bin of 6,000 cases is real evidence. A reliability diagram without sample sizes is a decoration, not proof.
Why use bins, and what does held-out mean?

A bin is a group of tickets with similar predicted probabilities, such as 75%–85%. The predictions need not repeat exactly for us to compare their mean with the group’s observed frequency.

Held-out cases were not used to fit the model or tune the calibration procedure. Checking these cases helps assess whether the probabilities remain useful on new data. Bin choices, sample size, and dependence between cases affect the evidence.

07 The reliability diagram

Reading the reliability diagram

Nine equally weighted synthetic forecast groups. Drag right: extreme predictions become overconfident and observed rates move toward 50%.

Teaching notes
Set the slider to zero. Every point sits on the diagonal: what was promised is what occurred. This is the target RLCD is named for.
Push it right. The curve flattens toward the base rate. A model saying “95 %” while the event happens 70 % of the time may still select the right answer — it is overstating how sure it is, and that is the error that breaks automated decisions.
How to read one point

A point at (0.80, 0.60) means the model predicted 80% on average for that group, while the event occurred in 60% of cases. The point lies below the diagonal because the event probability was overstated in that group.

Moving the slider changes the synthetic observed frequencies, keeping the forecast positions fixed. This lets you inspect different calibration patterns; it is not a simulation of Jev training.

When extreme forecasts are overconfident, high probabilities can lie below the diagonal and low probabilities above it. For example, a 10% event forecast with a 30% event rate is too certain that the event will not occur.

08 An honest pause

Goal known, recipe unspecified

TypeSafe names RLCD's goal — calibrated decisions. Reviewed sources do not disclose its exact reward, loss, or update algorithm.

So read the next few slides correctly

The scoring-rule examples that follow are a teaching illustration of how a calibration reward could work — standard, well-understood mathematics that makes the incentive visible. It is not a reconstruction of Jev's implementation, and you should not cite it as one.

An illustrative training loop

context → model → probability → observed outcome → feedback

In wordsContext enters a model, a probability comes out, and the outcome that eventually occurs supplies the feedback signal. This is a conceptual example, not a claim about Jev’s training pipeline.

Explanation & analogy
Why bother with an illustration at all Because the interesting question — what kind of reward could possibly make honesty the winning strategy? — has a beautiful and completely general answer. Understanding that answer tells you what RLCD is aiming at, even without the source code. That is worth four slides.
09 Setting up the training objective

From loss to reward

A loss measures prediction error: smaller is better. A reward scores the model’s behavior: larger is better.

R = −L   ⟹   maxθ 𝔼[R] = maxθ (−𝔼[L])

The same preference, with the sign reversedMaximizing expected negative loss is equivalent to minimizing expected loss.

Model A · reports 60%

Expected loss: 0.24
Expected reward: −0.24

Model B · reports 99%

Expected loss: 0.3921
Expected reward: −0.3921

These reuse slide 4’s example, where the true rate is 60%. Model A has the lower loss and the higher reward: −0.24 > −0.3921.

θ denotes the model’s trainable parameters. 𝔼 means an average over possible outcomes, weighted by their probabilities. This defines an objective, not the algorithm used to update θ.

Why use an average instead of a single outcome?

A model reporting 99% can be correct on one ticket and receive an excellent score. That does not establish that 99% is a realistic probability. Expected reward includes both correct and incorrect outcomes at their underlying rates.

Training estimates the objective from data. Finite data, imperfect labels, model limitations, and optimization can prevent it from reaching the population optimum. A proper scoring rule supplies an incentive; it does not by itself guarantee calibration.

10 A proper scoring rule

Scoring a probability

Let p be the model’s probability that returns is correct, and y be the observed result: 1 if returns is correct, 0 otherwise.

Eq. 3
L = (p − y)2

In wordsThe model’s reported probability minus the outcome — which is 0 or 1 — squared. Smaller is better. (This is the single binary squared-error convention; some texts double it.)

Predict 80 %, and wait

OutcomeyLoss (p − y)²
Event happens1(0.8 − 1)² = 0.04
No event0(0.8 − 0)² = 0.64

For p = 0.8, a non-event incurs 16 times the loss of an event. To understand the incentive, average over both outcomes, as the next slide does.

Explanation & analogy
Compare the same model’s possible forecastsWhen y = 1, reporting 99% gives loss 0.0001. When y = 0, that same report gives loss 0.9801. A report of 50% gives loss 0.25 in either case. Which report is better on average depends on how often the event actually occurs; 50% is optimal when the true rate is 50%.
Loss is a score, not a dollar cost

At p = 0.8, an observed event produces (0.8 − 1)² = 0.04. A non-event produces (0.8 − 0)² = 0.64. These are forecast losses, not probabilities and not business costs.

The model reports p before the outcome is known. The outcome y is then used to evaluate that report. The application’s financial cost of acting on the prediction is a separate question, covered later.

11 The central result

Why the true rate wins

One outcome isn't enough to judge a forecast. Let p be the model’s report and q the true event rate for cases like this one. Events occur with probability q, non-events with 1 − q. Average the two losses.

Eq. 4
𝔼[L] = q(1 − p)2 + (1 − q)p2   =   p2 − 2pq + q
𝔼[L] = (p − q)2 + q(1 − q)
forecast-dependent  ·  fixed for this q

In wordsComplete the square and the expected loss splits into two pieces. The second is irreducible — it is the outcome variance for this fixed q. Changing your report cannot reduce it; better information may change q. The first is the only part you control, and it is a squared distance from the truth.

The conclusion

A squared distance is smallest when it is zero. So expected loss is minimised at p = q.

Truthful reporting is the optimal strategy — not because we asked nicely, but because every alternative is arithmetically worse. The unique optimum makes Brier strictly proper. This is a population result; finite training does not guarantee calibration.

Explanation & analogy
What the result does and does not sayFor a fixed true rate q, reporting p = q has the smallest expected loss. This is a statement about the scoring rule. It does not imply that the model knows q, or that training will discover it exactly.
The derivation, one step at a time

1. Weight both outcomes. The event occurs with probability q, giving loss (1 − p)². The non-event occurs with probability 1 − q, giving loss p². Therefore 𝔼[L] = q(1 − p)² + (1 − q)p².

2. Expand. q(1 − 2p + p²) + p² − qp² = p² − 2pq + q.

3. Complete the square. Add and subtract q²: (p² − 2pq + q²) + q − q² = (p − q)² + q(1 − q).

4. Find the minimum. Holding q fixed, the second term does not depend on the model’s report. The first term is nonnegative and is zero only when p = q.

12 Equation 4, alive

Explore the scoring rule

The horizontal axis is the model’s reported probability p; the vertical axis is its expected loss. Set the assumed true rate q, then move p to compare reports.

The floor never moves as you drag p. That dashed line is q(1 − q) — the irreducible term. You cannot get under it by forecasting cleverly.
The valley sits at p = q, always. Slide q and the whole parabola follows. The optimum tracks the truth because it is the truth.
Note the symmetry. Overstating by 0.15 costs precisely what understating by 0.15 costs. Brier punishes distance, not direction — it has no opinion about optimism.
Try q = 0.80, then change p

Keep q at 0.80. Set the model’s report p to 0.80: expected loss is 0.16, the minimum. Move p to 0.95: loss rises to 0.1825. Move p to 0.50: loss rises to 0.25.

You are controlling the assumed population and the model’s report for illustration. In deployment, the true conditional rate q is generally unknown; we evaluate predictions using observed outcomes.

13 The same result, in numbers

Three forecasts, one population

Take a toy population whose true rate is q = 0.8. The irreducible term is 0.8 × 0.2 = 0.16. Now compare three ways of reporting it.

Model reportSquared distance (p − q)²Irreducible q(1−q)Expected Brier loss
p = 0.80 honest (0.80 − 0.80)² = 00.16 0.1600best
p = 0.95 overconfident (0.15)² = 0.02250.16 0.1825+14 %
p = 0.50 hedging (0.30)² = 0.090.16 0.2500+56 %
Same underlying outcome distribution in all three rows. The world did not change — only the reported number did. 80 % wins in expectation purely by matching the true rate. Note also that hedging toward 50 % is punished harder here than overconfidence: the penalty depends on distance from q, not on the direction of the error.

These are calculations, not measured Jev results. The value of the table is the mechanism it exposes — that a squared-error reward on the probability itself makes honesty the profit-maximising move. The reviewed sources do not specify whether Jev uses Brier, log loss, or another objective.

What is being compared?

Each row represents a possible model forecast for the same kind of case. These could be three models, or three candidate reports from one model. The true rate q = 0.8 stays fixed across the comparison.

The expected loss of 0.16 does not mean 16% classification errors. Brier loss measures the quality of probabilities. Even the optimal probability report has a nonzero expected loss when outcomes remain uncertain.

14 Back to documented ground

Jev’s three output types

Your application defines the question and what the answers mean. Jev returns one of three shapes.

Choice

Selects among supplied alternatives — returns, billing, shipping.

∑c pc = 1

The option with the highest probability is the selected category. Your application separately decides how to act on it.

Noul

A binary proposition. “Does the customer request a person?”

yes .50  ·  no .50  →  p = 0.5

0.5 means uncertain, not medium. It is not an intensity dial, and there is no separate confidence field here.

Score

Ordered levels you define.

0 cosmetic  ·  1 workaround  ·  2 blocking

The lecture’s example assigns probabilities 0, 0.7, 0.3. Next slide turns that into one number.

Noul is the documented name. The API identifier is noul; its value is the probability of “yes”. TypeSafe reference ↗
Choose the question before choosing an action

Choice: “Which queue should handle this ticket?” The model predicts a category and returns probabilities over the options.

Noul: “Does this customer request a human?” The model returns the probability of yes. A value of 0.5 means uncertainty about that proposition, not a halfway request.

Score: “Which severity level describes this bug?” The model returns a distribution over rubric levels and its mean index. The application can use these predictions in a routing or prioritization rule.

15 Reading a Score

Reading a Score

Ordered levels carry an index. Multiply each index by its probability and add them up.

Eq. 5
𝔼[level] = ∑i i · pi  =  0×0 + 1×0.7 + 2×0.3 = 1.3

In wordsA probability-weighted average of the level numbers. With most of the mass on “workaround” and 30% on “blocking”, the expectation lands just above 1.

What 1.3 is not

It is not a percentage. It is not a physical measurement. There is no level 1.3 — nothing in your system is “1.3 broken”.

It is an expected index: a weighted average under the model’s predicted distribution, not necessarily the true outcome distribution.

Explanation & analogy
Analogy “The average household has 1.8 children.” No household does. The number is still useful — for stocking a school, for sizing a budget — precisely because you remember it describes a population, not a case.

Practical consequence: ordering alone does not establish equal spacing. The mean summarizes the chosen index scale; it does not measure severity in natural units. Never compute this over an unordered Choice — the expected index of returns/billing/shipping is nonsense.
Work through the weighted average

The model assigns 0% to level 0, 70% to level 1, and 30% to level 2. Multiplying each index by its probability gives 0, 0.7, and 0.6. Their sum is 1.3.

That number summarizes the model’s uncertainty over defined levels. It does not mean the bug is “130% severe,” nor does it prove the model’s probabilities are calibrated. Different probability distributions can have the same mean, so inspect the distribution when the distinction affects your action.

16 A trap worth naming

Probability vs. confidence

The lecture’s Choice example uses this distribution, and alongside it a separate field called confidence.

Choice probabilities

returns0.60← the winner
billing0.38
shipping0.02

Reported confidence

0.39

Not 0.60. Not the winning probability. Not anything you can derive from the winner alone.

The lecture’s Score example reports confidence 0.54 — again unequal to any single probability in it.

Eq. 6
confidence need not equal pwinner     confidence is not defined as ℙ(the answer is correct)

In wordsDocumentation describes confidence as summarising the shape of the distribution — how concentrated it is — without specifying the formula. Two non-negotiables follow: don't invent a formula, and don't read it as a probability of correctness.

Teaching notes
The obvious guesses don't fit. Normalised entropy gives 0.32 for that distribution, the top-two margin gives 0.22, the sum of squared probabilities gives 0.50 — none is 0.39. Which is exactly the lecture's point: if you reverse-engineer a formula and ship logic that depends on it, a silent change upstream breaks you.
And concentration isn't correctness. A model can be confidently, concentratedly wrong. Use the probabilities for decision rules, and if you want to route on confidence, validate that routing against your own outcome data first.
Which number answers which question?

In this lecture example, 0.60 answers “What probability does the model assign to returns?” The separate 0.39 confidence value summarizes the distribution’s shape.

Do not interpret confidence 0.39 as “the selected category has a 39% chance of being correct.” The model’s probability for that category is 0.60. Whether 0.60 matches actual correctness rates still needs evaluation on outcomes.

A distribution concentrated on the wrong category can have high confidence. Confidence may help route cases, but its usefulness must be checked against the task’s observed results.

17 From probability to action

From prediction to action

Here the predicted event changes: p is the model’s probability that a transaction is fraudulent. The application chooses approval or review. Under this illustrative cost model: approving fraud costs $100, approving a legitimate transaction costs nothing, and review costs $5 and prevents all fraud loss.

Eq. 7
Capprove = 100p + 0(1 − p) = 100p     Creview = 5
review wins  ⟺  100p > 5  ⟺  p > 0.05

In wordsExpected cost is each outcome's cost weighted by its probability. Compare the two numbers and take the smaller one. The whole decision collapses to a single threshold — p* = Creview ⁄ Cfraud, here 5⁄100.

p = 0.02

100 × 0.02 = $2  vs  $5

Application: approve. Its expected cost is lower under these assumptions.

p = 0.05

100 × 0.05 = $5  vs  $5

A tie. Exactly the threshold; the cost rule is indifferent.

p = 0.12

100 × 0.12 = $12  vs  $5

Application: review. This saves $7 in expectation under the stated assumptions.

Real review is imperfect. It doesn't catch everything and it isn't free of knock-on cost. Measure your own review cost and effectiveness before wiring this up.
Why multiply the probability by $100?

If the application approves, it loses $100 when fraud occurs and $0 otherwise. With fraud probability p, expected approval cost is p × $100 + (1 − p) × $0 = $100p.

At p = 0.02, the model estimates a 2% chance of fraud. Approval costs $2 on average across comparable cases, while review costs $5. An individual transaction does not incur exactly $2: the average combines possible $100 and $0 outcomes.

At p = 0.12, expected approval cost is $12. Under the assumption that a $5 review prevents all fraud loss, review is cheaper. Both transactions are more likely legitimate than fraudulent, yet the application’s preferred actions differ.

Setting $100p = $5 gives p = 0.05. Below that value, approval is cheaper; above it, review is cheaper; at it, the expected costs tie.

18 The threshold, alive

Choose a cost threshold

Two expected-cost lines. Where they cross is your automation threshold. Change the costs and watch it move.

The lesson to take to your team

Probabilities describe uncertainty. Costs determine action. Keep them in separate parts of your code.

The model estimates fraud probability. The application supplies the costs and chooses the lower-cost action. Changing review costs can change the action without changing the prediction.

There is no universal automation threshold. Not 0.9, not 0.95, not “high confidence”. The correct threshold is a ratio of your costs, and it moves when your costs move. Anyone who quotes you a fixed number across products needs to justify the cost assumptions.
Try changing the application’s costs

With missed fraud costing $100 and review costing $5, the threshold is 5%. Raise the review cost to $10: the threshold becomes 10%, so review is justified only at higher fraud probabilities.

Keep review at $5 but raise missed-fraud cost to $200: the threshold falls to 2.5%. More expensive mistakes justify reviewing lower-risk transactions.

The sliders change the application’s cost assumptions, not the model’s probability. If review costs more than the maximum possible fraud loss, approval is cheaper for every probability in this simplified model.

19 The limit of the goal

Calibration is not enough

Suppose half your cases are positive. A model that says 50 % to everything is perfectly calibrated — among its 50 % forecasts, the event does occur half the time. It provides a useful base-rate forecast, but cannot distinguish one case from another.

Model A · constant forecast

always p = 0.5

Calibrated ✓  ·  Expected Brier 0.25

Model B · two calibrated groups

p = 0.2  and  p = 0.8

Calibrated ✓  ·  Expected Brier 0.16

Eq. 8
𝔼[L] = reliability − resolution + uncertainty
constant: 0 − 0 + 0.25 = 0.25     two groups: 0 − 0.09 + 0.25 = 0.16

In wordsBrier splits three ways. Reliability is miscalibration — both models score 0, both are honest. Uncertainty is the base rate's own noise — identical for both. The difference is entirely resolution: how much the conditional event rates differ across forecast groups, while staying honest.

Exact population definition

Population definitionLet m(p̂) = 𝔼[Y | p̂] and μ = 𝔼[Y]. Reliability = 𝔼[(p̂ − m(p̂))²]; resolution = 𝔼[(m(p̂) − μ)²]; uncertainty = μ(1 − μ). Arbitrary finite bins only approximate these population terms.

So the constant baseline was never miscalibrated — it simply offered no discrimination between cases. Calibration is necessary and not sufficient; you want a model that is honest and discriminating. (This three-way split is standard forecast-verification mathematics, included to make the lecture's point precise — it is not part of any stated Jev method.)

Two models evaluated on the same population

Imagine 100 cases. Model A reports 50% for every case, and 50 cases are positive. Model A is calibrated at its single forecast value, but it cannot distinguish higher-risk from lower-risk cases.

Model B reports 20% for 50 cases and 80% for the other 50. In this illustrative sample, the first group has 10 positive cases and the second has 40. Model B is calibrated in both groups, and there are still 50 positives overall.

Model B gives the application more useful distinctions. Its Brier score is 0.16, versus Model A’s 0.25. That improvement comes from separating groups whose actual outcome rates differ, not merely from making predictions more extreme.

20 Before you ship

Check predictions in practice

Valid does not mean true

Type safety prevents invalid schema answers — you will never get shippng or a stray paragraph where an enum belongs.

It does not prevent wrong valid choices. Ambiguous inputs and missing information remain exactly as hard as they were.

Yesterday's calibration may not survive tomorrow

Evaluate on held-out cases and subgroups — a model can be calibrated overall and badly miscalibrated for one customer segment. Log outcomes, check sample sizes, monitor distribution shift.

The loop that has to exist in your system

context

unstructured input + a defined question

→

model prediction

category + probabilities

→

application action

compare expected costs; act or ask

→

outcome

what actually happened

The arrow back from “outcome” is the one teams forget to build. Without it you can never check whether the probabilities mean anything in your setting — and monitoring belongs in the system, not in a quarterly review.
Explanation & analogy
Why separate prediction from decision Keeping the model's probability and your cost rule in different places is what exposes the business trade-off. When finance asks why you auto-approve at 4 % but not 6 %, you want to point at a threshold in your own code — not shrug at a model.
What to log and compare

Record the predicted event, its probability, the application’s action, and the eventual observed outcome. Then compare forecasts with event frequencies on new cases and on relevant subgroups.

Keep the event definition consistent. If p predicts fraud, the evaluation outcome must indicate fraud; it should not silently become “a loss happened after our intervention.” Review may prevent losses, and outcomes may be missing for some actions. A biased set of observed labels can give a misleading calibration picture.

A valid output such as returns still can be the wrong category. Type safety checks the output’s allowed form; outcome evaluation checks its accuracy and calibration.

21 Take this away

The eight ideas together

#EquationSay it in one sentenceStatus
1ℙ(Y=1 | p̂=p) = pAmong cases where you said p, the event happens p of the time.standard
2p̄k vs ȳkBin held-out forecasts and plot promised against observed.standard
3L = (p − y)²Brier: square the gap between claim and outcome.illustrative
4𝔼[L] = (p−q)² + q(1−q)Honesty is optimal — the only term you control is a squared distance from the truth.illustrative
5𝔼[level] = ∑ i·pi = 1.3An expected index over ordered levels — not a measurement.documented
6confidence ≠ probability (in general)A distribution summary; validate any routing rule on outcome data.undisclosed
7Capprove=100p vs Creview=5Act when expected cost is lower; the threshold is a cost ratio.illustrative
8reliability − resolution + uncertaintyUseful forecasts distinguish groups with different outcome rates.illustrative

Ask what the probabilities mean, test them against reality, and connect them to the consequences of acting.

How the ideas connect

Prediction: the model supplies an event probability. Calibration: we ask whether those probabilities agree with observed frequencies. Training: a proper score illustrates why reporting accurate probabilities can earn a higher expected reward.

Decision: the application combines probabilities with action costs. Evaluation: we verify both calibration and the ability to distinguish cases, then monitor outcomes over time.

RLCD names Jev’s calibration-oriented training approach. The mathematical examples explain the goal and possible incentives; they do not identify its implementation.

22 Sources & scope

Sources and scope

Jev & RLCD

Mathematical foundations

Video reviewed: What is RLCD in Jev AI? · 9:59. Calibration 1:07–2:27; scoring 2:33–5:19; output types 5:24–7:12; decision costs 7:17–8:19.

The 0.60 / 0.38 / 0.02 distribution, confidence 0.39, and score 1.3 are lecture examples; current documentation uses different numerical examples. Brier loss and the cost model are teaching illustrations, not Jev benchmarks or a reconstruction of its training algorithm.

← → slides · N notes · F fullscreen