RLCD · Reinforcement Learning for Calibrated Decisions
What an 80% forecast means, why honest probabilities can win, and how uncertainty becomes an action.
Among comparable 80% forecasts,
about 8 in 10 events should occur.
Eight mathematical ideas. Jev’s documented goal is distinguished from the teaching examples; the reviewed sources do not specify its training recipe.
A category tells your application what the model predicts. A probability tells it how uncertain that prediction is.
For two different support tickets, the model predicts the same category: returns.
98%
The model estimates a 98% chance that returns is the correct queue.
41%
Returns is the most likely category, but only has a 41% chance of being correct.
Your application can use these probabilities and the cost of mistakes to choose whether to route automatically or request a human check. The model’s category prediction is not the action.
Where RLCD comes inThose numbers are useful only if they reflect real outcomes. RLCD targets calibration: across many predictions assigned 98%, about 98% should be correct.
This example explains why probabilities matter. The next slides explain what makes them calibrated.
This example uses one model on two different tickets. On Ticket B, the probabilities could be returns 41%, billing 35%, and shipping 24%. Returns is the largest individual probability, so the model selects it. That does not make returns more likely than all the alternatives combined.
The model predicts the category. Your application chooses what to do with that prediction. It may route the ticket automatically or request a human check, depending on the probability and the costs of those actions.
Purpose of this slide: motivate the need for probabilities. RLCD’s stated goal is to make the probabilities calibrated; this routing example does not describe its training algorithm.
Reinforcement learning optimises whatever you reward. So the interesting question about any RL method is simply: what is it rewarding?
Rewards answers people prefer.
Can reflect usefulness, correctness, style, and other human preferences.
Rewards answers that pass a check.
Optimises for what a test can confirm — the code compiles, the sum is right.
Targets calibrated probabilities.
Optimises for forecasts whose stated rates actually occur.
During reinforcement learning, a training procedure adjusts a model to increase its expected reward. What the reward measures influences what the model learns to do.
A human-preference reward evaluates which response people favor. A correctness reward evaluates whether a prediction passes a check. A probability-sensitive reward can distinguish 60% from 99%, even when both forecasts select the same category.
These goals can overlap. The next slide compares two models to show what a category-only reward misses.
Two models classify the same ticket as returns. They attach different probabilities to that prediction.
60%
Model A predicts returns, with a 60% estimated chance of being correct.
99%
Model B predicts returns, with a 99% estimated chance of being correct.
If training rewards only the correct categoryIf returns is correct, both models receive +1. If it is incorrect, both receive 0. This reward cannot distinguish a realistic probability from an exaggerated one.
A possible alternative: score each model’s probability against the observed outcome. Brier loss, L = (p − y)², is one way to do this; lower loss means a better forecast score.
This illustrates the incentive behind calibrated probabilities. It does not establish which reward Jev uses.
Here p is a model’s probability that returns is the correct queue. The observed label supplies y = 1 when returns is correct, and y = 0 otherwise. Neither model chooses the observed label.
| Observed outcome | Model A · p = 0.60 | Model B · p = 0.99 |
|---|---|---|
| Returns is correct (y = 1) | (0.60 − 1)² = 0.16 | (0.99 − 1)² = 0.0001 |
| Returns is incorrect (y = 0) | (0.60 − 0)² = 0.36 | (0.99 − 0)² = 0.9801 |
Model B receives a lower loss when correct, but a much higher loss when wrong. One ticket does not tell us which probability was better calibrated. We must consider both possible outcomes and how often each occurs.
Suppose returns is truly correct in 60% of comparable cases, and the models keep reporting these probabilities. Model A’s expected loss is 0.60 × 0.16 + 0.40 × 0.36 = 0.24. Model B’s is 0.60 × 0.0001 + 0.40 × 0.9801 = 0.3921. Model A performs better on average because its probability matches the true rate.
Training could use reward = −loss: −0.24 is higher, and therefore better, than −0.3921. The later derivation shows why matching the true rate minimizes expected Brier loss.
One model predicts an 80% probability of returns for each of ten tickets. In this illustration, returns is the correct queue for eight tickets and incorrect for two.
Calibration is a property of a group, never of one case. It asks whether the agreement persists across similar forecasts: among all the 80 % predictions, does the event occur roughly 80 % of the time?
Each dot represents a different ticket, not a different model. A filled dot means the observed category was returns. An empty dot means the observed category was something else.
The model’s 80% prediction does not identify which tickets will be mistakes. Calibration asks whether the frequency matches the prediction across many comparable cases. Ten tickets make the idea visible, but are not enough to establish calibration reliably.
Y = 1 when returns is the correct queue, Y = 0 otherwise. The outcome, after the fact.
The model’s predicted probability that returns is the correct queue. A random variable — it changes case to case.
One particular forecast value we want to talk about, e.g. the number 0.8.
In wordsLook at every case where the model happened to say p. Among exactly those cases, the event's true rate should be p. Forecasts of 80 % should approach 80 % frequency across comparable cases.
Concrete exampleSuppose the model makes 100 predictions, and for each one it says “there is a 70 % chance this is correct.” If the model is well calibrated, roughly 70 of those 100 predictions should actually be correct. That is Equation 1 with p = 0.7 — and “roughly” matters: sampling uncertainty and dependence between cases affect the strength of the evidence.
ℙ means probability. Y = 1 means the event occurred. The vertical bar | means “given that.” p̂ = p restricts attention to cases where the model reported the value p.
Read the equation as: “Given that the model reported p, the actual chance of the event is p.” For p = 0.8, the target event rate is 80%. This is a population condition, not a requirement that every finite batch match exactly.
For continuously valued forecasts, the precise statement is 𝔼[Y | p̂] = p̂ almost surely. In a finite dataset, we estimate agreement using nearby probabilities, as on the next slide.
Group nearby held-out forecasts into bins. For each nonempty bin, compare the mean prediction with the observed event frequency.
Bucket number k — the set of cases whose forecasts fell in that range.
How many cases landed in the bucket. Just a count.
Sum the values, then divide by the number of cases. A bar over a symbol means “average”.
In wordsWhat the model said, on average, in this bin — against how often the event actually happened in that same bin. Plot one point per bin: said on the horizontal axis, happened on the vertical. Perfect population calibration corresponds to the diagonal; finite samples fluctuate around it.
Concrete exampleA bin holds five forecasts: 0.78, 0.80, 0.82, 0.81, 0.79. Their average — what the model said — is p̄k = 0.80. The outcomes were 1, 1, 0, 1, 0 — the event happened 3 times out of 5, so ȳk = 0.60. The model said 80 % but the event occurred 60 % of the time: plot the point (0.80, 0.60), below the diagonal — overstated in this sample. Five cases alone do not establish miscalibration.
A bin is a group of tickets with similar predicted probabilities, such as 75%–85%. The predictions need not repeat exactly for us to compare their mean with the group’s observed frequency.
Held-out cases were not used to fit the model or tune the calibration procedure. Checking these cases helps assess whether the probabilities remain useful on new data. Bin choices, sample size, and dependence between cases affect the evidence.
Nine equally weighted synthetic forecast groups. Drag right: extreme predictions become overconfident and observed rates move toward 50%.
A point at (0.80, 0.60) means the model predicted 80% on average for that group, while the event occurred in 60% of cases. The point lies below the diagonal because the event probability was overstated in that group.
Moving the slider changes the synthetic observed frequencies, keeping the forecast positions fixed. This lets you inspect different calibration patterns; it is not a simulation of Jev training.
When extreme forecasts are overconfident, high probabilities can lie below the diagonal and low probabilities above it. For example, a 10% event forecast with a 30% event rate is too certain that the event will not occur.
TypeSafe names RLCD's goal — calibrated decisions. Reviewed sources do not disclose its exact reward, loss, or update algorithm.
The scoring-rule examples that follow are a teaching illustration of how a calibration reward could work — standard, well-understood mathematics that makes the incentive visible. It is not a reconstruction of Jev's implementation, and you should not cite it as one.
In wordsContext enters a model, a probability comes out, and the outcome that eventually occurs supplies the feedback signal. This is a conceptual example, not a claim about Jev’s training pipeline.
A loss measures prediction error: smaller is better. A reward scores the model’s behavior: larger is better.
The same preference, with the sign reversedMaximizing expected negative loss is equivalent to minimizing expected loss.
Expected loss: 0.24
Expected reward: −0.24
Expected loss: 0.3921
Expected reward: −0.3921
These reuse slide 4’s example, where the true rate is 60%. Model A has the lower loss and the higher reward: −0.24 > −0.3921.
θ denotes the model’s trainable parameters. 𝔼 means an average over possible outcomes, weighted by their probabilities. This defines an objective, not the algorithm used to update θ.
A model reporting 99% can be correct on one ticket and receive an excellent score. That does not establish that 99% is a realistic probability. Expected reward includes both correct and incorrect outcomes at their underlying rates.
Training estimates the objective from data. Finite data, imperfect labels, model limitations, and optimization can prevent it from reaching the population optimum. A proper scoring rule supplies an incentive; it does not by itself guarantee calibration.
Let p be the model’s probability that returns is correct, and y be the observed result: 1 if returns is correct, 0 otherwise.
In wordsThe model’s reported probability minus the outcome — which is 0 or 1 — squared. Smaller is better. (This is the single binary squared-error convention; some texts double it.)
| Outcome | y | Loss (p − y)² |
|---|---|---|
| Event happens | 1 | (0.8 − 1)² = 0.04 |
| No event | 0 | (0.8 − 0)² = 0.64 |
For p = 0.8, a non-event incurs 16 times the loss of an event. To understand the incentive, average over both outcomes, as the next slide does.
At p = 0.8, an observed event produces (0.8 − 1)² = 0.04. A non-event produces (0.8 − 0)² = 0.64. These are forecast losses, not probabilities and not business costs.
The model reports p before the outcome is known. The outcome y is then used to evaluate that report. The application’s financial cost of acting on the prediction is a separate question, covered later.
One outcome isn't enough to judge a forecast. Let p be the model’s report and q the true event rate for cases like this one. Events occur with probability q, non-events with 1 − q. Average the two losses.
In wordsComplete the square and the expected loss splits into two pieces. The second is irreducible — it is the outcome variance for this fixed q. Changing your report cannot reduce it; better information may change q. The first is the only part you control, and it is a squared distance from the truth.
A squared distance is smallest when it is zero. So expected loss is minimised at p = q.
Truthful reporting is the optimal strategy — not because we asked nicely, but because every alternative is arithmetically worse. The unique optimum makes Brier strictly proper. This is a population result; finite training does not guarantee calibration.
1. Weight both outcomes. The event occurs with probability q, giving loss (1 − p)². The non-event occurs with probability 1 − q, giving loss p². Therefore 𝔼[L] = q(1 − p)² + (1 − q)p².
2. Expand. q(1 − 2p + p²) + p² − qp² = p² − 2pq + q.
3. Complete the square. Add and subtract q²: (p² − 2pq + q²) + q − q² = (p − q)² + q(1 − q).
4. Find the minimum. Holding q fixed, the second term does not depend on the model’s report. The first term is nonnegative and is zero only when p = q.
The horizontal axis is the model’s reported probability p; the vertical axis is its expected loss. Set the assumed true rate q, then move p to compare reports.
Keep q at 0.80. Set the model’s report p to 0.80: expected loss is 0.16, the minimum. Move p to 0.95: loss rises to 0.1825. Move p to 0.50: loss rises to 0.25.
You are controlling the assumed population and the model’s report for illustration. In deployment, the true conditional rate q is generally unknown; we evaluate predictions using observed outcomes.
Take a toy population whose true rate is q = 0.8. The irreducible term is 0.8 × 0.2 = 0.16. Now compare three ways of reporting it.
| Model report | Squared distance (p − q)² | Irreducible q(1−q) | Expected Brier loss | |
|---|---|---|---|---|
| p = 0.80 honest | (0.80 − 0.80)² = 0 | 0.16 | 0.1600 | best |
| p = 0.95 overconfident | (0.15)² = 0.0225 | 0.16 | 0.1825 | +14 % |
| p = 0.50 hedging | (0.30)² = 0.09 | 0.16 | 0.2500 | +56 % |
These are calculations, not measured Jev results. The value of the table is the mechanism it exposes — that a squared-error reward on the probability itself makes honesty the profit-maximising move. The reviewed sources do not specify whether Jev uses Brier, log loss, or another objective.
Each row represents a possible model forecast for the same kind of case. These could be three models, or three candidate reports from one model. The true rate q = 0.8 stays fixed across the comparison.
The expected loss of 0.16 does not mean 16% classification errors. Brier loss measures the quality of probabilities. Even the optimal probability report has a nonzero expected loss when outcomes remain uncertain.
Your application defines the question and what the answers mean. Jev returns one of three shapes.
Selects among supplied alternatives — returns, billing, shipping.
∑c pc = 1
The option with the highest probability is the selected category. Your application separately decides how to act on it.
A binary proposition. “Does the customer request a person?”
yes .50 · no .50 → p = 0.5
0.5 means uncertain, not medium. It is not an intensity dial, and there is no separate confidence field here.
Ordered levels you define.
0 cosmetic · 1 workaround · 2 blocking
The lecture’s example assigns probabilities 0, 0.7, 0.3. Next slide turns that into one number.
Choice: “Which queue should handle this ticket?” The model predicts a category and returns probabilities over the options.
Noul: “Does this customer request a human?” The model returns the probability of yes. A value of 0.5 means uncertainty about that proposition, not a halfway request.
Score: “Which severity level describes this bug?” The model returns a distribution over rubric levels and its mean index. The application can use these predictions in a routing or prioritization rule.
Ordered levels carry an index. Multiply each index by its probability and add them up.
In wordsA probability-weighted average of the level numbers. With most of the mass on “workaround” and 30% on “blocking”, the expectation lands just above 1.
It is not a percentage. It is not a physical measurement. There is no level 1.3 — nothing in your system is “1.3 broken”.
It is an expected index: a weighted average under the model’s predicted distribution, not necessarily the true outcome distribution.
The model assigns 0% to level 0, 70% to level 1, and 30% to level 2. Multiplying each index by its probability gives 0, 0.7, and 0.6. Their sum is 1.3.
That number summarizes the model’s uncertainty over defined levels. It does not mean the bug is “130% severe,” nor does it prove the model’s probabilities are calibrated. Different probability distributions can have the same mean, so inspect the distribution when the distinction affects your action.
The lecture’s Choice example uses this distribution, and alongside it a separate field called confidence.
| returns | 0.60 | ← the winner |
| billing | 0.38 | |
| shipping | 0.02 |
0.39
Not 0.60. Not the winning probability. Not anything you can derive from the winner alone.
The lecture’s Score example reports confidence 0.54 — again unequal to any single probability in it.
In wordsDocumentation describes confidence as summarising the shape of the distribution — how concentrated it is — without specifying the formula. Two non-negotiables follow: don't invent a formula, and don't read it as a probability of correctness.
In this lecture example, 0.60 answers “What probability does the model assign to returns?” The separate 0.39 confidence value summarizes the distribution’s shape.
Do not interpret confidence 0.39 as “the selected category has a 39% chance of being correct.” The model’s probability for that category is 0.60. Whether 0.60 matches actual correctness rates still needs evaluation on outcomes.
A distribution concentrated on the wrong category can have high confidence. Confidence may help route cases, but its usefulness must be checked against the task’s observed results.
Here the predicted event changes: p is the model’s probability that a transaction is fraudulent. The application chooses approval or review. Under this illustrative cost model: approving fraud costs $100, approving a legitimate transaction costs nothing, and review costs $5 and prevents all fraud loss.
In wordsExpected cost is each outcome's cost weighted by its probability. Compare the two numbers and take the smaller one. The whole decision collapses to a single threshold — p* = Creview ⁄ Cfraud, here 5⁄100.
100 × 0.02 = $2 vs $5
Application: approve. Its expected cost is lower under these assumptions.
100 × 0.05 = $5 vs $5
A tie. Exactly the threshold; the cost rule is indifferent.
100 × 0.12 = $12 vs $5
Application: review. This saves $7 in expectation under the stated assumptions.
If the application approves, it loses $100 when fraud occurs and $0 otherwise. With fraud probability p, expected approval cost is p × $100 + (1 − p) × $0 = $100p.
At p = 0.02, the model estimates a 2% chance of fraud. Approval costs $2 on average across comparable cases, while review costs $5. An individual transaction does not incur exactly $2: the average combines possible $100 and $0 outcomes.
At p = 0.12, expected approval cost is $12. Under the assumption that a $5 review prevents all fraud loss, review is cheaper. Both transactions are more likely legitimate than fraudulent, yet the application’s preferred actions differ.
Setting $100p = $5 gives p = 0.05. Below that value, approval is cheaper; above it, review is cheaper; at it, the expected costs tie.
Two expected-cost lines. Where they cross is your automation threshold. Change the costs and watch it move.
Probabilities describe uncertainty. Costs determine action. Keep them in separate parts of your code.
The model estimates fraud probability. The application supplies the costs and chooses the lower-cost action. Changing review costs can change the action without changing the prediction.
With missed fraud costing $100 and review costing $5, the threshold is 5%. Raise the review cost to $10: the threshold becomes 10%, so review is justified only at higher fraud probabilities.
Keep review at $5 but raise missed-fraud cost to $200: the threshold falls to 2.5%. More expensive mistakes justify reviewing lower-risk transactions.
The sliders change the application’s cost assumptions, not the model’s probability. If review costs more than the maximum possible fraud loss, approval is cheaper for every probability in this simplified model.
Suppose half your cases are positive. A model that says 50 % to everything is perfectly calibrated — among its 50 % forecasts, the event does occur half the time. It provides a useful base-rate forecast, but cannot distinguish one case from another.
always p = 0.5
Calibrated ✓ · Expected Brier 0.25
p = 0.2 and p = 0.8
Calibrated ✓ · Expected Brier 0.16
In wordsBrier splits three ways. Reliability is miscalibration — both models score 0, both are honest. Uncertainty is the base rate's own noise — identical for both. The difference is entirely resolution: how much the conditional event rates differ across forecast groups, while staying honest.
Population definitionLet m(p̂) = 𝔼[Y | p̂] and μ = 𝔼[Y]. Reliability = 𝔼[(p̂ − m(p̂))²]; resolution = 𝔼[(m(p̂) − μ)²]; uncertainty = μ(1 − μ). Arbitrary finite bins only approximate these population terms.
So the constant baseline was never miscalibrated — it simply offered no discrimination between cases. Calibration is necessary and not sufficient; you want a model that is honest and discriminating. (This three-way split is standard forecast-verification mathematics, included to make the lecture's point precise — it is not part of any stated Jev method.)
Imagine 100 cases. Model A reports 50% for every case, and 50 cases are positive. Model A is calibrated at its single forecast value, but it cannot distinguish higher-risk from lower-risk cases.
Model B reports 20% for 50 cases and 80% for the other 50. In this illustrative sample, the first group has 10 positive cases and the second has 40. Model B is calibrated in both groups, and there are still 50 positives overall.
Model B gives the application more useful distinctions. Its Brier score is 0.16, versus Model A’s 0.25. That improvement comes from separating groups whose actual outcome rates differ, not merely from making predictions more extreme.
Type safety prevents invalid schema answers — you will never get shippng or a stray paragraph where an enum belongs.
It does not prevent wrong valid choices. Ambiguous inputs and missing information remain exactly as hard as they were.
Evaluate on held-out cases and subgroups — a model can be calibrated overall and badly miscalibrated for one customer segment. Log outcomes, check sample sizes, monitor distribution shift.
unstructured input + a defined question
category + probabilities
compare expected costs; act or ask
what actually happened
Record the predicted event, its probability, the application’s action, and the eventual observed outcome. Then compare forecasts with event frequencies on new cases and on relevant subgroups.
Keep the event definition consistent. If p predicts fraud, the evaluation outcome must indicate fraud; it should not silently become “a loss happened after our intervention.” Review may prevent losses, and outcomes may be missing for some actions. A biased set of observed labels can give a misleading calibration picture.
A valid output such as returns still can be the wrong category. Type safety checks the output’s allowed form; outcome evaluation checks its accuracy and calibration.
| # | Equation | Say it in one sentence | Status |
|---|---|---|---|
| 1 | ℙ(Y=1 | p̂=p) = p | Among cases where you said p, the event happens p of the time. | standard |
| 2 | p̄k vs ȳk | Bin held-out forecasts and plot promised against observed. | standard |
| 3 | L = (p − y)² | Brier: square the gap between claim and outcome. | illustrative |
| 4 | 𝔼[L] = (p−q)² + q(1−q) | Honesty is optimal — the only term you control is a squared distance from the truth. | illustrative |
| 5 | 𝔼[level] = ∑ i·pi = 1.3 | An expected index over ordered levels — not a measurement. | documented |
| 6 | confidence ≠ probability (in general) | A distribution summary; validate any routing rule on outcome data. | undisclosed |
| 7 | Capprove=100p vs Creview=5 | Act when expected cost is lower; the threshold is a cost ratio. | illustrative |
| 8 | reliability − resolution + uncertainty | Useful forecasts distinguish groups with different outcome rates. | illustrative |
Ask what the probabilities mean, test them against reality, and connect them to the consequences of acting.
Sources: TypeSafe documentation as presented in the Jev RLCD lecture, plus standard statistical references for the Brier score and forecast verification. The sources reviewed here do not specify RLCD’s reward, loss, or update algorithm. See the source notes on the next slide.
Prediction: the model supplies an event probability. Calibration: we ask whether those probabilities agree with observed frequencies. Training: a proper score illustrates why reporting accurate probabilities can earn a higher expected reward.
Decision: the application combines probabilities with action costs. Evaluation: we verify both calibration and the ability to distinguish cases, then monitor outcomes over time.
RLCD names Jev’s calibration-oriented training approach. The mathematical examples explain the goal and possible incentives; they do not identify its implementation.
Video reviewed: What is RLCD in Jev AI? · 9:59. Calibration 1:07–2:27; scoring 2:33–5:19; output types 5:24–7:12; decision costs 7:17–8:19.
The 0.60 / 0.38 / 0.02 distribution, confidence 0.39, and score 1.3 are lecture examples; current documentation uses different numerical examples. Brier loss and the cost model are teaching illustrations, not Jev benchmarks or a reconstruction of its training algorithm.