From a typed judgment to a justified action: probability, confidence, calibration, Bayesian estimation, and reinforcement learning.
Source review: September 19, 2026 · Jev 1.13 documentation · examples are explicitly labeled.
A customer writes: “My running shoes are the wrong size, and my card was charged twice. Please swap them and refund the extra charge.” You are building the first routing step in a support system. Should the ticket go to Returns, Billing, or Shipping?
The ticket is our synthetic A-104 case file. Its local policy says Billing owns a duplicate-charge review even when the same message also requests an exchange. Billing is therefore the human-specified primary route for this exercise. That policy is part of our example, not a claim about Jev or any real support organization.
A free-text model might reply “billing team, please” or “send it to the people who handle charges.” A developer then has to parse those strings before code can route the ticket. Jev is documented as returning typed values from caller-defined answer spaces rather than generated prose; that solves the answer-shape problem. It does not make a chosen option correct by itself. TypeSafe Introduction — typed questions and outputs
Try the routing buttons below. “Returns” is a permitted label, so the schema gate passes. But our stated policy says Billing is the primary owner, so the correctness gate fails. The two gates answer different questions: can software read this value? and does it match the case's labeled target?
This is a local, deterministic teaching fixture. Change the response and watch the gates; no Jev API is called.
A schema specifies which forms of answer the program accepts. Here it is a simple set of three strings. A ground-truth label is the answer our policy or a resolved case says should have been chosen. Code can check schema membership without knowing the real-world answer; the second check needs a trustworthy label or review.
Count three illustrative responses to the same synthetic ticket: Billing, Billing, and Returns. Every string is in the allowed set, so schema validity is 3 ÷ 3 = 1.00, or 100%. Two routes agree with the policy truth, so case accuracy is 2 ÷ 3 ≈ 0.667, or 66.7%. The same three responses support both numbers because the tests measure different properties. These are arithmetic on a teaching fixture, not measured Jev rates.
Change the three responses yourself. The denominators stay at three, while the two numerators update separately. Selecting an unlisted label fails both checks; selecting the wrong allowed team passes only the first. This small audit is useful because a schema-only test suite could report a clean run while overlooking a routing regression.
These are repeated illustrative responses to the same A-104 case, not observed model trials.
| Response | Schema | Policy match |
|---|---|---|
| Billing | pass | match |
| Billing | pass | match |
| Returns | pass | mismatch |
Schema: 3 of 3. Policy match: 2 of 3.
One more option makes the difference vivid. “Escalations” is not in our three-option schema, so it fails the first gate even if a human might decide to escalate later. A valid value can be wrong; an invalid value cannot be safely consumed as this field. A production workflow may have a separate explicit escalation action in code.
Implement the two checks separately. The code below uses no model and no third-party package: it simply verifies the response contract and compares a valid response with the synthetic policy label. Run it from top to bottom to see all three cases and the two rates.
# Pure Python; no Jev API call.
ALLOWED = {"returns", "billing", "shipping"}
TRUTH = "billing" # Synthetic A-104 policy label.
def schema_valid(route):
return isinstance(route, str) and route in ALLOWED
def policy_correct(route):
return schema_valid(route) and route == TRUTH
responses = ["billing", "billing", "returns"]
valid_count = sum(schema_valid(route) for route in responses)
correct_count = sum(policy_correct(route) for route in responses)
validity = valid_count / len(responses)
accuracy = correct_count / len(responses)
print(valid_count, correct_count, validity, round(accuracy, 3))
assert (valid_count, correct_count) == (3, 2)
assert schema_valid("returns") and not policy_correct("returns")
assert not schema_valid("escalations")
Python's standard-library json.loads can parse an incoming JSON string, but
parsing alone answers only whether the text is valid JSON. The equivalent membership check
still needs route in ALLOWED, and semantic correctness still needs a policy
label. For example, json.loads('{"route":"returns"}')["route"] yields a readable
yet wrong route for A-104.
We now know why a bounded answer is helpful and why it is not enough. Next we will open the actual request contract: what state does Jev see, what questions can we ask, and what kinds of values come back?
Suppose the support program has only A-104's message. It needs a route, a rough urgency level, and a yes/no judgment about the duplicate charge. Those are different questions about the same evidence. A single unstructured “analyze this ticket” answer would hide which value belongs to which task.
TypeSafe calls the evidence state. A Jev request sends state alongside typed questions. The state can be a text string or a JSON object/array containing text values; Jev 1.13 does not directly inspect an image, audio clip, or video. We use a JSON-like case file so each question can point to the message, order, or policy by name. TypeSafe State documentation
message: wrong-size shoes, duplicate charge, exchange and extra-charge refund requested.policy: Billing owns duplicate-charge review.order: one order with two captured card charges in this teaching example.
A Choice question asks which of a fixed list fits. For the primary
department, the criteria are Returns, Billing, and Shipping. Its documented response has
choice, one probability per option, and a separate confidence field.
In the local A-104 fixture below, Returns has .60, Billing .38, and Shipping .02; these are
invented demonstration values, not a Jev call or a measured Jev prediction.
TypeSafe
Choice documentation
A Score question uses ordered levels defined by the caller, such as 0 = routine, 1 = soon, and 2 = urgent. The response includes a distribution over those levels and a numeric position along them. The local fixture assigns probabilities 0, .70, .30. By hand, 0 × 0 + 1 × .70 + 2 × .30 = 0 + .70 + .60 = 1.30. That is a place on our urgency rubric, not a 130% probability or an exact clock deadline. TypeSafe Score documentation
A Noul asks one yes/no question and returns the modeled probability of “yes”
between 0 and 1. Our fixture uses .80 for “Does the message report a duplicate charge?” A .50
Noul would mean yes and no have equal modeled probability; it would not mean medium urgency.
The documented Noul response has no separate confidence field.
TypeSafe
Noul documentation
The A-104 state remains the same. Select which question type the program asks. Bars and values are deterministic local fixtures, not live Jev output.
Move the urgent-mass slider while Score is selected. Its mass comes from the “soon” level, so the two nonzero masses still sum to one. The weighted position shifts continuously between level one and level two; this is the meaning of a score between labels. Move the Noul slider and the yes/no line moves independently. Neither slider is changing Jev's weights or running its inference service.
For Choice, the three fixture probabilities must sum to one: .60 + .38 + .02 = 1.00. The selected route is the largest probability, Returns. But the human-authored A-104 primary label is Billing, so this particular fixture is an example of a plausible, typed, wrong answer. The separate confidence formula is not disclosed in the cited API pages; do not infer it from our bars.
Try changing the primitive while keeping the ticket fixed. The question type controls what the result means and which fields your code may read. Choice is appropriate for mutually exclusive routes; Score is for a defined ordered rubric; Noul is for one statement that might be true. A production system can ask all three in one request, but they are not interchangeable gauges.
confidence field to Noul: the current primitive
documentation does not return one.
TypeSafe
primitive response tableThe following plain Python reproduces the fixture arithmetic and fields without implementing or pretending to run Jev. The keys illustrate some documented fields, not the full SDK serialization; the numbers are supplied by this lesson. Running it prints the same choice, score, and Noul values shown above.
# Local fixture only. No learned model and no network request.
import math
CHOICE_CRITERIA = {"returns", "billing", "shipping"}
route_probabilities = {
"returns": 0.60,
"billing": 0.38,
"shipping": 0.02,
}
level_probabilities = [0.0, 0.70, 0.30]
duplicate_yes_probability = 0.80
def make_choice(probabilities):
assert 1 <= len(probabilities) <= 255
assert set(probabilities) == CHOICE_CRITERIA
assert all(isinstance(key, str) for key in probabilities)
assert all(math.isfinite(p) and p >= 0 for p in probabilities.values())
assert abs(sum(probabilities.values()) - 1.0) < 1e-9
selected = max(probabilities, key=probabilities.get)
return {"choice": selected, "probabilities": probabilities}
def make_score(probabilities):
assert len(probabilities) == 3 # This fixture has three named levels.
assert all(math.isfinite(p) and p >= 0 for p in probabilities)
assert abs(sum(probabilities) - 1.0) < 1e-9
score = sum(level * p for level, p in enumerate(probabilities))
return {"score": score, "probabilities": probabilities}
def make_noul(yes_probability):
assert math.isfinite(yes_probability)
assert 0.0 <= yes_probability <= 1.0
return {"noul": yes_probability}
choice = make_choice(route_probabilities)
score = make_score(level_probabilities)
noul = make_noul(duplicate_yes_probability)
print(choice["choice"], score["score"], noul["noul"])
assert choice["choice"] == "returns"
assert abs(score["score"] - 1.30) < 1e-9
assert noul == {"noul": 0.80}
The short equivalent for the Score arithmetic is Python's built-in sum(i*p for i,p in
enumerate(level_probabilities)); it gives the same 1.30. The actual TypeSafe SDK offers
Choice, Score, and Noul question constructors, but only
the service supplies real model probabilities and the separate Choice/Score confidence field.
Our plain dictionaries intentionally leave that undisclosed confidence value out.
TypeSafe
SDK-shaped question examples
If you later replace this fixture with Jev, define the actual questions with the SDK rather
than asking the service for the displayed numbers. The following request shape follows
TypeSafe's Python SDK examples, including descriptions for each Choice criterion and a pinned
model version. It is an optional integration illustration: it requires
typesafe_sdk, an API key, and a network request; we did not execute it or claim its
response would match the fixture.
# Optional real SDK request shape; not run by this lesson.
from typesafe_sdk import Choice, Score, Noul, TypeSafeClient
state = {
"message": "Wrong-size shoes and two charges; swap and refund extra charge",
"policy": "Billing owns duplicate-charge review",
}
questions = {
"route": Choice(
instructions="Which team owns the primary review under policy?",
criteria={
"returns": "Exchange and wrong-size requests",
"billing": "Duplicate charges and payment review",
"shipping": "Delivery and missing-package issues",
},
),
"urgency": Score(
instructions="How urgent is this support request?",
criteria=["Routine", "Soon", "Urgent"],
),
"duplicate": Noul(
instructions="Does the message report a duplicate card charge?",
),
}
# Only run with your own configured credentials and network access:
# with TypeSafeClient(model="jev-1.13.0") as client:
# response = client.system_one(state=state, questions=questions)
The versioned model ID is documented as available at the time of this lesson; aliases such as
jev-latest may later point to a different version. This request defines
questions. It does not manufacture the answer probabilities. Our runnable
standard-library code above remains the reproducible comparison.
TypeSafe Python
multi-question example;
TypeSafe model IDs
and aliases
We have three useful judgments about one state. The next problem is coordination: if all three questions use A-104, should we ask them together, and does “independent evaluation” mean the facts themselves are statistically independent?
A-104 contains an exchange request and a duplicate charge. If we ask “What should happen with this customer?” we are hiding several decisions inside one sentence. A route, an urgency rating, and the existence of a duplicate-charge claim each need their own answer contract.
TypeSafe's interface lets a request carry multiple typed questions about one state. Each question sees the same state and is evaluated independently of the other answers. The company says those questions are evaluated in parallel, so one request can collect several narrow judgments. This is about the model's question evaluation, not a claim that facts in the ticket are statistically independent. TypeSafe Primitives — asking multiple questions
Click the lanes below. In the ordinary path, route, urgency, and duplicate-charge questions all see the same initial A-104 state; turning one off leaves the other local fixture values unchanged. Stepping to “code combines” shows our invented rule: if the duplicate-charge probability is at least .75, flag Billing review; otherwise use the route choice. The dependency switch illustrates a different path where urgency cannot be constructed yet.
Toggle questions or a dependency, then advance the two-stage flow. One lane per row keeps the same structure readable on a narrow screen.
An unfetched record alone does not force a second model call: code could fetch a known record before asking any questions. For a genuine dependency here, the first-stage Noul answer about the duplicate charge determines which record code may fetch and which urgency question it can construct. If the answer crosses our illustrative threshold, code fetches a payment record; otherwise it fetches order-fulfillment history. Only then can the application form the enriched state and ask urgency in a second request. Route and duplicate-charge questions in the first batch still see the same original state. TypeSafe Primitives — when one question depends on another
“Evaluated independently” is easy to mishear as “the events are independent.” To see the difference, use the local hypothetical marginal probabilities P(duplicate charge) = .80 and P(urgent) = .50. We have not been given a joint probability that both are true. The overlap cannot exceed the smaller marginal, so its upper bound is min(.80,.50) = .50.
At the low end, the two events together occupy .80 + .50 = 1.30 of a unit population if counted separately. Their overlap must be at least the excess over one: max(0, .80 + .50 − 1) = .30. Thus P(both) can lie from .30 to .50. Multiplying .80 × .50 = .40 selects one possible value only under the extra assumption that the events are statistically independent. Parallel computation did not grant us that assumption.
The bounds are real alternatives, not just algebra. Move the overlap below while both marginal probabilities stay fixed. When the overlap grows, the “neither” group grows too, and each one-event-only group shrinks. The four groups always add to one. This is a local probability table, not an output that Jev supplies for these two separate questions.
Drag the possible overlap from .30 to .50. Every cell is recalculated from P(duplicate)=.80 and P(urgent)=.50.
Here is a runnable local approximation of the software workflow, not a Jev implementation. The fixture lookup stands in for model results, and the branch belongs to code. Its assertion shows that adding urgency to the batch leaves the route and duplicate-charge fixture values unchanged.
# Deterministic fixture; no Jev API call.
STATE = {"ticket_id": "A-104", "text": "Wrong size and charged twice"}
ANSWERS = {
"route": {"choice": "returns",
"probabilities": {"returns": .60, "billing": .38, "shipping": .02}},
"urgency": {"score": 1.30},
"duplicate": {"noul": .80},
}
def ask_fixture(state, question_ids):
assert state["ticket_id"] == "A-104"
result = {}
for question_id in question_ids:
if question_id not in ANSWERS:
raise KeyError(question_id)
result[question_id] = ANSWERS[question_id]
return result
def choose_workflow_step(answers):
if answers.get("duplicate", {}).get("noul", 0.0) >= .75:
return "billing_review" # Local policy, not a model action.
if "route" in answers:
return answers["route"]["choice"]
return "request_more_evidence"
pair = ask_fixture(STATE, ["route", "duplicate"])
triple = ask_fixture(STATE, ["route", "urgency", "duplicate"])
assert pair["route"] == triple["route"]
assert pair["duplicate"] == triple["duplicate"]
assert choose_workflow_step(triple) == "billing_review"
print(choose_workflow_step(triple))
The short Python equivalent for gathering known answer IDs is {key: ANSWERS[key] for
key in question_ids}; it returns the same mapping. The actual SDK sends typed questions
with one state, rather than this lookup table. We keep the lookup local so the teaching flow
is reproducible without a credential, billing, or an invented claim about Jev's particular
answer to A-104.
TypeSafe multi-question request example
Our program now has a route probability distribution and a branch, but one more number appears in Choice and Score responses: confidence. The next chapter separates an option probability, the selected option, and that distribution-derived summary.
Our A-104 fixture gives Returns .60, Billing .38, and Shipping .02. Why does the program pick Returns? Because .60 is the largest of the three modeled option probabilities. But the customer also has a duplicate-charge issue, so the second option is not negligible. We need to see the whole distribution, not just the selected label.
For a TypeSafe Choice answer, probabilities lists masses over the caller's
options, and choice is the highest-probability one. The probabilities sum to one
across that supplied list. A separate confidence number summarizes how
concentrated the distribution is; the cited documentation does not provide a formula we should
reproduce. It also does not make a high value a guarantee that the chosen route is right.
TypeSafe
Choice — response structure;
TypeSafe
Confidence
Drag a bar or use the sliders. The display starts at .60, .38, .02 and always renormalizes to a total of one. Moving one option changes the other two proportionally, a rule of this teaching widget only. The adjacent readout names the top option and calculates our custom margin; it never presents a computed Jev confidence.
Local synthetic A-104 fixture. Drag a bar horizontally, adjust the accessible sliders, or select a preset. The human-authored primary route label remains Billing.
At the starting distribution, the top-two gap is .60 − .38 = .22. The selected option's own mass remains .60. At the peaked preset, the gap becomes .90 − .08 = .82. A larger gap can be useful as a custom warning signal, but it is not TypeSafe's reported confidence and cannot by itself tell us the route is correct under A-104's policy.
Try “Keep top .60, spread rest.” It sets Returns to .60, Billing to .20, and Shipping to .20. The selected option and its top probability have not changed, yet the runner-up shrinks and our margin becomes .60 − .20 = .40. This is a concrete reason to keep the full distribution: a single top number hides where the remaining mass went. It still cannot reveal TypeSafe's undisclosed confidence calculation.
The difference matters in practice. If code mistakes a confidence scalar for the probability of a particular event, it may compute the wrong expected cost. TypeSafe documents confidence as a summary derived from the distribution shape, and says thresholds should be adjusted for the risk and the task. A model can express a concentrated distribution and still choose the wrong allowed answer. TypeSafe Confidence — thresholds scale with risk
The documentation includes a Choice example with top route probability .61 and returned confidence .42. Those two values differ because they answer different questions; the .42 is a documented example output, not a number calculated by this widget or a result from A-104. We deliberately keep that evidence apart from the local editable bars. TypeSafe Choice — multi-issue ticket example
confidence of .42. The API page does not publish a formula that maps
an arbitrary bar set to that number. Do not apply .42 to the bars after you drag them.1 − confidence is not a documented
error probability. Nor is our top-two margin the vendor confidence field. For a decision with
real consequences, test probability calibration and the action's loss on labeled cases rather
than treating any one summary as a safety certificate.Run this plain Python example to see how a selected option and a custom top-two gap are
computed. The input weights are supplied locally. The function makes no Jev call, and the
result has no fabricated Jev confidence field.
# Local distribution arithmetic only; not Jev's confidence formula.
OPTIONS = ["returns", "billing", "shipping"]
def normalize(weights):
assert len(weights) == len(OPTIONS)
assert all(weight >= 0 for weight in weights)
total = sum(weights)
if total <= 0:
raise ValueError("At least one weight must be positive")
return [weight / total for weight in weights]
def selected_and_margin(weights):
probabilities = normalize(weights)
ranked = sorted(range(len(probabilities)),
key=lambda i: probabilities[i], reverse=True)
first, second = ranked[:2]
margin = probabilities[first] - probabilities[second]
return OPTIONS[first], probabilities, margin
name, probabilities, margin = selected_and_margin([60, 38, 2])
print(name, probabilities, round(margin, 2))
assert name == "returns"
assert abs(sum(probabilities) - 1.0) < 1e-12
assert abs(margin - .22) < 1e-12
peaked = selected_and_margin([90, 8, 2])
split = selected_and_margin([60, 20, 20])
assert abs(peaked[2] - .82) < 1e-12
assert split[0] == name and abs(split[2] - .40) < 1e-12
A short standard-library equivalent for finding the selected option is max(zip(OPTIONS,
probabilities), key=lambda pair: pair[1]); it returns the same Returns/.60 pair for the
initial fixture. The longer version shows the extra step needed for the runner-up and margin.
Neither computes TypeSafe's undisclosed confidence statistic.
Chapter 4 asks the empirical question we cannot answer from one ticket: when many cases receive a probability near .80, how often does the event actually occur? That is how we begin testing calibration instead of guessing correctness from the shape of one distribution.
Suppose a future routing ticket receives a forecast of 0.8 for the binary event “its selected route is policy-correct.” That number is a forecast. We cannot mark it calibrated from one ticket, even after the ticket is resolved: the selected route is either correct or not. Calibration is a pattern across comparable forecasts and independently resolved outcomes.
For a deliberately synthetic audit of that same event, collect 100 tickets, each forecast at p = 0.8 that its selected route will be policy-correct. If 80 of those events occur, the observed fraction in that bucket is 80/100 = 0.8. If only 60 occur, the fraction is 0.6 and the bucket misses the diagonal by 0.2. This exercise does not measure Jev; it shows how to audit any probability-bearing workflow.
A-104 remains the motivating case: its illustrative three-route fixture gives Returns .60, Billing .38, Shipping .02, and its policy label is Billing. The 100-ticket binary audit below is a separate constructed sample of whether selected routes are policy-correct; it does not silently replace A-104’s .60/.38/.02 masses. To audit such an event, compare each selected-route correctness forecast with an independently resolved policy label. Do not use the model’s own answer as the ground truth.
The Brier score for a binary event is the average of (forecast − outcome)². With 80 positives, each positive contributes (0.8−1)² = 0.04 and each negative contributes (0.8−0)² = 0.64. The full mean is (80×0.04 + 20×0.64)/100 = 0.16. At 60 positives and 40 negatives it becomes (60×0.04 + 40×0.64)/100 = 0.28. Smaller is better for this score on the same labeled cases. Brier's original probability-verification paper
Switch on the subgroup. It isolates ten illustrative tickets with 4 positives, or 0.4 observed frequency, while the overall 80/100 group still looks aligned. At the 80-positive slider setting, the other 90 contain 76 positives; at 60, they contain 56. Thus the remaining-group count is always the slider total minus 4. A single aggregate reliability point can hide variation that matters to a deployment population. Our subgroup is hand-constructed, not discovered in customer data.
Keep the same slider total. This table partitions the 100 tickets into a constructed ten-ticket subgroup and the remaining ninety. Each score is recomputed from its own positive and negative counts; the weighted average must recover the whole-group Brier score.
| Part | Positive | Negative | Frequency | Brier |
|---|---|---|---|---|
| Small group | ||||
| Other 90 | ||||
| All 100 |
At 80 positives overall, the small group has 4/10 correct routes at forecast .8, while the other group has 76/90. Its Brier score is (4×.04 + 6×.64)/10 = .40; the other group's score is (76×.04 + 14×.64)/90 ≈ .1333. Weight by group size: .10×.40 + .90×.1333 ≈ .16. The aggregate number can conceal a weak slice, even though its arithmetic is valid.
Even 80/100 is a noisy finite sample, not proof that all 0.8 forecasts are right 80% of the time. An audit needs enough independent resolved examples, several probability buckets, relevant segments and later time periods. It also needs a stable label definition: calibration to a weak proxy label cannot by itself establish that the route helps the customer.
Calibration and discrimination answer different questions. A system that always predicts a base rate can sometimes be calibrated, yet be poor at separating hard cases from easy ones. Brier loss provides one useful summary, while the reliability plot exposes where predicted and observed frequencies differ. For action decisions, we must also specify what different mistakes cost.
confidence scalar mean the chance that an action succeeds. TypeSafe confidence
documentationRun this complete Python example. The assertion checks the hand arithmetic; NumPy is an optional independent comparison on the identical fixture, not a second dataset or a Jev measurement.
# Standard-library calculation on synthetic resolved cases.
def brier(probabilities, outcomes):
assert len(probabilities) == len(outcomes) and outcomes
assert all(0 <= p <= 1 for p in probabilities)
assert all(y in (0, 1) for y in outcomes)
return sum((p-y)**2 for p, y in zip(probabilities, outcomes))/len(outcomes)
def frequency(outcomes):
return sum(outcomes)/len(outcomes)
probabilities = [0.8]*100
aligned = [1]*80 + [0]*20
shifted = [1]*60 + [0]*40
assert abs(frequency(aligned)-0.8) < 1e-12
assert abs(brier(probabilities, aligned)-0.16) < 1e-12
assert abs(brier(probabilities, shifted)-0.28) < 1e-12
print(frequency(aligned), brier(probabilities, aligned))
print(frequency(shifted), brier(probabilities, shifted))
small = [1]*4 + [0]*6
remaining = [1]*76 + [0]*14
assert len(small) + len(remaining) == 100
assert sum(small) + sum(remaining) == 80
small_score = brier([0.8]*10, small)
remaining_score = brier([0.8]*90, remaining)
assert abs(small_score-0.40) < 1e-12
assert abs(0.1*small_score + 0.9*remaining_score - 0.16) < 1e-12
print("partition", small_score, remaining_score)
try:
import numpy as np
except ImportError:
print("NumPy comparison skipped")
else:
print(float(np.mean((np.array(probabilities)-np.array(shifted))**2)))
In this fixture, the labels are supplied by the author. In practice, label collection, policy version, unresolved tickets, and changing traffic all affect what “observed frequency” means. Keep the forecast and outcome timestamps, then evaluate only cases whose truth has been established without copying the prediction.
A-104 requests both a shoe exchange and a duplicate-charge refund. In our synthetic policy, Billing is the primary owner because duplicate-charge review takes priority. A bounded route prediction can help the application decide, but the prediction is not itself the routing policy. The developer still chooses which errors matter and what to do when the evidence is weak.
Our local teaching fixture assigns Returns 0.60, Billing 0.38, and Shipping 0.02. Those are invented route masses for A-104; they are not a Jev call or a measured customer distribution. The allowed actions are route to one of those teams or send the ticket for review. A wrong automatic route costs 10 units, a correct route costs zero, and perfect review costs 2 units. These costs encode a preference, not a measured dollar amount.
Expected loss means averaging over possible outcomes using their probabilities. If we automatically choose Returns, its correct mass is .60 and wrong mass is .38+.02=.40. Thus loss = .60×0 + .40×10 = 4. For Billing, loss = (.60+.02)×10 = 6.2. For Shipping it is .98×10 = 9.8. Review's stipulated cost is 2. The minimum expected loss is review, despite Returns having the highest route mass.
Every bar is computed from the same three route masses. This is application-owned policy, not Jev output.
For a top route with probability p, the symmetric wrong-route rule gives automatic loss (1−p)C, where C is wrong-route cost. Perfect review wins when its cost R is lower. The break-even value is p = 1−R/C. With C=10 and R=2, that is 0.8: at p=.60 review wins; at p=.90, automatic loss 1 wins. The equality case is a tie. This threshold does not transfer to a different loss table or to a review process that sometimes makes errors.
The label “Billing” in our case file matters when we evaluate the eventual result. The probability-based policy can rationally request review because its own route distribution favors Returns while the author-supplied primary route truth is Billing. Reviewing the case in our toy reveals that truth. In a deployed system, a human queue might be slow, imperfect, or itself costly in ways our slider omits.
The action set also matters. With no review option, the minimum expected loss is the highest-mass route under equal wrong-route costs. With review available, it can be optimal to abstain from automatic routing. This is selective prediction under a stipulated cost model. It does not mean TypeSafe's scalar confidence has an inherent act/no-act cutoff: the cutoff belongs to the application and should be tested against labeled outcomes. TypeSafe on domain-dependent thresholds
The first comparison assumed review finds the correct team every time. That is a simplifying condition, not a property of real reviewers. Stress-test it with a second toy parameter: let a reviewer choose the wrong team with probability q, independently of the ticket type, while the base review cost still applies. Under the same symmetric wrong-route loss C, review's expected loss becomes R + qC. This extension remains application policy, not a Jev field.
The upper chart keeps the perfect-review baseline. This sensitivity chart shares its wrong-route and base-review cost sliders, then adds a hypothetical reviewer-error rate. It compares review against the best automatic route.
At our baseline C=10 and R=2, the best automatic route is Returns with expected loss 4. Review with error rate q costs 2+10q. It is cheaper while q<.20, ties at q=.20, and becomes more costly beyond that point. The action can change without the route distribution changing at all. In practice, reviewer errors may depend on case type, and delay can matter separately; estimate those effects rather than trusting this one-parameter sensitivity.
Run the entire loss table below. The NumPy comparison performs the same dot products with the same probabilities and cost rows. It is optional; the standard-library result and assertions work without dependencies.
# Synthetic A-104 fixture; no Jev API request.
ROUTES = ("returns", "billing", "shipping")
prob = (0.60, 0.38, 0.02)
assert abs(sum(prob)-1.0) < 1e-12
wrong_cost, review_cost = 10.0, 2.0
def route_loss(route, wrong_cost):
assert route in ROUTES
return sum(p * (0 if truth == route else wrong_cost)
for truth, p in zip(ROUTES, prob))
losses = {route: route_loss(route, wrong_cost) for route in ROUTES}
losses["review"] = review_cost # Perfect review by assumption.
choice = min(losses, key=losses.get)
print(losses, choice)
assert choice == "review"
assert abs(losses["returns"]-4.0) < 1e-12
assert abs(losses["billing"]-6.2) < 1e-12
assert abs(losses["shipping"]-9.8) < 1e-12
assert abs((1-review_cost/wrong_cost)-0.8) < 1e-12
def imperfect_review_loss(error_probability):
assert 0 <= error_probability <= 1
return review_cost + error_probability*wrong_cost
assert imperfect_review_loss(0.0) == 2.0
assert imperfect_review_loss(0.2) == losses["returns"]
assert imperfect_review_loss(0.3) > losses["returns"]
try:
import numpy as np
except ImportError:
print("NumPy comparison skipped")
else:
cost = np.array([[0,10,10],[10,0,10],[10,10,0]], dtype=float)
print("NumPy route losses:", np.array(prob) @ cost.T)
This table assumes one true primary route and the same penalty for every wrong team. A real support system may need a full asymmetric loss matrix: missing a charge dispute might cost more than sending an exchange to Billing first. The method stays the same—multiply each possible outcome's loss by its probability, sum, and compare available actions—but the simple top-probability threshold disappears.
A-104's route distribution, local policy label, and costs are deliberately distinct objects. You can version each one separately. If the fee for review rises while the case evidence stays fixed, code may choose Returns; the model did not “change its mind.” This separation makes a workflow debuggable and sets up the next question: what would a new record do to the probability before we act?
A-104 mentions a possible duplicate charge. Suppose a payment record arrives before we route it. That record can change how plausible Billing is, but only if we know how often such a record appears in genuinely Billing-owned cases and in other cases. This chapter builds a tiny Bayesian application-side teaching model; it does not claim Jev internally runs Bayes' rule.
Start with a different illustrative scenario from chapter 5's route fixture. Before seeing the record, assign probability 0.30 to “Billing is the primary route” and 0.70 to “another team is primary.” This is a synthetic prior over two hypotheses, not Jev's Choice output and not the .38 Billing mass in our earlier three-route example. The values differ on purpose so we can see what evidence adds.
Assume the record is positive 80% of the time when Billing truly owns the case, but 10% of the time when another team owns it. These are hypothetical likelihoods from a stipulated observation model, not TypeSafe benchmark statistics. The positive record has likelihood ratio 0.8/0.1 = 8 in favor of Billing.
Step through prior → joint weights → normalized posterior. The duplicate button copies the same record, so it must not create another likelihood update.
Multiply each prior mass by the probability of seeing this positive record under that hypothesis. Billing gets 0.30×0.80 = 0.24; other gets 0.70×0.10 = 0.07. These are joint weights, not posterior probabilities yet: they sum to 0.31, the probability of seeing a positive record in this toy world.
Normalize both weights by 0.31. Billing becomes 0.24/0.31 ≈ 0.774, and other becomes 0.07/0.31 ≈ 0.226; together they are 1. This is a posterior belief under the stated prior and likelihoods. It is not a guarantee that Billing is the correct team for this individual ticket.
Now tap “Duplicate same record.” A database copy, a second UI rendering, or a paraphrase of the same evidence is not a second independent observation. If we multiply by the same likelihood ratio again, we falsely get approximately 0.965 for Billing: prior odds 3/7 times 8² become 192/7, giving 192/199 ≈ 0.965. That impressive-looking rise is an error. Only a distinct new measurement, with a justified conditional likelihood given what we already saw, may be used for another update.
Can we decide whether the payment-record query itself is worth paying for? Keep the same prior and likelihoods, now add a decision: route to Billing, route to the other team, or use perfect review. Assume a wrong automatic route costs 10, a correct route costs 0, and perfect review costs 2. Before seeing the record, routing to Other costs 0.30×10 = 3, routing to Billing costs 0.70×10 = 7, and review costs 2. The baseline choice is review at expected loss 2.
This is a separate application-side value-of-information toy, using the same prior and test likelihoods. Move the information cost; the conditional actions and weighted risk are recomputed.
A positive record occurs with probability 0.30×0.80 + 0.70×0.10 = 0.31. Its Billing posterior is 24/31. Automatic Billing would risk 10×7/31, about 2.26, so review at 2 is cheaper. A negative record occurs with probability 0.69; its Billing posterior is (0.30×0.20)/0.69 = 2/23. Route to Other then risks 10×2/23 = 20/23, about 0.87, below review's 2.
The gross expected value of this sample information is 2−1.22 = 0.78. After a query cost of 0.20, the net expected benefit is 0.58. Under these toy assumptions, acquire the record only when its incremental expected benefit exceeds its cost: below 0.78, query; above 0.78, skip; at 0.78, tie. The query does not always make us auto-route: a positive still leads to review. This contingent policy is why averaging the two branches matters. MIT's expected value of sample information lesson
These numbers assume the record's sensitivity and false-positive rate are valid on this ticket population, the true route does not change while we wait, review is perfect, and query cost includes every delay or side effect. If any assumption fails, recompute the tree. This is ordinary decision-analysis logic around a hypothetical signal, not a documented Jev feature or a claim that Jev exposes value of information.
The word “independent” here has a precise condition. Two records may share the same bank event, so they are correlated even if they arrive through different services. To multiply two likelihood factors of the simple form, we would need the observations to be conditionally independent given the true route. Application code has to decide how to model dependence; TypeSafe's documented parallel evaluation of questions does not supply that statistical assumption. TypeSafe primitives and independent question evaluation
This full Python example uses exact fractions for both posterior branches and every action risk. Its assertions verify the hand calculation, including the optional query value. NumPy, if installed, checks the same weighted branch losses; no library or Jev call is needed.
# External Bayes and value-of-information analogy; no Jev call.
from fractions import Fraction as F
prior = F(3, 10) # P(Billing)
sensitivity = F(4, 5) # P(+ | Billing)
false_positive = F(1, 10) # P(+ | Other)
wrong_loss = F(10)
review_cost = F(2)
query_cost = F(1, 5)
def minimum_risk(p_billing):
risks = {"billing": (1-p_billing)*wrong_loss,
"other": p_billing*wrong_loss,
"review": review_cost}
action = min(risks, key=risks.get)
return action, risks[action]
p_positive = prior*sensitivity + (1-prior)*false_positive
p_negative = 1-p_positive
assert p_positive and p_negative # Both observed branches possible.
post_positive = prior*sensitivity/p_positive
post_negative = prior*(1-sensitivity)/p_negative
baseline_action, baseline = minimum_risk(prior)
pos_action, pos_risk = minimum_risk(post_positive)
neg_action, neg_risk = minimum_risk(post_negative)
post_observation = p_positive*pos_risk + p_negative*neg_risk
gross_value = baseline-post_observation
net_value = gross_value-query_cost
print("baseline", baseline_action, baseline)
print("positive", p_positive, post_positive, pos_action, pos_risk)
print("negative", p_negative, post_negative, neg_action, neg_risk)
print("post-observation", post_observation,
"with query", post_observation+query_cost,
"net benefit", net_value)
assert (baseline_action, baseline) == ("review", F(2))
assert (p_positive, post_positive) == (F(31,100), F(24,31))
assert (p_negative, post_negative) == (F(69,100), F(2,23))
assert (pos_action, pos_risk) == ("review", F(2))
assert (neg_action, neg_risk) == ("other", F(20,23))
assert (post_observation, gross_value, net_value) == (
F(61,50), F(39,50), F(29,50))
# Repeating the *same* record adds no new evidence.
assert post_positive == F(24,31)
try:
import numpy as np
except ImportError:
print("NumPy comparison skipped")
else:
branch = np.array([float(p_positive), float(p_negative)])
risk = np.array([float(pos_risk), float(neg_risk)])
assert abs(float(branch @ risk)-1.22) < 1e-12
print("NumPy post-observation risk", float(branch @ risk))
There is a useful connection to Bayes estimation and Bayes filtering: both distinguish a prior from evidence and a posterior. This static ticket example has no evolving state transition, so it is not a Bayes filter or a Kalman filter. The connection is the update structure, not a claim that Jev uses those algorithms.
Our code can use a posterior to inform a decision rule, but it must still define the loss of a wrong route or a review. The posterior is an estimate; the action belongs to the application. Next we examine TypeSafe's statement that Jev is trained for calibrated decisions, keeping the public claim separate from an external proper-scoring illustration.
TypeSafe calls its training approach Reinforcement Learning for Calibrated Decisions, or RLCD. Its public primer says the aim is decisions and probabilities rather than generated text, with a well-calibrated 0.8 occurring about 80% of the time across comparable cases. That is a product-level objective, not a published training algorithm. TypeSafe's AI primer
Our A-104 route example can make the idea concrete. Imagine many resolved support tickets where a precisely defined binary event, “Billing is the policy-correct primary route,” occurs on 80% of cases in a comparable group. A forecast of 0.8 for that event is better aligned with the group's frequency than a reflexive 0.5. This sentence describes an external synthetic dataset, not Jev's training data or a Jev output for A-104.
A proper scoring rule is one mathematical way to reward honest probabilistic forecasts. The binary Brier loss is (p−y)², where p is a forecast and y is the resolved 0-or-1 event. We use this familiar rule as a teaching analogy only. TypeSafe does not disclose that RLCD uses Brier loss, nor its exact reward, optimizer, dataset, model architecture, or sampling procedure. Founder's RLCD description
The curve shows expected toy Brier loss under a fixed synthetic 80/20 event mix. Move the slider or reveal individual outcomes; the underlying 80/20 group stays fixed.
At forecast p=0.8, the 80 positive cases each cost (0.8−1)²=0.04, and the 20 negatives each cost (0.8−0)²=0.64. Average over the group: 0.8×0.04 + 0.2×0.64 = 0.16. A p=0.5 forecast costs 0.25 on either outcome, so its group average is 0.25. These are computed on identical outcomes; lower loss is better.
Reveal outcomes above, then compare three forecasts on exactly that observed prefix. The sample column is a realized mean, while the smooth curve uses the stipulated eighty-twenty population. At zero revealed cases, a realized score does not exist.
| Forecast | Expected on 80/20 | Observed prefix |
|---|---|---|
| Current slider | ||
| 0.80 | ||
| 0.50 |
A forecast of .99 is closer than .8 when the next outcome is positive, but it is heavily penalized on a negative outcome: (.99−0)²=.9801. Across the stipulated 80/20 group, its expected Brier is .8×.0001 + .2×.9801 = .1961, above the honest .16. This contrast is about repeated outcomes; it does not say which forecast wins on every single case.
The minimum occurs at the group's true event frequency. Expanding the curve gives p²−1.6p+0.8, whose derivative is 2p−1.6; setting it to zero yields p=0.8. Or complete the square: (p−0.8)²+0.16. This explains why the illustrative score favors an honest probability over a fixed bluff across repeated events. It does not establish how TypeSafe optimizes Jev.
The same reasoning does not depend on eighty percent. If the true group frequency is q, expected binary Brier loss is q(p−1)² + (1−q)p² = (p−q)² + q(1−q). The square is smallest at p=q. For a different constructed group with q=.3, an honest p=.3 gives expected loss .21, while p=.8 gives .46. This is a property of the scoring rule under a correctly defined repeated event, not evidence about TypeSafe's undisclosed training reward.
Revealing a few outcomes can produce a noisy short-run average. Even if eight in ten is the long-run constructed frequency, the first three outcomes need not contain exactly 2.4 positives. The chart's expected curve uses the stipulated group distribution; the reveal control separately displays realized losses on a deterministic 10-outcome sequence. Keep those two quantities distinct.
“Reinforcement learning” in RLCD does not license drawing a specific PPO clip objective, DQN target network, actor-critic rollout, or customer-specific online update. The public pages identify a training goal, not enough implementation detail to reproduce the recipe. TypeSafe also documents that its standard model is not fine-tuned on each customer's requests. The application can still learn about its workflow by logging outcomes and retuning its own policy. TypeSafe models and customization
The executable program below enumerates forecasts from 0 to 1, computes their expected binary Brier loss, and confirms that 0.8 wins on this constructed event mix. NumPy, when installed, evaluates the same grid; no Jev inference is made.
# Synthetic proper-scoring demonstration, not Jev RLCD code.
def expected_brier(p, event_frequency=0.8):
assert 0 <= p <= 1 and 0 <= event_frequency <= 1
return event_frequency*(p-1)**2 + (1-event_frequency)*p**2
grid = [i/100 for i in range(101)]
losses = [expected_brier(p) for p in grid]
best = grid[min(range(len(grid)), key=lambda i: losses[i])]
print(best, expected_brier(0.8), expected_brier(0.5))
assert best == 0.8
assert abs(expected_brier(0.8)-0.16) < 1e-12
assert abs(expected_brier(0.5)-0.25) < 1e-12
assert abs(expected_brier(0.99)-0.1961) < 1e-12
assert abs(expected_brier(0.3, 0.3)-0.21) < 1e-12
assert abs(expected_brier(0.8, 0.3)-0.46) < 1e-12
outcomes = [1,1,0,1,1,1,0,1,1,1] # 8 positive, 2 negative.
assert sum(outcomes) == 8
for count in (1, 3, 10):
prefix = outcomes[:count]
realized = sum((0.8-y)**2 for y in prefix)/count
print("prefix", count, "realized Brier at 0.8", realized)
assert abs(sum((0.8-y)**2 for y in outcomes)/10-0.16) < 1e-12
try:
import numpy as np
except ImportError:
print("NumPy comparison skipped")
else:
p = np.linspace(0, 1, 101)
np_losses = .8*(p-1)**2 + .2*p**2
assert abs(float(p[np.argmin(np_losses)])-best) < 1e-12
print(float(np.min(np_losses)))
The event definition is as important as the score. Predicting a proxy label well does not prove that the system reduced customer harm. Calibration on one group does not guarantee subgroup calibration or that a future shifted population behaves the same. This is why the next chapter separates a classifier's event probability from a long-term action value in a sequential process.
A-104 still needs a primary owner. A typed Choice response can provide a probability distribution over Returns, Billing, and Shipping. The application can use that distribution in a routing rule. But a probability attached to a label does not say what will happen to tomorrow's support queue after today's action.
Keep our local synthetic fixture separate from the API: Returns .60, Billing .38, Shipping .02. Those three numbers sum to one and describe candidate labels for this one ticket. They are not rewards, action values, or measured Jev outputs for A-104. The case's local policy truth remains Billing because duplicate-charge review takes priority.
A sequential decision needs different ingredients. We must specify actions, a reward after
each action, a state transition, and a policy for what happens next. In reinforcement
learning, a discounted two-step return is G = r₀ + γr₁. The discount
γ controls the weight of the later reward; it does not change the Choice
probabilities. See Decision Making Under
Uncertainty Workbook — MDPs & Bellman Backup and Deep Reinforcement Learning — MDPs and value functions.
Work the default calculation by hand. At γ = .9, action A gives 1 + .9 × 0
= 1. Action B gives 0 + .9 × 5 = 4.5. B therefore has the larger two-step
return, even though A pays more immediately. If γ = 0, the comparison reverses: A
gives 1 and B gives 0. The slider recalculates these returns from the two reward pairs; the
route-probability card stays outside that calculation.
The diagram's B reward of 5 is deterministic by construction. Real future rewards may be
uncertain. In a separate stochastic variant, B pays 5 one step later with
probability s, otherwise 0. At s=.5 and γ=.9,
expected B return is 0 + .9 × (.5×5 + .5×0) = 2.25. At
s=.2, it is .9×1 = .9, below A's return of 1.
Move the success slider to see the crossover. This expectation requires an
action-conditioned outcome estimate; a Choice label mass is not that estimate.
A contextual bandit is a one-step relative: it sees a context, chooses an action, and
observes that action's outcome. A sequential MDP additionally models the next state and future
actions. Merely seeing a classifier's probability vector supplies neither action-conditioned
rewards nor transitions. To estimate a true Q(state, action), a learner would
need outcome data under relevant actions and a specified objective. This exercise is an
external decision analogy, not Jev's undisclosed training procedure.
P(Billing)=.38 is a belief about a
supplied answer option. Q(state, Billing) would be an expected future return
under a reward and policy. The numbers share neither units nor meaning, even if both happen to
lie between zero and one.Try the second control below the timeline. It computes the same single-row inverse-propensity contribution while you change the logged action propensity. When q gets smaller but stays positive, the weight grows. If that action had no chance of being logged at all, the “no support” switch correctly refuses to compute a weight. Zero support is missing evidence, not infinite evidence.
Logged outcomes bring a second trap: a policy may almost never try an action in some context,
so its alternative outcome is missing. In a separate one-step bandit example, suppose 100
logged decisions contain one target-policy action with observed reward 12 and logging
propensity .2. Its contribution to an inverse-propensity estimate is (12 ÷ .2) ÷ 100 =
.6. That correction is meaningful only when the logging propensity is known, positive
wherever the target policy acts, and the logged context/action/reward assumptions fit the
target. It is not a Jev metric; Bandits & Preference-Based Learning — Contextual Bandits develops the related setting.
Run the short program from scratch. It uses only Python's standard library. The
generator-expression and sum line provide an equivalent dot-product calculation
for the same two rewards, without introducing a library's unrelated learning algorithm.
# Constructed two-step rewards; not a Jev call.
rewards = {"A": (1.0, 0.0), "B": (0.0, 5.0)}
def two_step_return(pair, gamma):
assert 0 <= gamma <= 1
return pair[0] + gamma * pair[1]
def same_return_by_sum(pair, gamma):
return sum(r * weight for r, weight in zip(pair, (1.0, gamma)))
gamma = 0.9
values = {name: two_step_return(pair, gamma)
for name, pair in rewards.items()}
assert values == {"A": 1.0, "B": 4.5}
assert all(values[name] == same_return_by_sum(pair, gamma)
for name, pair in rewards.items())
assert max(values, key=values.get) == "B"
def stochastic_b_expected(gamma, success_probability):
assert 0 <= success_probability <= 1
outcomes = ((5.0, success_probability),
(0.0, 1 - success_probability))
expected_later = sum(reward * probability
for reward, probability in outcomes)
return gamma * expected_later
assert stochastic_b_expected(.9, .5) == 2.25
assert abs(stochastic_b_expected(.9, .2) - .9) < 1e-12
assert stochastic_b_expected(.9, .2) < values["A"]
assert two_step_return(rewards["A"], 0) > two_step_return(rewards["B"], 0)
print(values, "preferred:", max(values, key=values.get))
The choice between A and B belongs to this stipulated reward model. The next chapter will return to A-104 and put the typed-fixture judgment, a cost-based policy, and the labeled outcome into one controllable workflow. That step makes it possible to see which errors belong to the model, the policy, or the data.
Now join the lesson's pieces in a small, inspectable workflow. We will take a ticket state, ask a bounded routing question, feed a synthetic local fixture into code, choose auto-route or review, and compare the resulting action with a labeled outcome. No button on this page calls Jev or processes a real customer's ticket.
Our recurring A-104 message asks for a shoe exchange and a duplicate-charge refund. Its fixture gives Returns .60, Billing .38, Shipping .02; our stated local policy says Billing is primary. The response shape has three permitted choices. The policy that follows is authored here in code, not learned automatically from the shape. TypeSafe documents Choice as a bounded answer space, while the application owns downstream behavior. TypeSafe — Choice primitive
The harness has six stages, shown one at a time so a phone can display the complete current object. Step through them or press Play. Change the ticket, costs, or the fixture's shifted-prediction switch and repeat the run. Every figure in the summary is recalculated from the selected fixture and controls. The time slider does not hide any paid model request; the “fixture answer” stage simply reads an array stored locally.
Compute A-104 by hand at the initial settings. The highest local fixture mass is Returns .60,
leaving .38 + .02 = .40 mass on other labels. Under a deliberately simple
symmetric-loss model, auto-routing to Returns has estimated loss .40 × 10 =
4. Perfect human review is stipulated to cost 2 and to route correctly, so the policy
selects review. Its realized synthetic outcome is a Billing assignment at cost 2. The
fixture's .60 is not a verified chance of correctness; the estimated loss is conditional on
treating these numbers as useful probabilities.
Turn on the shifted fixture. For A-104 it now favors Returns .84, Billing .12, Shipping .04.
At the same costs, auto-route's estimated loss drops to (.12 + .04) × 10 =
1.6, so the code may auto-route to Returns. Yet the policy label remains Billing,
giving a realized wrong-route cost of 10. This is a controlled failure: a confident,
misspecified fixture can make a mathematically consistent rule choose badly. Moving the
review-cost slider can produce a different decision, and the stage panel explains which object
changed.
The outcome label is available to this teaching harness because we authored it; an operational system often learns it later from a resolved ticket or reviewer. Human review is assumed perfect here only to isolate the loss comparison. If reviewers make mistakes, their error rate and delays must enter the policy model. A forced binary auto/review switch also omits valid operational choices such as requesting more evidence or escalating a safety-sensitive case.
Run the same pipeline without a browser. The code exposes three pure functions so you can change a prediction, a cost, or a label and see which stage changes. The assertion for shifted A-104 is intentionally a wrong-control case: the expected calculation says auto, while the labeled outcome says that auto route was wrong.
# Synthetic, deterministic harness. No SDK or network request.
CASES = {
"A-104": {"truth": "billing", "base": (.60, .38, .02),
"shift": (.84, .12, .04)},
"A-106": {"truth": "returns", "base": (.82, .12, .06),
"shift": (.45, .48, .07)},
"A-107": {"truth": "shipping", "base": (.14, .12, .74),
"shift": (.38, .44, .18)},
}
ROUTES = ("returns", "billing", "shipping")
def classify_fixture(case_id, shifted=False):
case = CASES[case_id]
probs = case["shift" if shifted else "base"]
assert abs(sum(probs) - 1) < 1e-12
return dict(zip(ROUTES, probs))
def choose_action(probs, wrong_cost=10, review_cost=2):
route = max(ROUTES, key=lambda name: probs[name])
auto_estimate = (1 - probs[route]) * wrong_cost
if auto_estimate < review_cost:
return {"kind": "auto", "route": route,
"estimate": auto_estimate}
return {"kind": "review", "route": None,
"estimate": review_cost}
def evaluate_outcome(case_id, action, wrong_cost=10, review_cost=2):
truth = CASES[case_id]["truth"]
if action["kind"] == "review":
return {"assigned": truth, "correct": True, "cost": review_cost}
correct = action["route"] == truth
return {"assigned": action["route"], "correct": correct,
"cost": 0 if correct else wrong_cost}
base = choose_action(classify_fixture("A-104"))
shifted = choose_action(classify_fixture("A-104", shifted=True))
assert base == {"kind": "review", "route": None, "estimate": 2}
assert evaluate_outcome("A-104", base)["cost"] == 2
assert shifted["kind"] == "auto" and shifted["route"] == "returns"
assert abs(shifted["estimate"] - 1.6) < 1e-12
assert evaluate_outcome("A-104", shifted) == {
"assigned": "returns", "correct": False, "cost": 10}
print(base, shifted, evaluate_outcome("A-104", shifted))
A production harness can place a classifier at a focused decision point while deterministic
code still owns the route/review policy. LangChain documents TypeSafeClassifier
as a Runnable: configure named Choice criteria, call
.invoke(state), and read response.choices["route"]. This
shape-only integration sketch requires the package and an API key; it is
not run by this lesson, and the fixture numbers above are not its output.
LangChain — TypeSafe integration
# Integration sketch only: requires langchain-typesafe and API credentials.
from langchain_typesafe import Choice, TypeSafeClassifier
classifier = TypeSafeClassifier(questions={
"route": Choice(
instructions="Which team is the primary route?",
criteria={"returns": "Exchange or return without billing priority.",
"billing": "Duplicate-charge review takes priority.",
"shipping": "Shipment delivery or tracking issue."},
)
})
# response = classifier.invoke(ticket_text)
# probabilities = response.choices["route"].probabilities
# Application code then validates and applies its own route/review policy.
In this module, min over the two estimated losses is equivalent to the explicit
if branch; the branch makes the ownership of the rule easier to see. A real
integration could replace classify_fixture with a typed API response, but it
would still need the same cost model, fallbacks, logging, and outcome evaluation. The next
chapter asks what evidence from many labeled tickets would justify that integration.
The harness can be fast and tidy while sending a valuable ticket to the wrong team. To know whether its route policy helps, record predictions, decisions, and later resolved labels for many cases. Keep the label source explicit: a human-adjudicated policy target is different from another model's agreement with a prediction. The ten rows below are a constructed dataset, not a Jev benchmark.
The dataset is named jev-synthetic-routing-v1, the threshold rule
top-mass-threshold-v1, and the probabilities local-hand-authored-probabilities-v1. Those are deliberately separate version IDs. Changing labels, the routing threshold, or a model could change a metric for different reasons; versioning lets us reconstruct which choice caused the change.
For each case, code takes the top route only when its fixture mass reaches the slider threshold. Otherwise it sends the ticket to review. Coverage is the fraction auto-routed. Acted accuracy counts correct assignments among auto-routed cases, not among all ten. A policy can inflate acted accuracy by reviewing almost everything, so show coverage and error counts together. The review channel is stipulated perfect at cost 2; a wrong auto-route costs 10.
| Case | Top route / mass | Resolved label | Proxy label | Slice | Action |
|---|
Check the default .65 result by hand against the full log. A-104 at .60 and
A-110 at .50 go to review. The other eight are auto-routed. Of those, A-109 predicts Billing
while its resolved label is Returns, and A-112 predicts Returns while its resolved label is
Billing. Thus coverage = 8/10 = .80, acted accuracy = 6/8 = .75, and
the count of wrong auto-routes is 2. Those are the dataset's actual counts; they are not a
rounded story about an unshown experiment.
The shifted slice consists of A-109, A-110, and A-112. At .65, two of its three cases are auto-routed and both are wrong; the remaining one is reviewed. This small sample is not a reliable deployment estimate, but it exposes a failure that the overall 75% acted accuracy could conceal. A real evaluation should split by time, product, language, and other relevant conditions, with sufficiently many independently resolved examples.
The proxy column is not the resolved label. It is a synthetic stand-in for model consensus: for A-104 it says Returns while the policy label says Billing, and for A-109 it agrees with the wrong Billing prediction. Toggling the proxy shows disagreements rather than replacing ground truth in the main accuracy and cost metrics. TypeSafe's published workflow comparison uses reference probabilities averaged from GPT-6 Astra and Fable 5.1; those references are not human-adjudicated truth labels. Its 193.6× faster and 444.6× cheaper high-end workflow examples are vendor-reported comparisons under its benchmark setup, not universal latency, cost, or correctness guarantees. TypeSafe launch post — Workflow evals
To compare latency or cost fairly, fix the task schema, traffic mix, hardware/service settings, model identifiers, caching, retries, and quality target. TypeSafe documents model aliases; pinning an exact model version matters because an alias can move over time. Report uncertainty and error severity alongside throughput. A cheaper route that doubles damaging billing mistakes may be a poor operational trade. TypeSafe — Models and aliases
There are model-specific reasons to build these checks. TypeSafe's Jev 1.13 jaggedness guide says literal wording and contradictory instructions versus criteria can confuse a judgment; large irrelevant state can distract it. It also says counting, numerical precision, and date ordering belong in code. For this ticket, keep the duplicate-charge policy in clear Choice criteria, filter unrelated customer history, and compute money, dates, and loss arithmetic with ordinary program logic. TypeSafe — Jev 1.13 jaggedness
Independently asked questions need not obey cross-question identities automatically. If one answer says “duplicate charge likely” and another says “refund request unlikely,” inspect the wording and the record; do not force a joint interpretation by multiplying marginal probabilities. Test disagreements as explicit cases, then enforce any required structural invariant in code. The vendor guide itself recommends asking each decision directly. TypeSafe — structural invariants
Run the complete ten-row check. The standard-library sum forms are equivalent to
explicit count loops on the same rows. The assertions pin the displayed default values and a
wrong-control case; move threshold to see the trade-off.
# Ten synthetic cases: id, (Returns, Billing, Shipping), truth, proxy, slice.
ROUTES = ("returns", "billing", "shipping")
ROWS = [
("A-104",(.60,.38,.02),"billing","returns","standard"),
("A-105",(.08,.88,.04),"billing","billing","standard"),
("A-106",(.82,.12,.06),"returns","returns","standard"),
("A-107",(.14,.12,.74),"shipping","shipping","standard"),
("A-108",(.73,.20,.07),"returns","returns","standard"),
("A-109",(.12,.79,.09),"returns","billing","shifted"),
("A-110",(.45,.50,.05),"billing","billing","shifted"),
("A-111",(.06,.16,.78),"shipping","shipping","standard"),
("A-112",(.68,.28,.04),"billing","returns","shifted"),
("A-113",(.22,.66,.12),"billing","billing","standard"),
]
DATASET_ID = "jev-synthetic-routing-v1"
POLICY_ID = "top-mass-threshold-v1"
FIXTURE_MODEL_ID = "local-hand-authored-probabilities-v1"
def evaluate(rows, threshold=.65, wrong_cost=10, review_cost=2):
acted = correct = errors = reviewed = shifted_errors = 0
total_cost = 0
for case_id, probs, truth, proxy, segment in rows:
assert abs(sum(probs) - 1) < 1e-12
index = max(range(len(probs)), key=lambda i: probs[i])
if probs[index] < threshold:
reviewed += 1
total_cost += review_cost # Teaching assumption: perfect review.
continue
acted += 1
match = ROUTES[index] == truth
correct += int(match)
errors += int(not match)
shifted_errors += int(segment == "shifted" and not match)
total_cost += 0 if match else wrong_cost
return {"coverage": acted / len(rows),
"acted_accuracy": correct / acted if acted else None,
"acted": acted, "correct": correct, "errors": errors,
"reviewed": reviewed, "shifted_errors": shifted_errors,
"total_cost": total_cost}
result = evaluate(ROWS)
assert (result["acted"],result["correct"],result["errors"],
result["reviewed"],result["shifted_errors"],
result["total_cost"]) == (8,6,2,2,2,24)
assert result["coverage"] == .8 and result["acted_accuracy"] == .75
assert sum(1 for row in ROWS if row[2] != row[3]) == 3
print(DATASET_ID, POLICY_ID, FIXTURE_MODEL_ID, result)
These metrics tell us what the constructed workflow did on known cases; they do not identify why a model predicted badly. The final chapter connects this workflow to estimation, sequential control, and evaluation tools so you can choose the right next component instead of asking one typed classifier to do every job.
Return to A-104. The case state is a message about wrong-size shoes and a duplicate charge. A bounded routing question asks for one of three allowed labels. A synthetic fixture assigns Returns .60, Billing .38, Shipping .02. Code then compares its cost estimates, decides whether to route or review, and later checks a resolved label. Each arrow in this chain has a different owner.
Estimate: a typed judgment can provide option probabilities, but their usefulness must be tested on resolved examples. Jev's documented Choice primitive bounds the answer space; it does not write a customer reply or guarantee the chosen label. The probabilities in our A-104 exercise are local invented data, not a response from a live model. TypeSafe — Choice
Decide: the application owns allowed actions, losses, and fallback rules.
For A-104, .60 + .38 + .02 = 1. The top label is Returns, with other-label mass
.38 + .02 = .40. At wrong-route cost 10, its estimated auto loss is 10 ×
.40 = 4. A stipulated perfect review costs 2, so this particular rule chooses review.
That is a policy calculation, not a feature hidden inside a Choice value.
Learn and evaluate: a future model or policy can improve only if the workflow records suitable states, actions, propensities when relevant, and independently assessed outcomes. Review labels can help evaluate an auto-route rule, but observed data may be biased by which cases were sent to review. The ten-ticket exercise showed 8 auto-routes and only 6 correct ones at its default threshold; it is a teaching sample, not evidence about a deployed system.
The review choice depends on review quality too. If review itself were correct only
90% of the time and its mistakes had the same cost 10, expected review loss would
be 2 + (1 − .90) × 10 = 3, still below the estimated auto loss 4.
At 70% correct, it would be 2 + .30 × 10 = 5, making auto cheaper
under these particular assumptions. The displayed A-104 result retains the
explicit perfect-review assumption; this counterexample shows why that
assumption should be measured rather than silently treated as universal.
The map's audit card names the owner, required evidence, and a concrete failure check for each job. For example, a route probability can be syntactically valid while disagreeing with the resolved policy label; a policy can compute a correct minimum over the wrong cost table; and an evaluation can look strong overall while its shifted slice fails. Tap through the card to see how a single ambiguous ticket becomes several different engineering questions.
Choose Estimate to connect this lesson to priors and evidence updates in Bayes Estimation — Prior, Likelihood, Posterior and belief states in POMDP — Observations & Belief. Those are mathematical connections, not claims about Jev's internal algorithm. Choose Decide for Decision Making Under Uncertainty Workbook — Value of Information and MDPs, where costs, information, and future outcomes become explicit.
Choose Learn for Deep Reinforcement Learning — MDPs and policy objectives; its state/action/reward setting goes beyond a typed classifier. Choose Evaluate for Calibration: Knowing What the Model Knows — Cost-Coverage and Ask/Answer/Defer and AI Evaluation — measuring system behavior. Open the linked lesson and choose the named chapter.
A customer-facing explanation is another job again. A generative model or a human can compose and inspect that response after a route decision; a typed route label is not prose. Operational guardrails must also govern side effects such as issuing refunds or contacting a customer. The site’s Agents & Tool Use — Guardrails lesson develops that separation. Keep typed judgment, application policy, tool permissions, and human review visible rather than conflating them into one opaque “AI action.”
A minimal audit record should preserve the case and fixture version, selected route masses, policy version and costs, chosen action, and later label source. If a reviewer corrected a route, log the correction without silently treating an unreviewed case as correct. This separation prevents a model update, a policy edit, or a change in who gets reviewed from masquerading as the same experiment.
The complete pure-Python trace below reprises the case. Its assertion checks both the
probability total and the review decision; replacing the synthetic fixture with a real typed
response would still leave policy and outcome evaluation in application code. The
standard-library sum is equivalent to adding the three displayed masses manually.
# A-104 is a constructed fixture, not a Jev API result.
case = {
"id": "A-104", "truth": "billing",
"probabilities": {"returns": .60, "billing": .38, "shipping": .02},
}
def validate_fixture(case):
probs = case["probabilities"]
assert set(probs) == {"returns", "billing", "shipping"}
assert all(0 <= value <= 1 for value in probs.values())
assert abs(sum(probs.values()) - 1.0) < 1e-12
assert case["truth"] in probs
def decide_and_check(case, wrong_cost=10, review_cost=2):
probs = case["probabilities"]
assert abs(sum(probs.values()) - 1.0) < 1e-12
route = max(probs, key=probs.get)
wrong_mass = sum(p for label, p in probs.items() if label != route)
auto_estimate = wrong_cost * wrong_mass
action = "review" if review_cost <= auto_estimate else route
# This toy assumes review assigns the known policy label correctly.
assigned = case["truth"] if action == "review" else route
realized_cost = review_cost if action == "review" else (
0 if assigned == case["truth"] else wrong_cost)
return {"top_route": route, "wrong_mass": wrong_mass,
"auto_estimate": auto_estimate, "action": action,
"assigned": assigned, "realized_cost": realized_cost}
def review_estimated_loss(review_cost, review_accuracy, wrong_cost):
assert 0 <= review_accuracy <= 1
return review_cost + (1 - review_accuracy) * wrong_cost
assert abs(review_estimated_loss(2, .9, 10) - 3) < 1e-12
assert review_estimated_loss(2, .7, 10) > 4
bad_case = {**case, "probabilities": {"returns": .6, "billing": .6,
"shipping": .02}}
try:
validate_fixture(bad_case)
except AssertionError:
pass # Wrong control: probabilities must sum to one.
else:
raise AssertionError("invalid fixture accepted")
result = decide_and_check(case)
assert result["top_route"] == "returns"
assert abs(result["wrong_mass"] - .4) < 1e-12
assert abs(result["auto_estimate"] - 4) < 1e-12
assert result["action"] == "review"
assert result["assigned"] == "billing" and result["realized_cost"] == 2
print(result)
You can reuse this recipe: define the state and bounded question, keep the answer's provenance, specify a cost-based rule, log the action, and test against independently resolved outcomes. If you need future state control, turn to an MDP; if you need better beliefs after new evidence, turn to Bayesian estimation; if you need confidence-aware deployment, turn to calibration and evaluation. A typed judgment is a useful component because code can consume it precisely, and that precision makes the remaining responsibilities easier to inspect.