AI Harness Engineering

Jev & System One

From a typed judgment to a justified action: probability, confidence, calibration, Bayesian estimation, and reinforcement learning.

Prerequisites: basic arithmetic and code variables. Probability and decision theory are built from the ground up.
12
Chapters
12
Interactive Labs
1
Decision Harness

Source review: September 19, 2026 · Jev 1.13 documentation · examples are explicitly labeled.

Chapter 0: A valid answer can still be wrong

A customer writes: “My running shoes are the wrong size, and my card was charged twice. Please swap them and refund the extra charge.” You are building the first routing step in a support system. Should the ticket go to Returns, Billing, or Shipping?

The ticket is our synthetic A-104 case file. Its local policy says Billing owns a duplicate-charge review even when the same message also requests an exchange. Billing is therefore the human-specified primary route for this exercise. That policy is part of our example, not a claim about Jev or any real support organization.

Case file A-104 · synthetic fixture
State: wrong-size shoes + duplicate card charge + exchange/refund request.
Allowed route: Returns, Billing, or Shipping.
Local policy truth: Billing is primary because duplicate-charge review takes priority.

A free-text model might reply “billing team, please” or “send it to the people who handle charges.” A developer then has to parse those strings before code can route the ticket. Jev is documented as returning typed values from caller-defined answer spaces rather than generated prose; that solves the answer-shape problem. It does not make a chosen option correct by itself. TypeSafe Introduction — typed questions and outputs

Try the routing buttons below. “Returns” is a permitted label, so the schema gate passes. But our stated policy says Billing is the primary owner, so the correctness gate fails. The two gates answer different questions: can software read this value? and does it match the case's labeled target?

Two gates for one response

This is a local, deterministic teaching fixture. Change the response and watch the gates; no Jev API is called.

A schema specifies which forms of answer the program accepts. Here it is a simple set of three strings. A ground-truth label is the answer our policy or a resolved case says should have been chosen. Code can check schema membership without knowing the real-world answer; the second check needs a trustworthy label or review.

Count three illustrative responses to the same synthetic ticket: Billing, Billing, and Returns. Every string is in the allowed set, so schema validity is 3 ÷ 3 = 1.00, or 100%. Two routes agree with the policy truth, so case accuracy is 2 ÷ 3 ≈ 0.667, or 66.7%. The same three responses support both numbers because the tests measure different properties. These are arithmetic on a teaching fixture, not measured Jev rates.

validity = allowed responses / all responses = 3 / 3 = 1.00
accuracy = policy-matching responses / all responses = 2 / 3 ≈ 0.667

Change the three responses yourself. The denominators stay at three, while the two numerators update separately. Selecting an unlisted label fails both checks; selecting the wrong allowed team passes only the first. This small audit is useful because a schema-only test suite could report a clean run while overlooking a routing regression.

Audit three local responses

These are repeated illustrative responses to the same A-104 case, not observed model trials.

ResponseSchemaPolicy match
Billingpassmatch
Billingpassmatch
Returnspassmismatch

Schema: 3 of 3. Policy match: 2 of 3.

One more option makes the difference vivid. “Escalations” is not in our three-option schema, so it fails the first gate even if a human might decide to escalate later. A valid value can be wrong; an invalid value cannot be safely consumed as this field. A production workflow may have a separate explicit escalation action in code.

Common mistake: TypeSafe's reported zero schema/type-error figure is about staying inside an answer shape. It does not mean zero incorrect classifications, zero unsafe actions, or zero mistaken probabilities. The company's own discussion calls schema matching a guarantee rather than a ground-truth accuracy measurement. TypeSafe launch post — Hallucination and Type-safety

Implement the two checks separately. The code below uses no model and no third-party package: it simply verifies the response contract and compares a valid response with the synthetic policy label. Run it from top to bottom to see all three cases and the two rates.

# Pure Python; no Jev API call.
ALLOWED = {"returns", "billing", "shipping"}
TRUTH = "billing"  # Synthetic A-104 policy label.


def schema_valid(route):
    return isinstance(route, str) and route in ALLOWED


def policy_correct(route):
    return schema_valid(route) and route == TRUTH


responses = ["billing", "billing", "returns"]
valid_count = sum(schema_valid(route) for route in responses)
correct_count = sum(policy_correct(route) for route in responses)
validity = valid_count / len(responses)
accuracy = correct_count / len(responses)
print(valid_count, correct_count, validity, round(accuracy, 3))
assert (valid_count, correct_count) == (3, 2)
assert schema_valid("returns") and not policy_correct("returns")
assert not schema_valid("escalations")

Python's standard-library json.loads can parse an incoming JSON string, but parsing alone answers only whether the text is valid JSON. The equivalent membership check still needs route in ALLOWED, and semantic correctness still needs a policy label. For example, json.loads('{"route":"returns"}')["route"] yields a readable yet wrong route for A-104.

Why this matters: A typed model output is useful because it gives code a predictable input. The application still owns the rule that connects an answer to an action and the evaluation that tests that rule on labeled cases.

We now know why a bounded answer is helpful and why it is not enough. Next we will open the actual request contract: what state does Jev see, what questions can we ask, and what kinds of values come back?

A-104's synthetic policy truth is Billing. Which response is valid under the three-option schema but wrong under that policy?

Chapter 1: One state, three answer shapes

Suppose the support program has only A-104's message. It needs a route, a rough urgency level, and a yes/no judgment about the duplicate charge. Those are different questions about the same evidence. A single unstructured “analyze this ticket” answer would hide which value belongs to which task.

TypeSafe calls the evidence state. A Jev request sends state alongside typed questions. The state can be a text string or a JSON object/array containing text values; Jev 1.13 does not directly inspect an image, audio clip, or video. We use a JSON-like case file so each question can point to the message, order, or policy by name. TypeSafe State documentation

State · A-104 synthetic case file
message: wrong-size shoes, duplicate charge, exchange and extra-charge refund requested.
policy: Billing owns duplicate-charge review.
order: one order with two captured card charges in this teaching example.

A Choice question asks which of a fixed list fits. For the primary department, the criteria are Returns, Billing, and Shipping. Its documented response has choice, one probability per option, and a separate confidence field. In the local A-104 fixture below, Returns has .60, Billing .38, and Shipping .02; these are invented demonstration values, not a Jev call or a measured Jev prediction. TypeSafe Choice documentation

A Score question uses ordered levels defined by the caller, such as 0 = routine, 1 = soon, and 2 = urgent. The response includes a distribution over those levels and a numeric position along them. The local fixture assigns probabilities 0, .70, .30. By hand, 0 × 0 + 1 × .70 + 2 × .30 = 0 + .70 + .60 = 1.30. That is a place on our urgency rubric, not a 130% probability or an exact clock deadline. TypeSafe Score documentation

score = Σ (level × level probability) = 0×0 + 1×0.70 + 2×0.30 = 1.30

A Noul asks one yes/no question and returns the modeled probability of “yes” between 0 and 1. Our fixture uses .80 for “Does the message report a duplicate charge?” A .50 Noul would mean yes and no have equal modeled probability; it would not mean medium urgency. The documented Noul response has no separate confidence field. TypeSafe Noul documentation

Switch the answer contract

The A-104 state remains the same. Select which question type the program asks. Bars and values are deterministic local fixtures, not live Jev output.

Urgent mass.30
Noul yes.80

Move the urgent-mass slider while Score is selected. Its mass comes from the “soon” level, so the two nonzero masses still sum to one. The weighted position shifts continuously between level one and level two; this is the meaning of a score between labels. Move the Noul slider and the yes/no line moves independently. Neither slider is changing Jev's weights or running its inference service.

For Choice, the three fixture probabilities must sum to one: .60 + .38 + .02 = 1.00. The selected route is the largest probability, Returns. But the human-authored A-104 primary label is Billing, so this particular fixture is an example of a plausible, typed, wrong answer. The separate confidence formula is not disclosed in the cited API pages; do not infer it from our bars.

Try changing the primitive while keeping the ticket fixed. The question type controls what the result means and which fields your code may read. Choice is appropriate for mutually exclusive routes; Score is for a defined ordered rubric; Noul is for one statement that might be true. A production system can ask all three in one request, but they are not interchangeable gauges.

Common mistake: Do not read Score 1.30 or Noul .80 as Jev's probability that the whole routing decision is correct. They answer different, narrower questions. Also do not add a confidence field to Noul: the current primitive documentation does not return one. TypeSafe primitive response table

The following plain Python reproduces the fixture arithmetic and fields without implementing or pretending to run Jev. The keys illustrate some documented fields, not the full SDK serialization; the numbers are supplied by this lesson. Running it prints the same choice, score, and Noul values shown above.

# Local fixture only. No learned model and no network request.
import math

CHOICE_CRITERIA = {"returns", "billing", "shipping"}
route_probabilities = {
    "returns": 0.60,
    "billing": 0.38,
    "shipping": 0.02,
}
level_probabilities = [0.0, 0.70, 0.30]
duplicate_yes_probability = 0.80


def make_choice(probabilities):
    assert 1 <= len(probabilities) <= 255
    assert set(probabilities) == CHOICE_CRITERIA
    assert all(isinstance(key, str) for key in probabilities)
    assert all(math.isfinite(p) and p >= 0 for p in probabilities.values())
    assert abs(sum(probabilities.values()) - 1.0) < 1e-9
    selected = max(probabilities, key=probabilities.get)
    return {"choice": selected, "probabilities": probabilities}


def make_score(probabilities):
    assert len(probabilities) == 3  # This fixture has three named levels.
    assert all(math.isfinite(p) and p >= 0 for p in probabilities)
    assert abs(sum(probabilities) - 1.0) < 1e-9
    score = sum(level * p for level, p in enumerate(probabilities))
    return {"score": score, "probabilities": probabilities}


def make_noul(yes_probability):
    assert math.isfinite(yes_probability)
    assert 0.0 <= yes_probability <= 1.0
    return {"noul": yes_probability}


choice = make_choice(route_probabilities)
score = make_score(level_probabilities)
noul = make_noul(duplicate_yes_probability)
print(choice["choice"], score["score"], noul["noul"])
assert choice["choice"] == "returns"
assert abs(score["score"] - 1.30) < 1e-9
assert noul == {"noul": 0.80}

The short equivalent for the Score arithmetic is Python's built-in sum(i*p for i,p in enumerate(level_probabilities)); it gives the same 1.30. The actual TypeSafe SDK offers Choice, Score, and Noul question constructors, but only the service supplies real model probabilities and the separate Choice/Score confidence field. Our plain dictionaries intentionally leave that undisclosed confidence value out. TypeSafe SDK-shaped question examples

If you later replace this fixture with Jev, define the actual questions with the SDK rather than asking the service for the displayed numbers. The following request shape follows TypeSafe's Python SDK examples, including descriptions for each Choice criterion and a pinned model version. It is an optional integration illustration: it requires typesafe_sdk, an API key, and a network request; we did not execute it or claim its response would match the fixture.

# Optional real SDK request shape; not run by this lesson.
from typesafe_sdk import Choice, Score, Noul, TypeSafeClient

state = {
    "message": "Wrong-size shoes and two charges; swap and refund extra charge",
    "policy": "Billing owns duplicate-charge review",
}
questions = {
    "route": Choice(
        instructions="Which team owns the primary review under policy?",
        criteria={
            "returns": "Exchange and wrong-size requests",
            "billing": "Duplicate charges and payment review",
            "shipping": "Delivery and missing-package issues",
        },
    ),
    "urgency": Score(
        instructions="How urgent is this support request?",
        criteria=["Routine", "Soon", "Urgent"],
    ),
    "duplicate": Noul(
        instructions="Does the message report a duplicate card charge?",
    ),
}
# Only run with your own configured credentials and network access:
# with TypeSafeClient(model="jev-1.13.0") as client:
#     response = client.system_one(state=state, questions=questions)

The versioned model ID is documented as available at the time of this lesson; aliases such as jev-latest may later point to a different version. This request defines questions. It does not manufacture the answer probabilities. Our runnable standard-library code above remains the reproducible comparison. TypeSafe Python multi-question example; TypeSafe model IDs and aliases

Design choice: The developer defines the answer space before calling Jev. If none of the offered routes fits a case, a Choice needs an explicit “other” or “none of the above” option. The fixed shape keeps code predictable, but a poorly chosen option list can still force a bad choice. Choice criteria guidance

We have three useful judgments about one state. The next problem is coordination: if all three questions use A-104, should we ask them together, and does “independent evaluation” mean the facts themselves are statistically independent?

Which primitive should ask for A-104's primary route among Returns, Billing, and Shipping?

Chapter 2: Ask separately, compose in code

A-104 contains an exchange request and a duplicate charge. If we ask “What should happen with this customer?” we are hiding several decisions inside one sentence. A route, an urgency rating, and the existence of a duplicate-charge claim each need their own answer contract.

TypeSafe's interface lets a request carry multiple typed questions about one state. Each question sees the same state and is evaluated independently of the other answers. The company says those questions are evaluated in parallel, so one request can collect several narrow judgments. This is about the model's question evaluation, not a claim that facts in the ticket are statistically independent. TypeSafe Primitives — asking multiple questions

Three questions over one A-104 state
Choice: Which department is the primary route?
Score: How urgent is this case on a defined three-level rubric?
Noul: Does the message report a duplicate charge?
The values in this chapter are synthetic fixtures, not live Jev responses.

Click the lanes below. In the ordinary path, route, urgency, and duplicate-charge questions all see the same initial A-104 state; turning one off leaves the other local fixture values unchanged. Stepping to “code combines” shows our invented rule: if the duplicate-charge probability is at least .75, flag Billing review; otherwise use the route choice. The dependency switch illustrates a different path where urgency cannot be constructed yet.

Parallel questions, then a code-owned branch

Toggle questions or a dependency, then advance the two-stage flow. One lane per row keeps the same structure readable on a narrow screen.

An unfetched record alone does not force a second model call: code could fetch a known record before asking any questions. For a genuine dependency here, the first-stage Noul answer about the duplicate charge determines which record code may fetch and which urgency question it can construct. If the answer crosses our illustrative threshold, code fetches a payment record; otherwise it fetches order-fulfillment history. Only then can the application form the enriched state and ask urgency in a second request. Route and duplicate-charge questions in the first batch still see the same original state. TypeSafe Primitives — when one question depends on another

“Evaluated independently” is easy to mishear as “the events are independent.” To see the difference, use the local hypothetical marginal probabilities P(duplicate charge) = .80 and P(urgent) = .50. We have not been given a joint probability that both are true. The overlap cannot exceed the smaller marginal, so its upper bound is min(.80,.50) = .50.

At the low end, the two events together occupy .80 + .50 = 1.30 of a unit population if counted separately. Their overlap must be at least the excess over one: max(0, .80 + .50 − 1) = .30. Thus P(both) can lie from .30 to .50. Multiplying .80 × .50 = .40 selects one possible value only under the extra assumption that the events are statistically independent. Parallel computation did not grant us that assumption.

max(0, P(A)+P(B)−1) ≤ P(A and B) ≤ min(P(A),P(B))
max(0, .80+.50−1)=.30 ≤ P(both) ≤ .50; product .40 requires independence

The bounds are real alternatives, not just algebra. Move the overlap below while both marginal probabilities stay fixed. When the overlap grows, the “neither” group grows too, and each one-event-only group shrinks. The four groups always add to one. This is a local probability table, not an output that Jev supplies for these two separate questions.

Same marginals, different joint worlds

Drag the possible overlap from .30 to .50. Every cell is recalculated from P(duplicate)=.80 and P(urgent)=.50.

Both true.40

Common mistake: A model can answer two questions in parallel while the underlying facts are strongly related. “Duplicate charge” and “urgent support need” could co-occur for many real reasons. Never multiply the reported marginal probabilities into a joint probability just because the API batched the questions.

Here is a runnable local approximation of the software workflow, not a Jev implementation. The fixture lookup stands in for model results, and the branch belongs to code. Its assertion shows that adding urgency to the batch leaves the route and duplicate-charge fixture values unchanged.

# Deterministic fixture; no Jev API call.
STATE = {"ticket_id": "A-104", "text": "Wrong size and charged twice"}
ANSWERS = {
    "route": {"choice": "returns",
              "probabilities": {"returns": .60, "billing": .38, "shipping": .02}},
    "urgency": {"score": 1.30},
    "duplicate": {"noul": .80},
}


def ask_fixture(state, question_ids):
    assert state["ticket_id"] == "A-104"
    result = {}
    for question_id in question_ids:
        if question_id not in ANSWERS:
            raise KeyError(question_id)
        result[question_id] = ANSWERS[question_id]
    return result


def choose_workflow_step(answers):
    if answers.get("duplicate", {}).get("noul", 0.0) >= .75:
        return "billing_review"  # Local policy, not a model action.
    if "route" in answers:
        return answers["route"]["choice"]
    return "request_more_evidence"


pair = ask_fixture(STATE, ["route", "duplicate"])
triple = ask_fixture(STATE, ["route", "urgency", "duplicate"])
assert pair["route"] == triple["route"]
assert pair["duplicate"] == triple["duplicate"]
assert choose_workflow_step(triple) == "billing_review"
print(choose_workflow_step(triple))

The short Python equivalent for gathering known answer IDs is {key: ANSWERS[key] for key in question_ids}; it returns the same mapping. The actual SDK sends typed questions with one state, rather than this lookup table. We keep the lookup local so the teaching flow is reproducible without a credential, billing, or an invented claim about Jev's particular answer to A-104. TypeSafe multi-question request example

Useful boundary: The model supplies judgments; code owns the deterministic branch and any side effect. That boundary also lets a developer log which question and policy version led to a route. It does not certify that the probability inputs are calibrated for this support queue.

Our program now has a route probability distribution and a branch, but one more number appears in Choice and Score responses: confidence. The next chapter separates an option probability, the selected option, and that distribution-derived summary.

When does a second request become necessary for A-104?

Chapter 3: Probability is not confidence

Our A-104 fixture gives Returns .60, Billing .38, and Shipping .02. Why does the program pick Returns? Because .60 is the largest of the three modeled option probabilities. But the customer also has a duplicate-charge issue, so the second option is not negligible. We need to see the whole distribution, not just the selected label.

For a TypeSafe Choice answer, probabilities lists masses over the caller's options, and choice is the highest-probability one. The probabilities sum to one across that supplied list. A separate confidence number summarizes how concentrated the distribution is; the cited documentation does not provide a formula we should reproduce. It also does not make a high value a guarantee that the chosen route is right. TypeSafe Choice — response structure; TypeSafe Confidence

Keep three readings apart
Top option mass: the probability assigned to Returns in our local fixture.
Selected option: the option with the largest mass.
Top-two margin: our own illustrative statistic, not TypeSafe's confidence.

Drag a bar or use the sliders. The display starts at .60, .38, .02 and always renormalizes to a total of one. Moving one option changes the other two proportionally, a rule of this teaching widget only. The adjacent readout names the top option and calculates our custom margin; it never presents a computed Jev confidence.

Explore a bounded route distribution

Local synthetic A-104 fixture. Drag a bar horizontally, adjust the accessible sliders, or select a preset. The human-authored primary route label remains Billing.

Returns.60
Billing.38

At the starting distribution, the top-two gap is .60 − .38 = .22. The selected option's own mass remains .60. At the peaked preset, the gap becomes .90 − .08 = .82. A larger gap can be useful as a custom warning signal, but it is not TypeSafe's reported confidence and cannot by itself tell us the route is correct under A-104's policy.

Try “Keep top .60, spread rest.” It sets Returns to .60, Billing to .20, and Shipping to .20. The selected option and its top probability have not changed, yet the runner-up shrinks and our margin becomes .60 − .20 = .40. This is a concrete reason to keep the full distribution: a single top number hides where the remaining mass went. It still cannot reveal TypeSafe's undisclosed confidence calculation.

our top-two margin = largest option probability − second-largest option probability
starting fixture: .60 − .38 = .22; peaked fixture: .90 − .08 = .82

The difference matters in practice. If code mistakes a confidence scalar for the probability of a particular event, it may compute the wrong expected cost. TypeSafe documents confidence as a summary derived from the distribution shape, and says thresholds should be adjusted for the risk and the task. A model can express a concentrated distribution and still choose the wrong allowed answer. TypeSafe Confidence — thresholds scale with risk

The documentation includes a Choice example with top route probability .61 and returned confidence .42. Those two values differ because they answer different questions; the .42 is a documented example output, not a number calculated by this widget or a result from A-104. We deliberately keep that evidence apart from the local editable bars. TypeSafe Choice — multi-issue ticket example

Documented example, separate from the simulation: one TypeSafe Choice response has a highest option probability of .61 and a separate confidence of .42. The API page does not publish a formula that maps an arbitrary bar set to that number. Do not apply .42 to the bars after you drag them.
Common mistake: 1 − confidence is not a documented error probability. Nor is our top-two margin the vendor confidence field. For a decision with real consequences, test probability calibration and the action's loss on labeled cases rather than treating any one summary as a safety certificate.

Run this plain Python example to see how a selected option and a custom top-two gap are computed. The input weights are supplied locally. The function makes no Jev call, and the result has no fabricated Jev confidence field.

# Local distribution arithmetic only; not Jev's confidence formula.
OPTIONS = ["returns", "billing", "shipping"]


def normalize(weights):
    assert len(weights) == len(OPTIONS)
    assert all(weight >= 0 for weight in weights)
    total = sum(weights)
    if total <= 0:
        raise ValueError("At least one weight must be positive")
    return [weight / total for weight in weights]


def selected_and_margin(weights):
    probabilities = normalize(weights)
    ranked = sorted(range(len(probabilities)),
                    key=lambda i: probabilities[i], reverse=True)
    first, second = ranked[:2]
    margin = probabilities[first] - probabilities[second]
    return OPTIONS[first], probabilities, margin


name, probabilities, margin = selected_and_margin([60, 38, 2])
print(name, probabilities, round(margin, 2))
assert name == "returns"
assert abs(sum(probabilities) - 1.0) < 1e-12
assert abs(margin - .22) < 1e-12

peaked = selected_and_margin([90, 8, 2])
split = selected_and_margin([60, 20, 20])
assert abs(peaked[2] - .82) < 1e-12
assert split[0] == name and abs(split[2] - .40) < 1e-12

A short standard-library equivalent for finding the selected option is max(zip(OPTIONS, probabilities), key=lambda pair: pair[1]); it returns the same Returns/.60 pair for the initial fixture. The longer version shows the extra step needed for the runner-up and margin. Neither computes TypeSafe's undisclosed confidence statistic.

Chapter 4 asks the empirical question we cannot answer from one ticket: when many cases receive a probability near .80, how often does the event actually occur? That is how we begin testing calibration instead of guessing correctness from the shape of one distribution.

The A-104 local fixture starts at Returns .60, Billing .38, Shipping .02. Which statement is warranted?

Chapter 4: Test calibration on groups

Suppose a future routing ticket receives a forecast of 0.8 for the binary event “its selected route is policy-correct.” That number is a forecast. We cannot mark it calibrated from one ticket, even after the ticket is resolved: the selected route is either correct or not. Calibration is a pattern across comparable forecasts and independently resolved outcomes.

For a deliberately synthetic audit of that same event, collect 100 tickets, each forecast at p = 0.8 that its selected route will be policy-correct. If 80 of those events occur, the observed fraction in that bucket is 80/100 = 0.8. If only 60 occur, the fraction is 0.6 and the bucket misses the diagonal by 0.2. This exercise does not measure Jev; it shows how to audit any probability-bearing workflow.

A-104 remains the motivating case: its illustrative three-route fixture gives Returns .60, Billing .38, Shipping .02, and its policy label is Billing. The 100-ticket binary audit below is a separate constructed sample of whether selected routes are policy-correct; it does not silently replace A-104’s .60/.38/.02 masses. To audit such an event, compare each selected-route correctness forecast with an independently resolved policy label. Do not use the model’s own answer as the ground truth.

Synthetic audit fixture · v1
100 comparable route-correctness forecasts at p = 0.8. Drag the count of resolved positives from 60 to 80. The Brier score is the mean squared distance between each forecast and its 0-or-1 outcome.
A reliability point and a proper score

The Brier score for a binary event is the average of (forecast − outcome)². With 80 positives, each positive contributes (0.8−1)² = 0.04 and each negative contributes (0.8−0)² = 0.64. The full mean is (80×0.04 + 20×0.64)/100 = 0.16. At 60 positives and 40 negatives it becomes (60×0.04 + 40×0.64)/100 = 0.28. Smaller is better for this score on the same labeled cases. Brier's original probability-verification paper

Brier(80/20) = (80·0.04 + 20·0.64)/100 = 0.16
Brier(60/40) = (60·0.04 + 40·0.64)/100 = 0.28

Switch on the subgroup. It isolates ten illustrative tickets with 4 positives, or 0.4 observed frequency, while the overall 80/100 group still looks aligned. At the 80-positive slider setting, the other 90 contain 76 positives; at 60, they contain 56. Thus the remaining-group count is always the slider total minus 4. A single aggregate reliability point can hide variation that matters to a deployment population. Our subgroup is hand-constructed, not discovered in customer data.

Where did the aggregate score come from?

Keep the same slider total. This table partitions the 100 tickets into a constructed ten-ticket subgroup and the remaining ninety. Each score is recomputed from its own positive and negative counts; the weighted average must recover the whole-group Brier score.

PartPositiveNegativeFrequencyBrier
Small group
Other 90
All 100

At 80 positives overall, the small group has 4/10 correct routes at forecast .8, while the other group has 76/90. Its Brier score is (4×.04 + 6×.64)/10 = .40; the other group's score is (76×.04 + 14×.64)/90 ≈ .1333. Weight by group size: .10×.40 + .90×.1333 ≈ .16. The aggregate number can conceal a weak slice, even though its arithmetic is valid.

Even 80/100 is a noisy finite sample, not proof that all 0.8 forecasts are right 80% of the time. An audit needs enough independent resolved examples, several probability buckets, relevant segments and later time periods. It also needs a stable label definition: calibration to a weak proxy label cannot by itself establish that the route helps the customer.

Calibration and discrimination answer different questions. A system that always predicts a base rate can sometimes be calibrated, yet be poor at separating hard cases from easy ones. Brier loss provides one useful summary, while the reliability plot exposes where predicted and observed frequencies differ. For action decisions, we must also specify what different mistakes cost.

Common mistake: “This ticket has a confident-looking 0.99, so the model is calibrated.” A single forecast cannot establish a group-frequency property. Nor does TypeSafe's reported confidence scalar mean the chance that an action succeeds. TypeSafe confidence documentation

Run this complete Python example. The assertion checks the hand arithmetic; NumPy is an optional independent comparison on the identical fixture, not a second dataset or a Jev measurement.

# Standard-library calculation on synthetic resolved cases.
def brier(probabilities, outcomes):
    assert len(probabilities) == len(outcomes) and outcomes
    assert all(0 <= p <= 1 for p in probabilities)
    assert all(y in (0, 1) for y in outcomes)
    return sum((p-y)**2 for p, y in zip(probabilities, outcomes))/len(outcomes)

def frequency(outcomes):
    return sum(outcomes)/len(outcomes)

probabilities = [0.8]*100
aligned = [1]*80 + [0]*20
shifted = [1]*60 + [0]*40
assert abs(frequency(aligned)-0.8) < 1e-12
assert abs(brier(probabilities, aligned)-0.16) < 1e-12
assert abs(brier(probabilities, shifted)-0.28) < 1e-12
print(frequency(aligned), brier(probabilities, aligned))
print(frequency(shifted), brier(probabilities, shifted))
small = [1]*4 + [0]*6
remaining = [1]*76 + [0]*14
assert len(small) + len(remaining) == 100
assert sum(small) + sum(remaining) == 80
small_score = brier([0.8]*10, small)
remaining_score = brier([0.8]*90, remaining)
assert abs(small_score-0.40) < 1e-12
assert abs(0.1*small_score + 0.9*remaining_score - 0.16) < 1e-12
print("partition", small_score, remaining_score)
try:
    import numpy as np
except ImportError:
    print("NumPy comparison skipped")
else:
    print(float(np.mean((np.array(probabilities)-np.array(shifted))**2)))

In this fixture, the labels are supplied by the author. In practice, label collection, policy version, unresolved tickets, and changing traffic all affect what “observed frequency” means. Keep the forecast and outcome timestamps, then evaluate only cases whose truth has been established without copying the prediction.

Next: A calibrated probability is an ingredient, not an action. Chapter 5 assigns costs to a wrong route and to review, then lets application code choose the lower expected loss.
Among 100 synthetic tickets all forecast at p = 0.8, only 60 resolve positive. What is the strongest justified conclusion?

Chapter 5: A probability becomes a decision only after we name a loss

A-104 requests both a shoe exchange and a duplicate-charge refund. In our synthetic policy, Billing is the primary owner because duplicate-charge review takes priority. A bounded route prediction can help the application decide, but the prediction is not itself the routing policy. The developer still chooses which errors matter and what to do when the evidence is weak.

Our local teaching fixture assigns Returns 0.60, Billing 0.38, and Shipping 0.02. Those are invented route masses for A-104; they are not a Jev call or a measured customer distribution. The allowed actions are route to one of those teams or send the ticket for review. A wrong automatic route costs 10 units, a correct route costs zero, and perfect review costs 2 units. These costs encode a preference, not a measured dollar amount.

Synthetic A-104 decision · fixture v1
Route masses: Returns .60 · Billing .38 · Shipping .02. Wrong route loss 10; correct route loss 0. Review cost 2 and, in this simplified model, review always finds the right route. The actual synthetic policy label remains Billing.

Expected loss means averaging over possible outcomes using their probabilities. If we automatically choose Returns, its correct mass is .60 and wrong mass is .38+.02=.40. Thus loss = .60×0 + .40×10 = 4. For Billing, loss = (.60+.02)×10 = 6.2. For Shipping it is .98×10 = 9.8. Review's stipulated cost is 2. The minimum expected loss is review, despite Returns having the highest route mass.

L(route i) = (1 − pi) × wrong-route loss
L(review) = review cost
A-104: L(Returns)=4, L(Billing)=6.2, L(Shipping)=9.8, L(review)=2
Move the costs; watch the chosen action

Every bar is computed from the same three route masses. This is application-owned policy, not Jev output.

For a top route with probability p, the symmetric wrong-route rule gives automatic loss (1−p)C, where C is wrong-route cost. Perfect review wins when its cost R is lower. The break-even value is p = 1−R/C. With C=10 and R=2, that is 0.8: at p=.60 review wins; at p=.90, automatic loss 1 wins. The equality case is a tie. This threshold does not transfer to a different loss table or to a review process that sometimes makes errors.

The label “Billing” in our case file matters when we evaluate the eventual result. The probability-based policy can rationally request review because its own route distribution favors Returns while the author-supplied primary route truth is Billing. Reviewing the case in our toy reveals that truth. In a deployed system, a human queue might be slow, imperfect, or itself costly in ways our slider omits.

The action set also matters. With no review option, the minimum expected loss is the highest-mass route under equal wrong-route costs. With review available, it can be optimal to abstain from automatic routing. This is selective prediction under a stipulated cost model. It does not mean TypeSafe's scalar confidence has an inherent act/no-act cutoff: the cutoff belongs to the application and should be tested against labeled outcomes. TypeSafe on domain-dependent thresholds

The first comparison assumed review finds the correct team every time. That is a simplifying condition, not a property of real reviewers. Stress-test it with a second toy parameter: let a reviewer choose the wrong team with probability q, independently of the ticket type, while the base review cost still applies. Under the same symmetric wrong-route loss C, review's expected loss becomes R + qC. This extension remains application policy, not a Jev field.

What if review sometimes makes an error?

The upper chart keeps the perfect-review baseline. This sensitivity chart shares its wrong-route and base-review cost sliders, then adds a hypothetical reviewer-error rate. It compares review against the best automatic route.

At our baseline C=10 and R=2, the best automatic route is Returns with expected loss 4. Review with error rate q costs 2+10q. It is cheaper while q<.20, ties at q=.20, and becomes more costly beyond that point. The action can change without the route distribution changing at all. In practice, reviewer errors may depend on case type, and delay can matter separately; estimate those effects rather than trusting this one-parameter sensitivity.

review with error = R + qC; baseline example: 2 + 10q
compare Returns 4 ⇒ review wins for q < .20, ties at .20
Common mistake: “The model says Returns 0.60, so send it to Returns.” That omits the other route masses, the cost of being wrong, and the review alternative. Conversely, choosing review in this toy is not a production-safety guarantee; the perfect-review assumption must be checked.

Run the entire loss table below. The NumPy comparison performs the same dot products with the same probabilities and cost rows. It is optional; the standard-library result and assertions work without dependencies.

# Synthetic A-104 fixture; no Jev API request.
ROUTES = ("returns", "billing", "shipping")
prob = (0.60, 0.38, 0.02)
assert abs(sum(prob)-1.0) < 1e-12
wrong_cost, review_cost = 10.0, 2.0

def route_loss(route, wrong_cost):
    assert route in ROUTES
    return sum(p * (0 if truth == route else wrong_cost)
               for truth, p in zip(ROUTES, prob))

losses = {route: route_loss(route, wrong_cost) for route in ROUTES}
losses["review"] = review_cost  # Perfect review by assumption.
choice = min(losses, key=losses.get)
print(losses, choice)
assert choice == "review"
assert abs(losses["returns"]-4.0) < 1e-12
assert abs(losses["billing"]-6.2) < 1e-12
assert abs(losses["shipping"]-9.8) < 1e-12
assert abs((1-review_cost/wrong_cost)-0.8) < 1e-12

def imperfect_review_loss(error_probability):
    assert 0 <= error_probability <= 1
    return review_cost + error_probability*wrong_cost

assert imperfect_review_loss(0.0) == 2.0
assert imperfect_review_loss(0.2) == losses["returns"]
assert imperfect_review_loss(0.3) > losses["returns"]
try:
    import numpy as np
except ImportError:
    print("NumPy comparison skipped")
else:
    cost = np.array([[0,10,10],[10,0,10],[10,10,0]], dtype=float)
    print("NumPy route losses:", np.array(prob) @ cost.T)

This table assumes one true primary route and the same penalty for every wrong team. A real support system may need a full asymmetric loss matrix: missing a charge dispute might cost more than sending an exchange to Billing first. The method stays the same—multiply each possible outcome's loss by its probability, sum, and compare available actions—but the simple top-probability threshold disappears.

A-104's route distribution, local policy label, and costs are deliberately distinct objects. You can version each one separately. If the fee for review rises while the case evidence stays fixed, code may choose Returns; the model did not “change its mind.” This separation makes a workflow debuggable and sets up the next question: what would a new record do to the probability before we act?

A route has probability 0.90, wrong-route loss 10, and perfect review costs 2. Which action has lower expected loss?

Chapter 6: Update a belief with new evidence

A-104 mentions a possible duplicate charge. Suppose a payment record arrives before we route it. That record can change how plausible Billing is, but only if we know how often such a record appears in genuinely Billing-owned cases and in other cases. This chapter builds a tiny Bayesian application-side teaching model; it does not claim Jev internally runs Bayes' rule.

Start with a different illustrative scenario from chapter 5's route fixture. Before seeing the record, assign probability 0.30 to “Billing is the primary route” and 0.70 to “another team is primary.” This is a synthetic prior over two hypotheses, not Jev's Choice output and not the .38 Billing mass in our earlier three-route example. The values differ on purpose so we can see what evidence adds.

Assume the record is positive 80% of the time when Billing truly owns the case, but 10% of the time when another team owns it. These are hypothetical likelihoods from a stipulated observation model, not TypeSafe benchmark statistics. The positive record has likelihood ratio 0.8/0.1 = 8 in favor of Billing.

A-104 evidence exercise · separate synthetic prior
Prior Billing .30; other .70. P(positive payment record | Billing)=.80; P(positive record | other)=.10. A single record is observed. No Jev API or hidden Jev update is being represented.
One observation, three named stages

Step through prior → joint weights → normalized posterior. The duplicate button copies the same record, so it must not create another likelihood update.

Multiply each prior mass by the probability of seeing this positive record under that hypothesis. Billing gets 0.30×0.80 = 0.24; other gets 0.70×0.10 = 0.07. These are joint weights, not posterior probabilities yet: they sum to 0.31, the probability of seeing a positive record in this toy world.

P(Billing | +) = [P(+ | Billing) × P(Billing)] / P(+)
= (0.80×0.30) / (0.80×0.30 + 0.10×0.70)
= 0.24 / 0.31 ≈ 0.77419

Normalize both weights by 0.31. Billing becomes 0.24/0.31 ≈ 0.774, and other becomes 0.07/0.31 ≈ 0.226; together they are 1. This is a posterior belief under the stated prior and likelihoods. It is not a guarantee that Billing is the correct team for this individual ticket.

Now tap “Duplicate same record.” A database copy, a second UI rendering, or a paraphrase of the same evidence is not a second independent observation. If we multiply by the same likelihood ratio again, we falsely get approximately 0.965 for Billing: prior odds 3/7 times 8² become 192/7, giving 192/199 ≈ 0.965. That impressive-looking rise is an error. Only a distinct new measurement, with a justified conditional likelihood given what we already saw, may be used for another update.

Can we decide whether the payment-record query itself is worth paying for? Keep the same prior and likelihoods, now add a decision: route to Billing, route to the other team, or use perfect review. Assume a wrong automatic route costs 10, a correct route costs 0, and perfect review costs 2. Before seeing the record, routing to Other costs 0.30×10 = 3, routing to Billing costs 0.70×10 = 7, and review costs 2. The baseline choice is review at expected loss 2.

Would you pay to acquire the record?

This is a separate application-side value-of-information toy, using the same prior and test likelihoods. Move the information cost; the conditional actions and weighted risk are recomputed.

A positive record occurs with probability 0.30×0.80 + 0.70×0.10 = 0.31. Its Billing posterior is 24/31. Automatic Billing would risk 10×7/31, about 2.26, so review at 2 is cheaper. A negative record occurs with probability 0.69; its Billing posterior is (0.30×0.20)/0.69 = 2/23. Route to Other then risks 10×2/23 = 20/23, about 0.87, below review's 2.

Expected loss after observing = 0.31×2 + 0.69×(20/23) = 1.22
With record cost 0.20: total 1.42; baseline 2.00; net value 0.58

The gross expected value of this sample information is 2−1.22 = 0.78. After a query cost of 0.20, the net expected benefit is 0.58. Under these toy assumptions, acquire the record only when its incremental expected benefit exceeds its cost: below 0.78, query; above 0.78, skip; at 0.78, tie. The query does not always make us auto-route: a positive still leads to review. This contingent policy is why averaging the two branches matters. MIT's expected value of sample information lesson

These numbers assume the record's sensitivity and false-positive rate are valid on this ticket population, the true route does not change while we wait, review is perfect, and query cost includes every delay or side effect. If any assumption fails, recompute the tree. This is ordinary decision-analysis logic around a hypothetical signal, not a documented Jev feature or a claim that Jev exposes value of information.

The word “independent” here has a precise condition. Two records may share the same bank event, so they are correlated even if they arrive through different services. To multiply two likelihood factors of the simple form, we would need the observations to be conditionally independent given the true route. Application code has to decide how to model dependence; TypeSafe's documented parallel evaluation of questions does not supply that statistical assumption. TypeSafe primitives and independent question evaluation

Common mistake: “I asked the same question twice and received the same answer, so my odds doubled.” Repeated presentation of the same evidence carries no new information. A high posterior from a wrong dependence model is not a better estimate.

This full Python example uses exact fractions for both posterior branches and every action risk. Its assertions verify the hand calculation, including the optional query value. NumPy, if installed, checks the same weighted branch losses; no library or Jev call is needed.

# External Bayes and value-of-information analogy; no Jev call.
from fractions import Fraction as F

prior = F(3, 10)             # P(Billing)
sensitivity = F(4, 5)       # P(+ | Billing)
false_positive = F(1, 10)   # P(+ | Other)
wrong_loss = F(10)
review_cost = F(2)
query_cost = F(1, 5)

def minimum_risk(p_billing):
    risks = {"billing": (1-p_billing)*wrong_loss,
             "other": p_billing*wrong_loss,
             "review": review_cost}
    action = min(risks, key=risks.get)
    return action, risks[action]

p_positive = prior*sensitivity + (1-prior)*false_positive
p_negative = 1-p_positive
assert p_positive and p_negative  # Both observed branches possible.
post_positive = prior*sensitivity/p_positive
post_negative = prior*(1-sensitivity)/p_negative
baseline_action, baseline = minimum_risk(prior)
pos_action, pos_risk = minimum_risk(post_positive)
neg_action, neg_risk = minimum_risk(post_negative)
post_observation = p_positive*pos_risk + p_negative*neg_risk
gross_value = baseline-post_observation
net_value = gross_value-query_cost
print("baseline", baseline_action, baseline)
print("positive", p_positive, post_positive, pos_action, pos_risk)
print("negative", p_negative, post_negative, neg_action, neg_risk)
print("post-observation", post_observation,
      "with query", post_observation+query_cost,
      "net benefit", net_value)
assert (baseline_action, baseline) == ("review", F(2))
assert (p_positive, post_positive) == (F(31,100), F(24,31))
assert (p_negative, post_negative) == (F(69,100), F(2,23))
assert (pos_action, pos_risk) == ("review", F(2))
assert (neg_action, neg_risk) == ("other", F(20,23))
assert (post_observation, gross_value, net_value) == (
    F(61,50), F(39,50), F(29,50))
# Repeating the *same* record adds no new evidence.
assert post_positive == F(24,31)
try:
    import numpy as np
except ImportError:
    print("NumPy comparison skipped")
else:
    branch = np.array([float(p_positive), float(p_negative)])
    risk = np.array([float(pos_risk), float(neg_risk)])
    assert abs(float(branch @ risk)-1.22) < 1e-12
    print("NumPy post-observation risk", float(branch @ risk))

There is a useful connection to Bayes estimation and Bayes filtering: both distinguish a prior from evidence and a posterior. This static ticket example has no evolving state transition, so it is not a Bayes filter or a Kalman filter. The connection is the update structure, not a claim that Jev uses those algorithms.

Our code can use a posterior to inform a decision rule, but it must still define the loss of a wrong route or a review. The posterior is an estimate; the action belongs to the application. Next we examine TypeSafe's statement that Jev is trained for calibrated decisions, keeping the public claim separate from an external proper-scoring illustration.

The same positive payment record is copied into a second field. What should this Bayesian toy do?

Chapter 7: What RLCD says, and what it leaves undisclosed

TypeSafe calls its training approach Reinforcement Learning for Calibrated Decisions, or RLCD. Its public primer says the aim is decisions and probabilities rather than generated text, with a well-calibrated 0.8 occurring about 80% of the time across comparable cases. That is a product-level objective, not a published training algorithm. TypeSafe's AI primer

Our A-104 route example can make the idea concrete. Imagine many resolved support tickets where a precisely defined binary event, “Billing is the policy-correct primary route,” occurs on 80% of cases in a comparable group. A forecast of 0.8 for that event is better aligned with the group's frequency than a reflexive 0.5. This sentence describes an external synthetic dataset, not Jev's training data or a Jev output for A-104.

A proper scoring rule is one mathematical way to reward honest probabilistic forecasts. The binary Brier loss is (p−y)², where p is a forecast and y is the resolved 0-or-1 event. We use this familiar rule as a teaching analogy only. TypeSafe does not disclose that RLCD uses Brier loss, nor its exact reward, optimizer, dataset, model architecture, or sampling procedure. Founder's RLCD description

Boundary of the diagram
Documented: TypeSafe says RLCD targets calibrated decisions and probabilities. Constructed here: a Brier-score teaching curve on a binary event with true group frequency 0.8. The curve is not a Jev loss curve, training log, or reward specification.
Choose a forecast, then reveal the outcome mix

The curve shows expected toy Brier loss under a fixed synthetic 80/20 event mix. Move the slider or reveal individual outcomes; the underlying 80/20 group stays fixed.

At forecast p=0.8, the 80 positive cases each cost (0.8−1)²=0.04, and the 20 negatives each cost (0.8−0)²=0.64. Average over the group: 0.8×0.04 + 0.2×0.64 = 0.16. A p=0.5 forecast costs 0.25 on either outcome, so its group average is 0.25. These are computed on identical outcomes; lower loss is better.

E[Brier at p] = 0.8(p−1)² + 0.2p²
E[Brier at 0.8] = 0.16; E[Brier at 0.5] = 0.25
Expected loss versus the revealed sample

Reveal outcomes above, then compare three forecasts on exactly that observed prefix. The sample column is a realized mean, while the smooth curve uses the stipulated eighty-twenty population. At zero revealed cases, a realized score does not exist.

ForecastExpected on 80/20 Observed prefix
Current slider
0.80
0.50

A forecast of .99 is closer than .8 when the next outcome is positive, but it is heavily penalized on a negative outcome: (.99−0)²=.9801. Across the stipulated 80/20 group, its expected Brier is .8×.0001 + .2×.9801 = .1961, above the honest .16. This contrast is about repeated outcomes; it does not say which forecast wins on every single case.

The minimum occurs at the group's true event frequency. Expanding the curve gives p²−1.6p+0.8, whose derivative is 2p−1.6; setting it to zero yields p=0.8. Or complete the square: (p−0.8)²+0.16. This explains why the illustrative score favors an honest probability over a fixed bluff across repeated events. It does not establish how TypeSafe optimizes Jev.

The same reasoning does not depend on eighty percent. If the true group frequency is q, expected binary Brier loss is q(p−1)² + (1−q)p² = (p−q)² + q(1−q). The square is smallest at p=q. For a different constructed group with q=.3, an honest p=.3 gives expected loss .21, while p=.8 gives .46. This is a property of the scoring rule under a correctly defined repeated event, not evidence about TypeSafe's undisclosed training reward.

At q=.3: E[Brier at p=.3]=.21; E[Brier at p=.8]=.3×.04 + .7×.64=.46

Revealing a few outcomes can produce a noisy short-run average. Even if eight in ten is the long-run constructed frequency, the first three outcomes need not contain exactly 2.4 positives. The chart's expected curve uses the stipulated group distribution; the reveal control separately displays realized losses on a deterministic 10-outcome sequence. Keep those two quantities distinct.

“Reinforcement learning” in RLCD does not license drawing a specific PPO clip objective, DQN target network, actor-critic rollout, or customer-specific online update. The public pages identify a training goal, not enough implementation detail to reproduce the recipe. TypeSafe also documents that its standard model is not fine-tuned on each customer's requests. The application can still learn about its workflow by logging outcomes and retuning its own policy. TypeSafe models and customization

Common mistake: “This Brier curve must be Jev's RLCD objective.” The curve is ours. It illustrates a reason calibrated probabilities are useful while the vendor's actual reward and optimizer remain undisclosed.

The executable program below enumerates forecasts from 0 to 1, computes their expected binary Brier loss, and confirms that 0.8 wins on this constructed event mix. NumPy, when installed, evaluates the same grid; no Jev inference is made.

# Synthetic proper-scoring demonstration, not Jev RLCD code.
def expected_brier(p, event_frequency=0.8):
    assert 0 <= p <= 1 and 0 <= event_frequency <= 1
    return event_frequency*(p-1)**2 + (1-event_frequency)*p**2

grid = [i/100 for i in range(101)]
losses = [expected_brier(p) for p in grid]
best = grid[min(range(len(grid)), key=lambda i: losses[i])]
print(best, expected_brier(0.8), expected_brier(0.5))
assert best == 0.8
assert abs(expected_brier(0.8)-0.16) < 1e-12
assert abs(expected_brier(0.5)-0.25) < 1e-12
assert abs(expected_brier(0.99)-0.1961) < 1e-12
assert abs(expected_brier(0.3, 0.3)-0.21) < 1e-12
assert abs(expected_brier(0.8, 0.3)-0.46) < 1e-12
outcomes = [1,1,0,1,1,1,0,1,1,1]  # 8 positive, 2 negative.
assert sum(outcomes) == 8
for count in (1, 3, 10):
    prefix = outcomes[:count]
    realized = sum((0.8-y)**2 for y in prefix)/count
    print("prefix", count, "realized Brier at 0.8", realized)
assert abs(sum((0.8-y)**2 for y in outcomes)/10-0.16) < 1e-12
try:
    import numpy as np
except ImportError:
    print("NumPy comparison skipped")
else:
    p = np.linspace(0, 1, 101)
    np_losses = .8*(p-1)**2 + .2*p**2
    assert abs(float(p[np.argmin(np_losses)])-best) < 1e-12
    print(float(np.min(np_losses)))

The event definition is as important as the score. Predicting a proxy label well does not prove that the system reduced customer harm. Calibration on one group does not guarantee subgroup calibration or that a future shifted population behaves the same. This is why the next chapter separates a classifier's event probability from a long-term action value in a sequential process.

Which statement is supported by the public material and this chapter's boundary?

Chapter 8: A class probability is not a future return

A-104 still needs a primary owner. A typed Choice response can provide a probability distribution over Returns, Billing, and Shipping. The application can use that distribution in a routing rule. But a probability attached to a label does not say what will happen to tomorrow's support queue after today's action.

Keep our local synthetic fixture separate from the API: Returns .60, Billing .38, Shipping .02. Those three numbers sum to one and describe candidate labels for this one ticket. They are not rewards, action values, or measured Jev outputs for A-104. The case's local policy truth remains Billing because duplicate-charge review takes priority.

A sequential decision needs different ingredients. We must specify actions, a reward after each action, a state transition, and a policy for what happens next. In reinforcement learning, a discounted two-step return is G = r₀ + γr₁. The discount γ controls the weight of the later reward; it does not change the Choice probabilities. See Decision Making Under Uncertainty Workbook — MDPs & Bellman Backup and Deep Reinforcement Learning — MDPs and value functions.

Constructed two-step queue exercise. Action A earns 1 now and 0 next step. Action B earns 0 now and 5 next step. These rewards are invented to teach time horizon; neither action is a real A-104 routing result.
Move the discount; watch the preferred action change

Work the default calculation by hand. At γ = .9, action A gives 1 + .9 × 0 = 1. Action B gives 0 + .9 × 5 = 4.5. B therefore has the larger two-step return, even though A pays more immediately. If γ = 0, the comparison reverses: A gives 1 and B gives 0. The slider recalculates these returns from the two reward pairs; the route-probability card stays outside that calculation.

G(A) = 1 + γ × 0 = 1;   G(B) = 0 + γ × 5 = 5γ

The diagram's B reward of 5 is deterministic by construction. Real future rewards may be uncertain. In a separate stochastic variant, B pays 5 one step later with probability s, otherwise 0. At s=.5 and γ=.9, expected B return is 0 + .9 × (.5×5 + .5×0) = 2.25. At s=.2, it is .9×1 = .9, below A's return of 1. Move the success slider to see the crossover. This expectation requires an action-conditioned outcome estimate; a Choice label mass is not that estimate.

E[G(B)] = γ × (s × 5 + (1 − s) × 0) = 5γs

A contextual bandit is a one-step relative: it sees a context, chooses an action, and observes that action's outcome. A sequential MDP additionally models the next state and future actions. Merely seeing a classifier's probability vector supplies neither action-conditioned rewards nor transitions. To estimate a true Q(state, action), a learner would need outcome data under relevant actions and a specified objective. This exercise is an external decision analogy, not Jev's undisclosed training procedure.

Common mistake: P(Billing)=.38 is a belief about a supplied answer option. Q(state, Billing) would be an expected future return under a reward and policy. The numbers share neither units nor meaning, even if both happen to lie between zero and one.

Try the second control below the timeline. It computes the same single-row inverse-propensity contribution while you change the logged action propensity. When q gets smaller but stays positive, the weight grows. If that action had no chance of being logged at all, the “no support” switch correctly refuses to compute a weight. Zero support is missing evidence, not infinite evidence.

Logged outcomes bring a second trap: a policy may almost never try an action in some context, so its alternative outcome is missing. In a separate one-step bandit example, suppose 100 logged decisions contain one target-policy action with observed reward 12 and logging propensity .2. Its contribution to an inverse-propensity estimate is (12 ÷ .2) ÷ 100 = .6. That correction is meaningful only when the logging propensity is known, positive wherever the target policy acts, and the logged context/action/reward assumptions fit the target. It is not a Jev metric; Bandits & Preference-Based Learning — Contextual Bandits develops the related setting.

Run the short program from scratch. It uses only Python's standard library. The generator-expression and sum line provide an equivalent dot-product calculation for the same two rewards, without introducing a library's unrelated learning algorithm.

# Constructed two-step rewards; not a Jev call.
rewards = {"A": (1.0, 0.0), "B": (0.0, 5.0)}

def two_step_return(pair, gamma):
    assert 0 <= gamma <= 1
    return pair[0] + gamma * pair[1]

def same_return_by_sum(pair, gamma):
    return sum(r * weight for r, weight in zip(pair, (1.0, gamma)))

gamma = 0.9
values = {name: two_step_return(pair, gamma)
          for name, pair in rewards.items()}
assert values == {"A": 1.0, "B": 4.5}
assert all(values[name] == same_return_by_sum(pair, gamma)
           for name, pair in rewards.items())
assert max(values, key=values.get) == "B"

def stochastic_b_expected(gamma, success_probability):
    assert 0 <= success_probability <= 1
    outcomes = ((5.0, success_probability),
                (0.0, 1 - success_probability))
    expected_later = sum(reward * probability
                         for reward, probability in outcomes)
    return gamma * expected_later

assert stochastic_b_expected(.9, .5) == 2.25
assert abs(stochastic_b_expected(.9, .2) - .9) < 1e-12
assert stochastic_b_expected(.9, .2) < values["A"]
assert two_step_return(rewards["A"], 0) > two_step_return(rewards["B"], 0)
print(values, "preferred:", max(values, key=values.get))

The choice between A and B belongs to this stipulated reward model. The next chapter will return to A-104 and put the typed-fixture judgment, a cost-based policy, and the labeled outcome into one controllable workflow. That step makes it possible to see which errors belong to the model, the policy, or the data.

Which value describes a future, reward-weighted two-step consequence in this exercise?

Chapter 9: Put the decision into a test harness

Now join the lesson's pieces in a small, inspectable workflow. We will take a ticket state, ask a bounded routing question, feed a synthetic local fixture into code, choose auto-route or review, and compare the resulting action with a labeled outcome. No button on this page calls Jev or processes a real customer's ticket.

Our recurring A-104 message asks for a shoe exchange and a duplicate-charge refund. Its fixture gives Returns .60, Billing .38, Shipping .02; our stated local policy says Billing is primary. The response shape has three permitted choices. The policy that follows is authored here in code, not learned automatically from the shape. TypeSafe documents Choice as a bounded answer space, while the application owns downstream behavior. TypeSafe — Choice primitive

The harness has six stages, shown one at a time so a phone can display the complete current object. Step through them or press Play. Change the ticket, costs, or the fixture's shifted-prediction switch and repeat the run. Every figure in the summary is recalculated from the selected fixture and controls. The time slider does not hide any paid model request; the “fixture answer” stage simply reads an array stored locally.

A-104 decision lab · constructed data and policy

Compute A-104 by hand at the initial settings. The highest local fixture mass is Returns .60, leaving .38 + .02 = .40 mass on other labels. Under a deliberately simple symmetric-loss model, auto-routing to Returns has estimated loss .40 × 10 = 4. Perfect human review is stipulated to cost 2 and to route correctly, so the policy selects review. Its realized synthetic outcome is a Billing assignment at cost 2. The fixture's .60 is not a verified chance of correctness; the estimated loss is conditional on treating these numbers as useful probabilities.

auto estimated loss = (1 − p(top route)) × wrong-route cost;   review estimated loss = review cost

Turn on the shifted fixture. For A-104 it now favors Returns .84, Billing .12, Shipping .04. At the same costs, auto-route's estimated loss drops to (.12 + .04) × 10 = 1.6, so the code may auto-route to Returns. Yet the policy label remains Billing, giving a realized wrong-route cost of 10. This is a controlled failure: a confident, misspecified fixture can make a mathematically consistent rule choose badly. Moving the review-cost slider can produce a different decision, and the stage panel explains which object changed.

The outcome label is available to this teaching harness because we authored it; an operational system often learns it later from a resolved ticket or reviewer. Human review is assumed perfect here only to isolate the loss comparison. If reviewers make mistakes, their error rate and delays must enter the policy model. A forced binary auto/review switch also omits valid operational choices such as requesting more evidence or escalating a safety-sensitive case.

Common mistake: A polished stepper does not validate the underlying fixture probabilities, local route policy, or review assumption. The current screen demonstrates how a workflow could be wired; it is not a Jev performance test or a guarantee that a top-ranked answer is right.

Run the same pipeline without a browser. The code exposes three pure functions so you can change a prediction, a cost, or a label and see which stage changes. The assertion for shifted A-104 is intentionally a wrong-control case: the expected calculation says auto, while the labeled outcome says that auto route was wrong.

# Synthetic, deterministic harness. No SDK or network request.
CASES = {
    "A-104": {"truth": "billing", "base": (.60, .38, .02),
              "shift": (.84, .12, .04)},
    "A-106": {"truth": "returns", "base": (.82, .12, .06),
              "shift": (.45, .48, .07)},
    "A-107": {"truth": "shipping", "base": (.14, .12, .74),
              "shift": (.38, .44, .18)},
}
ROUTES = ("returns", "billing", "shipping")

def classify_fixture(case_id, shifted=False):
    case = CASES[case_id]
    probs = case["shift" if shifted else "base"]
    assert abs(sum(probs) - 1) < 1e-12
    return dict(zip(ROUTES, probs))

def choose_action(probs, wrong_cost=10, review_cost=2):
    route = max(ROUTES, key=lambda name: probs[name])
    auto_estimate = (1 - probs[route]) * wrong_cost
    if auto_estimate < review_cost:
        return {"kind": "auto", "route": route,
                "estimate": auto_estimate}
    return {"kind": "review", "route": None,
            "estimate": review_cost}

def evaluate_outcome(case_id, action, wrong_cost=10, review_cost=2):
    truth = CASES[case_id]["truth"]
    if action["kind"] == "review":
        return {"assigned": truth, "correct": True, "cost": review_cost}
    correct = action["route"] == truth
    return {"assigned": action["route"], "correct": correct,
            "cost": 0 if correct else wrong_cost}

base = choose_action(classify_fixture("A-104"))
shifted = choose_action(classify_fixture("A-104", shifted=True))
assert base == {"kind": "review", "route": None, "estimate": 2}
assert evaluate_outcome("A-104", base)["cost"] == 2
assert shifted["kind"] == "auto" and shifted["route"] == "returns"
assert abs(shifted["estimate"] - 1.6) < 1e-12
assert evaluate_outcome("A-104", shifted) == {
    "assigned": "returns", "correct": False, "cost": 10}
print(base, shifted, evaluate_outcome("A-104", shifted))

A production harness can place a classifier at a focused decision point while deterministic code still owns the route/review policy. LangChain documents TypeSafeClassifier as a Runnable: configure named Choice criteria, call .invoke(state), and read response.choices["route"]. This shape-only integration sketch requires the package and an API key; it is not run by this lesson, and the fixture numbers above are not its output. LangChain — TypeSafe integration

# Integration sketch only: requires langchain-typesafe and API credentials.
from langchain_typesafe import Choice, TypeSafeClassifier
classifier = TypeSafeClassifier(questions={
    "route": Choice(
        instructions="Which team is the primary route?",
        criteria={"returns": "Exchange or return without billing priority.",
                  "billing": "Duplicate-charge review takes priority.",
                  "shipping": "Shipment delivery or tracking issue."},
    )
})
# response = classifier.invoke(ticket_text)
# probabilities = response.choices["route"].probabilities
# Application code then validates and applies its own route/review policy.

In this module, min over the two estimated losses is equivalent to the explicit if branch; the branch makes the ownership of the rule easier to see. A real integration could replace classify_fixture with a typed API response, but it would still need the same cost model, fallbacks, logging, and outcome evaluation. The next chapter asks what evidence from many labeled tickets would justify that integration.

For shifted A-104 at wrong-route cost 10 and review cost 2, what does this local policy do, and what does the labeled outcome reveal?

Chapter 10: Evaluate actions against resolved cases

The harness can be fast and tidy while sending a valuable ticket to the wrong team. To know whether its route policy helps, record predictions, decisions, and later resolved labels for many cases. Keep the label source explicit: a human-adjudicated policy target is different from another model's agreement with a prediction. The ten rows below are a constructed dataset, not a Jev benchmark.

The dataset is named jev-synthetic-routing-v1, the threshold rule top-mass-threshold-v1, and the probabilities local-hand-authored-probabilities-v1. Those are deliberately separate version IDs. Changing labels, the routing threshold, or a model could change a metric for different reasons; versioning lets us reconstruct which choice caused the change.

For each case, code takes the top route only when its fixture mass reaches the slider threshold. Otherwise it sends the ticket to review. Coverage is the fraction auto-routed. Acted accuracy counts correct assignments among auto-routed cases, not among all ten. A policy can inflate acted accuracy by reviewing almost everything, so show coverage and error counts together. The review channel is stipulated perfect at cost 2; a wrong auto-route costs 10.

Fixed 10-ticket evaluation · synthetic data

CaseTop route / massResolved labelProxy labelSliceAction

Check the default .65 result by hand against the full log. A-104 at .60 and A-110 at .50 go to review. The other eight are auto-routed. Of those, A-109 predicts Billing while its resolved label is Returns, and A-112 predicts Returns while its resolved label is Billing. Thus coverage = 8/10 = .80, acted accuracy = 6/8 = .75, and the count of wrong auto-routes is 2. Those are the dataset's actual counts; they are not a rounded story about an unshown experiment.

total constructed cost = 2 wrong × 10 + 2 reviewed × 2 = 24;   mean = 24 / 10 = 2.4

The shifted slice consists of A-109, A-110, and A-112. At .65, two of its three cases are auto-routed and both are wrong; the remaining one is reviewed. This small sample is not a reliable deployment estimate, but it exposes a failure that the overall 75% acted accuracy could conceal. A real evaluation should split by time, product, language, and other relevant conditions, with sufficiently many independently resolved examples.

The proxy column is not the resolved label. It is a synthetic stand-in for model consensus: for A-104 it says Returns while the policy label says Billing, and for A-109 it agrees with the wrong Billing prediction. Toggling the proxy shows disagreements rather than replacing ground truth in the main accuracy and cost metrics. TypeSafe's published workflow comparison uses reference probabilities averaged from GPT-6 Astra and Fable 5.1; those references are not human-adjudicated truth labels. Its 193.6× faster and 444.6× cheaper high-end workflow examples are vendor-reported comparisons under its benchmark setup, not universal latency, cost, or correctness guarantees. TypeSafe launch post — Workflow evals

To compare latency or cost fairly, fix the task schema, traffic mix, hardware/service settings, model identifiers, caching, retries, and quality target. TypeSafe documents model aliases; pinning an exact model version matters because an alias can move over time. Report uncertainty and error severity alongside throughput. A cheaper route that doubles damaging billing mistakes may be a poor operational trade. TypeSafe — Models and aliases

There are model-specific reasons to build these checks. TypeSafe's Jev 1.13 jaggedness guide says literal wording and contradictory instructions versus criteria can confuse a judgment; large irrelevant state can distract it. It also says counting, numerical precision, and date ordering belong in code. For this ticket, keep the duplicate-charge policy in clear Choice criteria, filter unrelated customer history, and compute money, dates, and loss arithmetic with ordinary program logic. TypeSafe — Jev 1.13 jaggedness

Independently asked questions need not obey cross-question identities automatically. If one answer says “duplicate charge likely” and another says “refund request unlikely,” inspect the wording and the record; do not force a joint interpretation by multiplying marginal probabilities. Test disagreements as explicit cases, then enforce any required structural invariant in code. The vendor guide itself recommends asking each decision directly. TypeSafe — structural invariants

Common mistake: Neither a top probability nor agreement with a second model is a resolved outcome. “80% coverage” also does not mean “80% correct.” Here the 8 acted cases contain 2 errors, so acted accuracy is 75%, and the three-case shifted slice is worse.

Run the complete ten-row check. The standard-library sum forms are equivalent to explicit count loops on the same rows. The assertions pin the displayed default values and a wrong-control case; move threshold to see the trade-off.

# Ten synthetic cases: id, (Returns, Billing, Shipping), truth, proxy, slice.
ROUTES = ("returns", "billing", "shipping")
ROWS = [
 ("A-104",(.60,.38,.02),"billing","returns","standard"),
 ("A-105",(.08,.88,.04),"billing","billing","standard"),
 ("A-106",(.82,.12,.06),"returns","returns","standard"),
 ("A-107",(.14,.12,.74),"shipping","shipping","standard"),
 ("A-108",(.73,.20,.07),"returns","returns","standard"),
 ("A-109",(.12,.79,.09),"returns","billing","shifted"),
 ("A-110",(.45,.50,.05),"billing","billing","shifted"),
 ("A-111",(.06,.16,.78),"shipping","shipping","standard"),
 ("A-112",(.68,.28,.04),"billing","returns","shifted"),
 ("A-113",(.22,.66,.12),"billing","billing","standard"),
]
DATASET_ID = "jev-synthetic-routing-v1"
POLICY_ID = "top-mass-threshold-v1"
FIXTURE_MODEL_ID = "local-hand-authored-probabilities-v1"

def evaluate(rows, threshold=.65, wrong_cost=10, review_cost=2):
    acted = correct = errors = reviewed = shifted_errors = 0
    total_cost = 0
    for case_id, probs, truth, proxy, segment in rows:
        assert abs(sum(probs) - 1) < 1e-12
        index = max(range(len(probs)), key=lambda i: probs[i])
        if probs[index] < threshold:
            reviewed += 1
            total_cost += review_cost  # Teaching assumption: perfect review.
            continue
        acted += 1
        match = ROUTES[index] == truth
        correct += int(match)
        errors += int(not match)
        shifted_errors += int(segment == "shifted" and not match)
        total_cost += 0 if match else wrong_cost
    return {"coverage": acted / len(rows),
            "acted_accuracy": correct / acted if acted else None,
            "acted": acted, "correct": correct, "errors": errors,
            "reviewed": reviewed, "shifted_errors": shifted_errors,
            "total_cost": total_cost}

result = evaluate(ROWS)
assert (result["acted"],result["correct"],result["errors"],
        result["reviewed"],result["shifted_errors"],
        result["total_cost"]) == (8,6,2,2,2,24)
assert result["coverage"] == .8 and result["acted_accuracy"] == .75
assert sum(1 for row in ROWS if row[2] != row[3]) == 3
print(DATASET_ID, POLICY_ID, FIXTURE_MODEL_ID, result)

These metrics tell us what the constructed workflow did on known cases; they do not identify why a model predicted badly. The final chapter connects this workflow to estimation, sequential control, and evaluation tools so you can choose the right next component instead of asking one typed classifier to do every job.

What evidence would support a claim that a new routing policy improved real route quality?

Chapter 11: Know which job each component owns

Return to A-104. The case state is a message about wrong-size shoes and a duplicate charge. A bounded routing question asks for one of three allowed labels. A synthetic fixture assigns Returns .60, Billing .38, Shipping .02. Code then compares its cost estimates, decides whether to route or review, and later checks a resolved label. Each arrow in this chain has a different owner.

Estimate: a typed judgment can provide option probabilities, but their usefulness must be tested on resolved examples. Jev's documented Choice primitive bounds the answer space; it does not write a customer reply or guarantee the chosen label. The probabilities in our A-104 exercise are local invented data, not a response from a live model. TypeSafe — Choice

Decide: the application owns allowed actions, losses, and fallback rules. For A-104, .60 + .38 + .02 = 1. The top label is Returns, with other-label mass .38 + .02 = .40. At wrong-route cost 10, its estimated auto loss is 10 × .40 = 4. A stipulated perfect review costs 2, so this particular rule chooses review. That is a policy calculation, not a feature hidden inside a Choice value.

state → typed fixture → policy: min(4 auto, 2 review) → reviewed Billing → labeled result

Learn and evaluate: a future model or policy can improve only if the workflow records suitable states, actions, propensities when relevant, and independently assessed outcomes. Review labels can help evaluate an auto-route rule, but observed data may be biased by which cases were sent to review. The ten-ticket exercise showed 8 auto-routes and only 6 correct ones at its default threshold; it is a teaching sample, not evidence about a deployed system.

The review choice depends on review quality too. If review itself were correct only 90% of the time and its mistakes had the same cost 10, expected review loss would be 2 + (1 − .90) × 10 = 3, still below the estimated auto loss 4. At 70% correct, it would be 2 + .30 × 10 = 5, making auto cheaper under these particular assumptions. The displayed A-104 result retains the explicit perfect-review assumption; this counterexample shows why that assumption should be measured rather than silently treated as universal.

Explore the four jobs in the pipeline

The map's audit card names the owner, required evidence, and a concrete failure check for each job. For example, a route probability can be syntactically valid while disagreeing with the resolved policy label; a policy can compute a correct minimum over the wrong cost table; and an evaluation can look strong overall while its shifted slice fails. Tap through the card to see how a single ambiguous ticket becomes several different engineering questions.

Choose Estimate to connect this lesson to priors and evidence updates in Bayes Estimation — Prior, Likelihood, Posterior and belief states in POMDP — Observations & Belief. Those are mathematical connections, not claims about Jev's internal algorithm. Choose Decide for Decision Making Under Uncertainty Workbook — Value of Information and MDPs, where costs, information, and future outcomes become explicit.

Choose Learn for Deep Reinforcement Learning — MDPs and policy objectives; its state/action/reward setting goes beyond a typed classifier. Choose Evaluate for Calibration: Knowing What the Model Knows — Cost-Coverage and Ask/Answer/Defer and AI Evaluation — measuring system behavior. Open the linked lesson and choose the named chapter.

A customer-facing explanation is another job again. A generative model or a human can compose and inspect that response after a route decision; a typed route label is not prose. Operational guardrails must also govern side effects such as issuing refunds or contacting a customer. The site’s Agents & Tool Use — Guardrails lesson develops that separation. Keep typed judgment, application policy, tool permissions, and human review visible rather than conflating them into one opaque “AI action.”

Common mistake: A fast, typed probability interface does not by itself guarantee correctness, write a good explanation, choose a risk policy, or solve a sequential-control problem. Those require labels, rules, objectives, and evidence appropriate to each job.

A minimal audit record should preserve the case and fixture version, selected route masses, policy version and costs, chosen action, and later label source. If a reviewer corrected a route, log the correction without silently treating an unreviewed case as correct. This separation prevents a model update, a policy edit, or a change in who gets reviewed from masquerading as the same experiment.

The complete pure-Python trace below reprises the case. Its assertion checks both the probability total and the review decision; replacing the synthetic fixture with a real typed response would still leave policy and outcome evaluation in application code. The standard-library sum is equivalent to adding the three displayed masses manually.

# A-104 is a constructed fixture, not a Jev API result.
case = {
    "id": "A-104", "truth": "billing",
    "probabilities": {"returns": .60, "billing": .38, "shipping": .02},
}

def validate_fixture(case):
    probs = case["probabilities"]
    assert set(probs) == {"returns", "billing", "shipping"}
    assert all(0 <= value <= 1 for value in probs.values())
    assert abs(sum(probs.values()) - 1.0) < 1e-12
    assert case["truth"] in probs


def decide_and_check(case, wrong_cost=10, review_cost=2):
    probs = case["probabilities"]
    assert abs(sum(probs.values()) - 1.0) < 1e-12
    route = max(probs, key=probs.get)
    wrong_mass = sum(p for label, p in probs.items() if label != route)
    auto_estimate = wrong_cost * wrong_mass
    action = "review" if review_cost <= auto_estimate else route
    # This toy assumes review assigns the known policy label correctly.
    assigned = case["truth"] if action == "review" else route
    realized_cost = review_cost if action == "review" else (
        0 if assigned == case["truth"] else wrong_cost)
    return {"top_route": route, "wrong_mass": wrong_mass,
            "auto_estimate": auto_estimate, "action": action,
            "assigned": assigned, "realized_cost": realized_cost}

def review_estimated_loss(review_cost, review_accuracy, wrong_cost):
    assert 0 <= review_accuracy <= 1
    return review_cost + (1 - review_accuracy) * wrong_cost

assert abs(review_estimated_loss(2, .9, 10) - 3) < 1e-12
assert review_estimated_loss(2, .7, 10) > 4

bad_case = {**case, "probabilities": {"returns": .6, "billing": .6,
                                       "shipping": .02}}
try:
    validate_fixture(bad_case)
except AssertionError:
    pass  # Wrong control: probabilities must sum to one.
else:
    raise AssertionError("invalid fixture accepted")

result = decide_and_check(case)
assert result["top_route"] == "returns"
assert abs(result["wrong_mass"] - .4) < 1e-12
assert abs(result["auto_estimate"] - 4) < 1e-12
assert result["action"] == "review"
assert result["assigned"] == "billing" and result["realized_cost"] == 2
print(result)

You can reuse this recipe: define the state and bounded question, keep the answer's provenance, specify a cost-based rule, log the action, and test against independently resolved outcomes. If you need future state control, turn to an MDP; if you need better beliefs after new evidence, turn to Bayesian estimation; if you need confidence-aware deployment, turn to calibration and evaluation. A typed judgment is a useful component because code can consume it precisely, and that precision makes the remaining responsibilities easier to inspect.

Who owns the rule that sends A-104 to review when estimated auto loss 4 exceeds review cost 2?