CS224R · HOMEWORK AS FORGE · IMITATION LEARNING

HW1 — Imitation Learning

A Flappy Bird, a 4-number observation, twenty future moves to predict. The simplest imitation-learning algorithm there is — why it flies clean on easy mode, why it drives straight into the wall on hard mode, and the two orthogonal fixes that rescue it.

Prerequisites: basic algebra + comfort with the idea of a neural network as a function. No prior RL or imitation learning assumed. PyTorch primer included.
10
Chapters
10
Live Sims
3
Code Labs
1
Forge Studio

A companion & practice forge for Stanford's CS 224R Homework 1. It credits the public course materials and teaches you to implement the core math yourself — it is not a copy-paste answer bank. The real starter-code contracts (BCPolicy, mse_loss, FlowMatchingSchedule, DeterministicExpert) are the ones you meet here, shrunk to a scale you can run in your browser.

Chapter 0: The Wall

Imagine training a bird from demonstrations of target heights. A single opening gives the expert one route to follow. This illustration isolates what changes when the demonstrations contain two incompatible routes.

On hard mode, pipes alternate between one and two openings. Train the same architecture and squared-error objective on hard-mode demonstrations. When equally likely expert routes share an observation, their mean target can point straight into the solid wall between the openings. This is an idealized mechanism, not a measured claim that every trained policy crashes.

Even a correct implementation can produce this result: minimizing prediction error is not the same objective as avoiding collisions.

The mean can be unsafe

Two valid demonstrations from the same state: one expert aims at the upper gap (y = 0.30), another aims at the lower gap (y = 0.70). Press play. Watch where a mean-seeking policy — one trained to minimize squared error — decides to fly.

upper gap y 0.30
lower gap y 0.70
The big reveal. The averaging policy does not pick a gap. It averages the two heights the expert used and aims for the point exactly between them — which is precisely where the wall is. Two perfectly good answers, blended into one catastrophic one.

This is not a bug in your code. It is a mathematical property of the loss function you chose. Fixing it is what launches the rest of imitation learning — and this lesson walks the whole road: how behavior cloning works, why the wall appears, and the two very different escapes (a richer model called flow matching, and a data trick called DAgger).

On the hard course, the expert sometimes aims at the upper gap and sometimes at the lower gap — from the very same observation. Before you know any of the math, what is the most likely reason a "predict the average action" policy crashes?
Where we are headed. Ten chapters. We build the setup (the bird, the observation, the paradigm), implement behavior cloning by hand, watch it fail on hard mode, then fix it two ways. There is a Forge Studio (the ⚒ button in the mode row) where you practice the core operations in HW1 — the policy MLP, MSE, the flow-matching loss and sampler, the deterministic expert, and the DAgger relabel loop — on a live Flappy Bird instrument, and see how each correct kernel changes the illustrative policy.

Chapter 1: The Bird & The Paradigm

Before we can fix anything we need to know the machine. The environment is a physics-based Flappy Bird. A bird falls under gravity. Pipes scroll left. Each pipe has one gap (easy) or two (hard). The bird must thread the gaps without touching a pipe, the ceiling, or the floor.

The action: a target height, not a flap

The agent does not press "flap." At every timestep it outputs a single number: a target y-position, normalized to [0, 1], where 0 is the top of the screen and 1 is the bottom. A built-in PD controller — a proportional-derivative feedback loop — converts that target into thrust, so the bird has momentum and cannot teleport. Aim for a height; the controller flies you there with realistic lag.

Why this matters for learning. Because the bird has momentum, a good policy must anticipate: to be centered on a gap when the pipe arrives, you have to start climbing several frames early. Learning action sequences can capture anticipation, while execution horizon controls feedback frequency; performance still requires evaluation.

The observation: four numbers

The policy sees a 4-dimensional vector, every entry normalized to roughly [0, 1]:

IndexMeaningNote
obs[0]distance to the next pipecounts down as the pipe approaches
obs[1]y-position of gap 1 (upper)the two gaps are ~0.40 normalized (about 179 px) apart
obs[2]y-position of gap 2 (lower)on easy mode, equal to gap 1
obs[3]the bird's current y-positionthe bird's velocity is not observed
The hidden hard part is already here. On hard mode obs[1] ≠ obs[2] — there are two gaps — but nothing in the observation tells the policy which gap the expert will pick. Same four numbers, two valid answers. That is the seed of the whole multimodality problem.

Two difficulty modes

ModePipesExpertPlain BC
Easyone gap (gap1 == gap2)raw target is the one gap; returned command is smoothedevaluate performance
Hardalternating single- and double-gap pipesrandomly picks a gap → multimodalcan average incompatible routes

Imitation learning vs reinforcement learning

Behavior cloning learns from expert action labels. Reinforcement learning optimizes a reward signal, often using environment interaction. These signals can be combined. Imitation learning includes interactive methods such as DAgger.

Reinforcement LearningImitation Learning
Data sourcetrial and error in the environmentexpert demonstrations; interactive methods also query experts on learner rollouts
Signala reward functionthe expert's actions, treated as labels
Algorithm familyQ-learning, policy gradients, …BC uses supervised learning; other IL methods differ
Hard partexploration, sparse reward, credit assignmentdistribution shift, multimodal experts
Analogylearn chess by self-playlearn chess by watching humans
The imitation-learning loop, in three boxes
1 · Collect
Run the expert. Log every (state, action) pair into a dataset D.
↓
2 · Fit
Train a network πθ(s) to output the expert's action at each state — plain supervised learning.
↓
3 · Deploy
Run πθ in the environment. No reward, no exploration, no Bellman equation.
Why start here? BC is a useful supervised baseline that exposes distribution shift, multimodality, and representation choices. ACT/ALOHA is an example of learning action sequences from demonstrations. These ideas connect to other robot-learning methods without implying that every method starts with BC pretraining.
In this environment, what does the policy's action actually control?

Chapter 2: Behavior Cloning

Strip behavior cloning (BC) down to its essence and it is literally just supervised learning. If you have ever trained an image classifier, you have already done every mechanical step of BC. The only thing that changes is the data: states in, expert actions out.

The dataset

Run the expert in the environment N times. Record every (state, action) pair. You get

D = { (s1, a1*), (s2, a2*), …, (sN, aN*) }

where each ai* is the height the expert aimed for at state si. In the real homework the default is roughly 500 demonstration episodes, with an episode of length L yielding max(0, L−19) overlapping 20-step training windows.

The objective

Find the network parameters θ that make the policy's action match the expert's action, averaged over the whole dataset:

θ* = arg minθ   (1/N) ∑i   L( πθ(si), ai* )

Here L is a loss that measures how far off the prediction is. For Problem 1 of the homework, L is mean squared error. Problem 2 swaps in a generative loss (flow matching). Same overall structure — different L.

Concept → realization. The data flow is exactly a regressor's. Input tensor s of shape [B, 4]. Output tensor πθ(s) of shape [B, 20] (twenty future target heights — Chapter 3 explains the 20). Target tensor a* of shape [B, 20]. The loss reduces the two to a single scalar. Backprop, Adam step, repeat. Nothing exotic.

Why BC is not trivially perfect

The optimization is easy. The hardness lives in three failure modes — and each of the follow-up problems attacks one:

Failure modeWhat goes wrongFixed by
Multimodal expertstwo valid actions per state get averaged into an invalid one (the wall)Problem 2 · flow matching  or  Problem 3 · deterministic expert
Distribution shiftsmall errors compound until the bird is in states the expert never visitedProblem 3 · DAgger
Causal confusionthe network learns a shortcut that works on training data but not in deploymentbetter data / representations (beyond HW1)
Regression fits the demonstrations — that is all it does

A single state, several expert demonstrations (dots), and the line a mean-squared-error fit draws through them. Drag the slider to add spread. On unimodal data the fit sits right on the cluster. Keep this picture — Chapter 5 shows what happens when the dots split into two clusters.

demo spread 0.04
What is behavior cloning, in one sentence?

Chapter 3: Action Chunking

Most textbook BC predicts one action at a time: see a state, output an action, repeat. Modern robot-learning policies — Diffusion Policy, ACT/ALOHA, π0 — do something that looks strange at first: they predict a whole chunk of future actions per policy query. Generative policies may use several network evaluations to produce that chunk.

πθ(st) = ( at, at+1, at+2, …, at+T−1 )

In this homework ACTION_CHUNK = 20: the policy predicts twenty future target heights at once. But it does not execute all twenty. At rollout time it executes only the first EXECUTE_STEPS = 10 before firing again. That "predict 20, execute 10, re-plan" pattern is called receding-horizon control.

Receding horizon: predict a plan, walk half of it, re-plan

The blue bar is the 20-step plan the policy just predicted. The solid segment is the 10 steps actually executed; the faded segment is discarded. Press play: watch the plan slide forward, always re-planning before the executed part runs out.

Why predict more than you execute?

Three concrete reasons, each a real engineering decision:

1 · Temporal consistency. Joint sequence prediction can model relationships between future actions and reduce abrupt changes between independently sampled commands. It does not guarantee coherence: an MSE chunk can still average incompatible routes.
2 · Fewer network queries. Executing ten commands per query reduces network calls. It does not imply ten times less control error: disturbances continue between queries, and prediction errors and dynamics are correlated.
3 · Learning temporal structure. Demonstration chunks contain how targets evolve over time. Predicting those commands can capture anticipation, but is not itself an explicit dynamics planner or a guarantee of collision avoidance.

The chunk-size tradeoff

Chunk vs executeEffect
fixed prediction horizon, execute fewermore frequent policy feedback, with more inference calls
fixed prediction horizon, execute morelonger open-loop commitment; fewer calls but slower reaction to disturbances
predict 20, execute 10 (this HW)the standard compromise — predict more than you execute, re-query before the buffer runs dry

How chunking shapes the network

This is the whole reason the policy's output has 20 dimensions. Input: a state of shape [B, 4]. Output: a chunk of shape [B, 20] — each output number is one future target height. The action_dim argument of BCPolicy defaults to 20 because that is the chunk size. Keep that number in mind; it is the shape everything downstream must match.

Predict 20, execute 10. The last 10 predicted actions are computed and then thrown away every re-plan. They are not wasted — asking the network to think 20 steps out is what makes the executed 10 coherent. It is cheaper to over-predict than to under-plan.
The policy predicts ACTION_CHUNK=20 actions but only executes EXECUTE_STEPS=10 before re-querying. What happens to the other 10?

Chapter 4: The MLP & The MSE Loss

Now the two design decisions that are Problem 1: the network architecture and the loss. Both are textbook — the interesting part (Chapter 5) is what they do together on bimodal data.

A three-layer MLP

The policy, BCPolicy, is a plain multilayer perceptron — the simplest neural network, a stack of matrix multiplies with nonlinearities between them:

state (4-D)
the observation [dist, gap1, gap2, bird_y]
↓
Linear(4 → 256) · ReLU
first hidden layer
↓
Linear(256 → 256) · ReLU
second hidden layer
↓
Linear(256 → 20) · Sigmoid
output: 20 target heights, each in (0, 1)

Why an MLP and nothing fancier? The problem is small — a 4-D input, a 20-D output, simple physics. A convolutional net would be overkill (no spatial structure). A transformer would be overkill (no sequential input structure). Two hidden layers of 256 units is the supplied baseline choice. (Compare: Problem 2's flow-matching policy is a 1D U-Net, which models structure across the action chunk while conditioning velocity predictions on the observation and generation time.)

Why ReLU? It keeps positive inputs and zeros negative ones, adding a nonlinearity between affine layers. It does not saturate on the positive side, but inactive units have zero derivative; healthy gradients are not guaranteed.

Why Sigmoid? It maps real outputs into (0,1), matching the normalized height range. This constraint does not ensure good control or strong gradients: sigmoid saturates near the ends of its range.

In PyTorch. nn.Sequential(nn.Linear(4,256), nn.ReLU(), nn.Linear(256,256), nn.ReLU(), nn.Linear(256,20), nn.Sigmoid()). Call super().__init__() before assigning submodules. Otherwise normal module assignment raises an initialization error; it does not silently train with unregistered layers.

The MSE loss

Mean squared error is the average, over the dataset and over all 20 chunk dimensions, of the squared gap between prediction and expert action:

LMSE(θ) = (1/(20N)) ∑i   || πθ(si) − ai* ||2  =  (1/(20N)) ∑i ∑k=1..20 ( πθ(si)[k] − ai*[k] )2

The one fact you must carry forward

What is the unrestricted population optimum of MSE regression? A finite network may only approximate it. Take the expected squared error at a state and set the derivative with respect to the prediction p to zero:

d/dp   E[ (p − a)2 ] = 2 · E[ p − a ] = 2 · ( p − E[a] ) = 0  ⇒  p = E[a]

The MSE-minimizing prediction is the conditional mean of the expert's actions at that state:

πθ*(s) = Ea* ~ expert[ a* | s ]
Remember this. The unrestricted squared-error optimum is the conditional mean given the observation. A mean can be unsafe when it lies between separated valid actions. Even a unimodal target distribution does not by itself guarantee successful control.

Why squared error and not absolute error?

LossMinimizerTradeoff
L2 · MSE · (p−a)2conditional meansmooth gradient; but averages multimodal targets
L1 · MAE · |p−a|conditional medianrobust to outliers; non-smooth at 0; a median need not be unique or safe

MSE is a common regression objective for continuous commands. Its smoothness is a gift for optimization — and its mean-seeking is a curse for multimodal data.

What does an MSE-trained regressor output, in expectation, at a given state?

Chapter 5: The Multimodality Trap

This is the conceptual climax of Problem 1. Everything before was setup; everything after is a fix. Spend time here.

The setup

Hard mode alternates single-gap and double-gap pipes. When the expert sees a double-gap pipe it hovers at the midpoint while far away, then — when it gets within a commit distance of the pipe — randomly picks one of the two gaps and aims there. Same observation, two different choices, decided by a coin flip inside the expert. To isolate the averaging mechanism, consider two equally likely raw targets at one observation:

The random gap choice is internal to the expert. The actual returned command is smoothed using the expert’s history, so these two raw target values are an idealized example, not a claim that every recorded label is exactly a gap center.

What MSE does with that

From Chapter 4: the MSE minimizer is the conditional mean. For equally probable targets, the conditional mean of 0.30 and 0.70 is

(0.30 + 0.70) / 2 = 0.50

So the ideal squared-error predictor outputs 0.50 at this state — which, because the two gaps are about 0.40 apart (~179 px) and each opening is only about 0.167 tall (75 px, i.e. half-opening ±0.084), is exactly the solid wall between them. The policy is not doing anything wrong. It is faithfully reproducing the conditional mean of the expert distribution. The conditional mean is a valid normalized command, but its target lies in the wall.

Watch the mean walk into the wall as the demos split

Start with all demos aiming at one gap (unimodal). Drag the slider to send more and more of them to the other gap. The dashed line is the MSE optimum — the mean of the demos. Watch it leave the safe gap and drift into the wall as the data becomes bimodal.

fraction aiming lower gap 0.50

The fundamental issue, stated precisely

MSE predicts a mean; it does not require unimodality. Squared loss is defined for multimodal targets. The danger is executing their conditional mean when that mean lies in a wall. Equal mixture weights put the mean halfway; unequal weights move it toward the more frequent target.

Why this is the lesson of HW1

Problem 1 exposes a possible multimodal-regression failure on hard mode. Evaluate mean episode length against the 1000-step cap, and the homework asks you to explain why in a couple of sentences. The answer is exactly: multimodality of the expert + mean-seeking of MSE.

And there are two clean escapes, which take orthogonal routes:

FixIdeaChanges
Problem 2 · Flow matchingmake the model richer — learn the full distribution p(a | s), then sample one modethe model (data untouched)
Problem 3 · DAggeradd consistent new labels on learner-visited states — using a deterministic gap choicethe data (model untouched)

These address different failure mechanisms and can be combined. Measure performance rather than assuming either change guarantees success.

The writeup, in three sentences. "On hard mode the expert chooses between two valid actions (upper or lower gap) randomly at the same state, making the action distribution multimodal. The unrestricted MSE optimum is the conditional mean of expert actions, so the policy outputs the average of the two gaps — a target between them, where the wall is. This can cause crashes; measure the actual episode lengths."
If demo A aims at the upper gap (y=0.3) and demo B aims at the lower gap (y=0.7), both from the same state, what does the MSE-trained policy output there, and why is it a problem?

Chapter 6: Flow Matching — fix the model

The first escape keeps the data exactly as-is and makes the model richer. Instead of a regressor that outputs one number per state, we train a generative model that learns the whole distribution of expert actions at each state — and then samples from it. Same state, different samples can represent either route; an imperfect model can still produce unsafe commands. This is flow matching, the tool of Problem 2. (For a gentler standalone tour see the site's Flow Matching Gleam.)

Regressor vs generative model

Regressor (Problem 1)Generative model (Problem 2)
Given a state, returnsone deterministic actiona sample from p(a | s)
Same state twicesame actionpossibly different actions
On bimodal datathe mean (the wall)can represent either mode; no safety guarantee
The deep idea. A deterministic regressor returns one prediction; a generative policy samples an action chunk from a learned conditional distribution. Even unimodal distributions contain variation that their mean discards. Fresh noise is sampled per policy query, not once per whole episode.

The flow-matching idea: transport noise into data

Flow matching generates a sample from a complex distribution by continuously carrying a sample from a simple one — Gaussian noise — along a learned vector field. Picture the 20-D action space (think of it as a plane). At every point and time, an arrow says "flow this way." Follow the arrows from time τ=0 to τ=1 and a noise sample is carried to a data sample.

dx/dτ = vθ(x, s, τ),   x(0) ~ N(0, I),   x(1) ~ p(a | s)

Different noise points can produce different modes under the learned field. Noise lives in 20-dimensional chunk space; its position is not the physical altitude of the bird. The projected illustration is not a guarantee about every learned trajectory.

Training: regress the velocity along straight lines

The training objective is startlingly simple. Take a clean expert action a1. Draw independent noise a0 ~ N(0, I). Pick a random time τ ~ U(0,1) and interpolate along the straight line between them:

aτ = τ · a1 + (1 − τ) · a0

Differentiate that straight line and the velocity is constant — it does not even depend on τ:

d aτ / dτ = a1 − a0

So the loss is: regress the network's output toward that constant difference.

LFM(θ) = E [   || vθ(aτ, s, τ) − (a1 − a0) ||2   ]
It is still MSE — on the right thing. Plain BC does MSE on actions, which collapses to the mean. Flow matching does MSE on velocities at random interpolation points. Averaged over many noise/data pairs, the network learns E[a1 − a0 | aτ=x, s, τ] — the population velocity associated with the probability path. Finite training and numerical integration only approximate that transport.

Sampling: Euler integration

To generate an action for state s, start from noise and take small steps in the direction of the learned velocity. With num_steps = 20, the step size is h = 1/20 = 0.05:

xτ+h = xτ + h · vθ(xτ, s, τ)

Repeat 20 times from τ=0 to τ=1, then clamp to [0,1]. Straight conditional training paths do not ensure straight marginal sampling trajectories. Step count trades numerical error against computation; flow matching is not universally faster than every diffusion sampler.

Noise splits into two gaps — the vector field at work

Schematic transport, not a trained-vector-field measurement: twenty noise samples on the left (τ=0) move rightward in a schematic field toward a bimodal target — upper gap (teal) and lower gap (blue). Press play. Notice: the illustrated samples reach different modes; real learned samples are not guaranteed safe.

An ideal one-step limit. With independent noise and target chunks, the population-optimal field at τ=0 is E[a|s]−x. One full Euler step therefore yields E[a|s]. This calculation explains a possible failure of this particular sampler; it does not rule out one-step models trained with different objectives or distillation.
Why does flow matching solve the multimodality problem that broke plain MSE?

Chapter 7: DAgger — fix the data

The second escape keeps the model and the loss exactly as Problem 1 — plain MLP, plain MSE — and changes the data. It is the tool of Problem 3, and the homework combines interventions for: the multimodality you just met, and a second, more fundamental failure of BC called distribution shift.

The distribution-shift problem

BC learns from expert-visited observations, but deployment observations depend on the learner's actions. An error can move it into poorly covered regions, where later errors can compound. This is a possible failure mechanism, not a fixed fifty-step deadline.

Errors compound: the drift off the expert's states

The green band is the states the expert visited (the only states BC was trained on). Press play: the agent starts inside it, but each small error nudges it out, and the further out it goes the worse it acts — a runaway. Raise the per-step error to see the drift accelerate.

per-step error 0.020
A worst-case bound, not a crash probability. Under bounded-cost classification assumptions, BC can incur an excess cost of order T²ε. DAgger obtains improved dependence under suitable learning and recoverability assumptions. These are not empirical laws for raw action MSE, and lower error can help. The 1000-step cap does not turn a 1% error into a guaranteed crash.

DAgger: iteratively relabel the policy's own states

DAgger (Dataset Aggregation; Ross, Gordon, Bagnell 2011) closes the gap directly. Each round: roll out the current policy to collect the states it actually visits, ask the expert what to do at those states, add those (state, expert-action) pairs to the dataset, and retrain. Aggregation adds supervision on learner-induced states; finite rounds do not guarantee matching distributions or zero error.

1 · Roll out
Run the current policy πk; collect the states it visits.
↓
2 · Relabel
Query the expert at each visited state — expert actions, not policy actions.
↓
3 · Aggregate
Add the new pairs to the growing dataset D.
↓
4 · Retrain
Plain MSE BC on all of D. Repeat for 5 rounds.
The one-sentence insight. The learner supplies visited observations and the expert supplies labels there. Adding these examples can reduce the coverage mismatch. Linear-in-horizon guarantees require appropriate loss, learning, and recovery assumptions; they are not an unconditional promise of successful flights.

The deterministic-expert trick

Here is the cleverest part of HW1. The original expert is multimodal — it randomly picks a gap. If we relabeled with it, we would keep adding multimodal labels, and MSE would keep averaging into the wall. DAgger would fix distribution shift but not multimodality.

So for relabeling we swap in a deterministic expert that always picks the same gap (gap 1, the upper one) when it commits:

Original expert (multimodal):

if dist < commit_dist:
    target_gap = np.random.choice([0, 1])  # coin flip!
    committed = True

Deterministic expert (unimodal):

if dist < commit_dist:
    committed = True
    raw_target = float(gap1_y)  # ALWAYS gap 1
Two separate interventions. Consistent gap choice removes one source of ambiguity in new labels. The expert still hovers while far away and EMA-smooths its targets, so labels need not equal the gap center. The original demonstrations remain in the aggregate. DAgger improves coverage; this particular labeling strategy also reduces conflicting choices. Neither guarantees perfection.

DAgger with action chunking

One wrinkle: the policy predicts 20-step chunks but executes 10 (receding horizon). During a rollout you query the expert at every step, storing a per-step list of (state, expert-action). Afterward you window that list into chunks: state st gets paired with the next 20 expert actions [a*t, …, a*t+19]. Window only within one episode and reset the expert between episodes. These labels are corrections along the learner trajectory, not necessarily the actions on an expert-controlled rollout from the first state.

In DAgger, who provides the states and who provides the action labels — and why does it matter?

Chapter 8: Showcase — Three Methods, One Bird

This is a scripted illustration, not an evaluation of trained policies. Three methods — plain BC regression, flow matching, DAgger — attacking the same problem: a multimodal expert on a long-horizon control task. Pick a method, press play, and watch the bird fly the hard course. In this simplified animation, methods select scripted aims; real policies also differ in training and computation. The displayed difference is where the bird aims when it sees a double-gap pipe.

Schematic hard course — compare mechanisms

Buttons choose the policy. BC-MSE aims for the mean of the two gaps → the wall → crash. Flow samples one gap and commits. DAgger was retrained on deterministic-expert labels and always takes gap 1. Watch the "aiming at" readout and the crash counter.

What the three methods teach

MethodHard-mode resultHow it solves the problem
BC regression (P1)measure mean ± stdestimates the conditional mean, which can lie between safe routes.
Flow matching (P2)measure mean ± stdkeep the data, enrich the model — a generative policy samples a chunk per query.
DAgger (P3, final round)measure mean ± stdkeep the model, fix the data — consistent new labels reduce one ambiguity while aggregation improves state coverage.
Two orthogonal cures. Flow matching makes the model richer; DAgger makes the data cleaner. Real robot-learning systems often combine them — a diffusion or flow policy trained with DAgger-style relabeling. The three problems of HW1 are the three ideas you will keep meeting for the rest of robot learning.
The costs are not equal. Flow matching pays at inference (20 velocity-network evaluations per predicted chunk, versus 1 for the MLP). DAgger needs an interactive expert you can query at any state — cheap in simulation, expensive when the expert is a human teleoperator. The MLP is cheaper per query, but its conditional mean can be unsafe in this illustration. Pick your fix by your constraint.
Practice the core operations. The Showcase animates the outcome; the Forge Studio (the ⚒ button in the mode row) walks you through core operations practiced in HW1 — BCPolicy.forward, mse_loss, flow_matching_loss, the schedule’s interpolate + sample, DeterministicExpert.act, and rollout_and_relabel — on a live Flappy Bird stage. The bird crashes until your code goes green, then clears the gap. The browser kernels practice these concepts; completing the full starter requires the actual framework and rollout contracts.

Chapter 9: Field Guide — Cheat Sheet & Self-Quiz

Everything worth carrying out of HW1, on one page. If you can reconstruct this from memory, you can teach it.

The equations

BC objective:   θ* = arg minθ E(s,a*)~D [ || πθ(s) − a* ||2 ]
MSE minimizer:   πθ*(s) = E[ a* | s ]   (the conditional mean)
Flow training:   aτ = τ a1 + (1−τ) a0,   target v = a1 − a0
Flow sampling (Euler):   xτ+h = xτ + h · vθ(xτ, s, τ)
Under appropriate loss/learning assumptions: linear DAgger vs quadratic worst-case BC horizon dependence; not measured crash probabilities

The three problems, one table

P1 · BC-MSEP2 · Flow MatchingP3 · DAgger
What you change— (baseline)the modelthe data
Model3-layer MLP1-D U-Net (given)3-layer MLP (same as P1)
LossMSE on actionsMSE on velocitiesMSE on actions (same as P1)
Fixes multimodality?nocan represent multiple modesreduces new-label ambiguity (consistent strategy)
Fixes distribution shift?nonocan improve coverage (expert labels on learner states)
Inference cost1 forward pass20 (Euler)1 forward pass
Needs interactive expert?nonoyes

The real starter-code contracts

You implementFileIn one line
BCPolicy.__init__ / forwardnetworks.pythe nn.Sequential MLP; return self.net(state)
mse_losslosses.pyF.mse_loss(policy(s), a)
FlowMatchingSchedule.interpolatenetworks.pyx_t = t*x1+(1-t)*eps; v = x1-eps
FlowMatchingSchedule.samplenetworks.pyEuler loop over num_steps, then clamp(0,1)
flow_matching_losslosses.pyinterpolate → model → mse_loss(v_pred, v_target)
DeterministicExpert.actdagger.pyraw_target = float(gap1_y) when committed
rollout_episode / rollout_and_relabeldagger.pyroll out policy, relabel with expert, window into chunks

Named failure modes & their symptoms

Failure modeSymptom you would actually see
Multimodal collapselow hard-mode episode length despite better easy-mode results; the bird hits the wall between two open gaps
Forgot super().__init__()assigning submodules before base initialization normally raises an error
Forgot the output Sigmoidpredictions leave [0,1]; unstable, random-looking performance even on easy mode
Wrong action_dimoutputs one number instead of a 20-chunk; the shape mismatch or unintended broadcasting can invalidate the loss
Flow: querying model at τ+hsubtly off-distribution samples; performance drops with no crash
Flow: num_steps = 1ideal independent-noise endpoint calculation gives the mean; finite models may differ
DAgger: labels from policy not expertno improvement over rounds; the policy just reinforces its own mistakes
DAgger: forgot obs.copy()if the environment reuses mutable arrays, stored states can alias the latest value

Self-quiz

For independent noise and the exact population-optimal field at time zero, where does one full Euler step land?
Suppose you ran DAgger but relabeled with the original (multimodal) expert instead of the deterministic one. Would it fix distribution shift? Would it fix multimodality?
What motivates quadratic-versus-linear horizon bounds under the relevant loss and learning assumptions?

Take it back to class

If a friend asks "Why does behavior cloning fail on hard mode?" — you say:

An MSE predictor estimates the conditional mean, which can lie between valid routes. A generative policy can represent multiple choices. DAgger adds expert labels on learner-visited states; this homework also makes new gap choices consistent while retaining old demonstrations. Evaluate the resulting closed-loop behavior.

Where this connects

This forge is the entry point to the CS224R arc. Related on the site:

"What I cannot create, I do not understand." — Feynman. You have now created all three: the mean-seeking regressor, the noise-transporting flow, and the self-correcting relabeler. Build them in the Studio and the understanding is yours.

Companion & practice forge for Stanford CS 224R Homework 1 (Imitation Learning). Credits the public course materials. Implement the core math yourself; this is not a solution answer bank.

Sources and illustration scope. The official assignment and starter code define the contracts. See DAgger theory and flow matching. Canvases and browser kernels are simplified demonstrations, not trained-policy benchmarks. Scalar, unsmoothed expert examples isolate gap choice; the real expert maintains history and smooths targets. Illustrative successful samples do not establish safety guarantees.