Understanding is the entry ticket; building is the destination. Forges are multi-milestone expeditions on real data and free compute where every step has been machine-verified runnable, every milestone ends in a checkable artifact, and every expedition ends with something public.
Commands are badged executed (a machine ran them; the golden artifact is committed) or docs+pinned (verified against exact-version docs). Versions are locked, the verification date is stamped on every page.
Each milestone ends in one JSON artifact your own run prints. The page validates it in your browser — schema, relational sanity, tolerance vs the golden run — and your progress syncs to your account.
Every Forge prescribes at least one failure you must cause and diagnose yourself. You don’t own a tool until you have watched it fail in a way it can measure.
Milestones earn capability bricks that recur across expeditions — the eval harness you build here is the same brick the next Forge starts from. Expertise compounds.
Stanford’s Assignment 1 as a voice-led build session on 6,000 real CIFAR-10 photos: vectorize five million distances into one matmul, derive p − y and let finite differences audit it, backpropagate a hidden layer by hand, hand-craft HOG features, race four optimizers — then your champion classifies YOUR photos and sketches.
A full first week on a perception-data team, in your browser: derive the softmax gradient and let finite differences audit it, convict a labeling vendor with a rate table, run a real 24-photo labeling shift, quantize your head to a provable error bound — then point your camera at the world and watch YOUR engine call it.
A host walks you through the finished machine running on real driving data — then it powers down, and you rebuild it act by act, in the browser, with the coach in your ear. Your code becomes the engine.
The full 8-milestone build: KF → EKF → HMM → trained emissions → IMM → a Claude-powered analyst with tools — in-page labs, evidence checkpoints, golden artifacts.
vLLM + SGLang on free GPU tiers: quantization ablations, prefix caching, load testing, observability, a public serving benchmark.
LeRobot ACT vs Diffusion Policy on PushT, trained free with checkpoint-resume, evaluated at n=100+ with fixed seeds, shipped as a model card.