Understanding is the entry ticket; building is the destination. Forges are multi-milestone expeditions on real data and free compute where every step has been machine-verified runnable, every milestone ends in a checkable artifact, and every expedition ends with something public.
Commands are badged executed (a machine ran them; the golden artifact is committed) or docs+pinned (verified against exact-version docs). Versions are locked, the verification date is stamped on every page.
Each milestone ends in one JSON artifact your own run prints. The page validates it in your browser — schema, relational sanity, tolerance vs the golden run — and your progress syncs to your account.
Every Forge prescribes at least one failure you must cause and diagnose yourself. You don’t own a tool until you have watched it fail in a way it can measure.
Milestones earn capability bricks that recur across expeditions — the eval harness you build here is the same brick the next Forge starts from. Expertise compounds.
One real minute of Highway 280. Build KF → EKF → HMM → trained emissions → IMM → a Claude-powered analyst with tools — and discover they are one idea: allocating trust across four kinds of uncertainty.
vLLM + SGLang on free GPU tiers: quantization ablations, prefix caching, load testing, observability, a public serving benchmark.
LeRobot ACT vs Diffusion Policy on PushT, trained free with checkpoint-resume, evaluated at n=100+ with fixed seeds, shipped as a model card.