A New Tier · Guided Practical Projects

Forges

Understanding is the entry ticket; building is the destination. Forges are multi-milestone expeditions on real data and free compute where every step has been machine-verified runnable, every milestone ends in a checkable artifact, and every expedition ends with something public.

1
Expedition Live
15
Capability Bricks
$0
Compute Required

How a Forge works

ReproduceModifyBuildBreak & DebugBenchmarkPublish

● Proof-carrying steps

Commands are badged executed (a machine ran them; the golden artifact is committed) or docs+pinned (verified against exact-version docs). Versions are locked, the verification date is stamped on every page.

✓ Evidence checkpoints

Each milestone ends in one JSON artifact your own run prints. The page validates it in your browser — schema, relational sanity, tolerance vs the golden run — and your progress syncs to your account.

⚠ Break it on purpose

Every Forge prescribes at least one failure you must cause and diagnose yourself. You don’t own a tool until you have watched it fail in a way it can measure.

⬢ Lego bricks

Milestones earn capability bricks that recur across expeditions — the eval harness you build here is the same brick the next Forge starts from. Expertise compounds.

Expeditions

✦ STUDIO SESSION · NEW · THE FAMOUS HOMEWORK
The Gauntlet — CS231n, Studio Edition

Stanford’s Assignment 1 as a voice-led build session on 6,000 real CIFAR-10 photos: vectorize five million distances into one matmul, derive p − y and let finite differences audit it, backpropagate a hidden layer by hand, hand-craft HOG features, race four optimizers — then your champion classifies YOUR photos and sketches.

ASSIGNMENT 1 COMPLETE · 7 acts · the accuracy ladder 10% → 48.4%, every rung yours · all on-device
✦ STUDIO SESSION · NEW · COMPUTER VISION
The Vision Engine — Studio

A full first week on a perception-data team, in your browser: derive the softmax gradient and let finite differences audit it, convict a labeling vendor with a rate table, run a real 24-photo labeling shift, quantize your head to a provable error bound — then point your camera at the world and watch YOUR engine call it.

THE FULL WEEK LIVE · 7 acts · real model + 489 real photos, all on-device · CS231n × the actual job
✦ STUDIO SESSION · COACHED BUILD
The Trip Analyst — Studio

A host walks you through the finished machine running on real driving data — then it powers down, and you rebuild it act by act, in the browser, with the coach in your ear. Your code becomes the engine.

Acts 0–2 live · voice + stage + workbench · predictions, form correction, say-it-back
EXPEDITION 01 · ESTIMATION × ML × LLM
The Trip Analyst — Field Manual

The full 8-milestone build: KF → EKF → HMM → trained emissions → IMM → a Claude-powered analyst with tools — in-page labs, evidence checkpoints, golden artifacts.

8 milestones · ~14 h · laptop-only · verified 2026-07-28
EXPEDITION 02 · INFERENCE SYSTEMS
The Serving Foundry

vLLM + SGLang on free GPU tiers: quantization ablations, prefix caching, load testing, observability, a public serving benchmark.

In the forge — feasibility ledger verified, authoring next
EXPEDITION 03 · ROBOT LEARNING
The Policy Proving Ground

LeRobot ACT vs Diffusion Policy on PushT, trained free with checkpoint-resume, evaluated at n=100+ with fixed seeds, shipped as a model card.

In the forge — feasibility ledger verified, authoring next
Why the name: a forge is where material that was merely understood gets worked under real heat into something that holds weight. Same motto as everywhere on this site — what I cannot create, I do not understand — taken literally.