Mulligan datasets
Five campaigns, with original recordings, training views, and evaluations connected by explicit source and parent links. Browse the verified release copies and their recorded ancestry.
Audited · Hugging Face organization · Policy Arena · Download mapping CSV · Full manifest · Release plan · Training recipes
Mixed teleop and DAgger sessions preserve every recorded attempt. Evaluation parents preserve their policy roster and shared starts.
Views select episodes from parents. They are useful inputs, but adding their episode counts would count recordings twice.
Evaluation blocks stay distinct. A historical eval can feed a later critic; consult the recipe before treating it as held out.
How to read rounds and names
cNN is a collection increment; rNN is an evaluated model round; bNN is a recorded evaluation block. All policy rollout datasets end in policy-rollouts.
Routing maps C00 → R0, C01+C02 → R1, C03+C04 → R2, C05+C06 → R3, C07+C08 → R4, and C09 → R5. Its final R0–R5 evaluation is one 15-policy repository. Required older collector evaluations retain legacy-rNN names.
Mainline contains the primary collections and selected evaluations. Supporting contains policy rollouts and their source recordings. Ablation contains same-campaign no-CF, Sobol, and first100 alternatives. Older real D1/pilot campaigns are excluded, apart from Routing ancestry used by D2.
Simulation evaluation bundles
mulligan/sim-square-narrow-r00-r03-evalmulligan/sim-square-broad-r00-r03-eval
These are two new result bundles, separate from the existing HF copies below. Both bundles provide verified per-state outcomes, exact grid definitions, policy artifact IDs and matching throughput. All 320 mainline seed results are included, alongside 230 named DIVL seed evaluations. Exact cell and source-file plan.
Source-to-release mapping
| Release dataset and pinned source | Role / group | Round | Episodes |
|---|