Datasets:
arm stringclasses 4
values | size stringclasses 6
values | tokens float64 0.01 68.7 | suite stringclasses 3
values | seed stringclasses 6
values | value float64 0 0.77 |
|---|---|---|---|---|---|
dclm | 100k | 0.1 | raw | seed-0 | 0.00952 |
dclm | 100k | 0.1 | text | seed-0 | 0 |
dclm | 100k | 0.1 | word | seed-0 | 0 |
dclm | 100k | 0.25 | raw | seed-0 | 0.00403 |
dclm | 100k | 0.25 | text | seed-0 | 0 |
dclm | 100k | 0.25 | word | seed-0 | 0 |
dclm | 100k | 0.5 | raw | seed-0 | 0.01013 |
dclm | 100k | 0.5 | text | seed-0 | 0.0026 |
dclm | 100k | 0.5 | word | seed-0 | 0 |
dclm | 100k | 1 | raw | seed-0 | 0.01135 |
dclm | 100k | 1 | text | seed-0 | 0.00087 |
dclm | 100k | 1 | word | seed-0 | 0 |
dclm | 100k | 1.617 | raw | seed-0 | 0.00549 |
dclm | 100k | 1.617 | text | seed-0 | 0.00694 |
dclm | 100k | 1.617 | word | seed-0 | 0 |
dclm | 100k | 3.228 | raw | seed-0 | 0.01318 |
dclm | 100k | 3.228 | text | seed-0 | 0.00434 |
dclm | 100k | 3.228 | word | seed-0 | 0 |
dclm | 100k | 6.449 | raw | seed-0 | 0.00928 |
dclm | 100k | 6.449 | text | seed-0 | 0 |
dclm | 100k | 6.449 | word | seed-0 | 0 |
dclm | 100k | 12.891 | raw | seed-0 | 0.00952 |
dclm | 100k | 12.891 | text | seed-0 | 0.00347 |
dclm | 100k | 12.891 | word | seed-0 | 0 |
dclm | 100k | 17.723 | raw | seed-0 | 0.0083 |
dclm | 100k | 17.723 | text | seed-0 | 0.00347 |
dclm | 100k | 17.723 | word | seed-0 | 0 |
dclm | 1M | 0.1 | raw | seed-0 | 0.00513 |
dclm | 1M | 0.1 | raw | seed-1 | 0.01086 |
dclm | 1M | 0.1 | text | seed-0 | 0 |
dclm | 1M | 0.1 | text | seed-1 | 0 |
dclm | 1M | 0.1 | word | seed-0 | 0.00521 |
dclm | 1M | 0.1 | word | seed-1 | 0 |
dclm | 1M | 0.25 | raw | seed-0 | 0.00806 |
dclm | 1M | 0.25 | raw | seed-1 | 0.01453 |
dclm | 1M | 0.25 | text | seed-0 | 0 |
dclm | 1M | 0.25 | text | seed-1 | 0 |
dclm | 1M | 0.25 | word | seed-0 | 0 |
dclm | 1M | 0.25 | word | seed-1 | 0 |
dclm | 1M | 0.5 | raw | seed-0 | 0.00793 |
dclm | 1M | 0.5 | raw | seed-1 | 0.00757 |
dclm | 1M | 0.5 | text | seed-0 | 0.00087 |
dclm | 1M | 0.5 | text | seed-1 | 0 |
dclm | 1M | 0.5 | word | seed-0 | 0 |
dclm | 1M | 0.5 | word | seed-1 | 0 |
dclm | 1M | 1 | raw | seed-0 | 0.00952 |
dclm | 1M | 1 | raw | seed-1 | 0.00732 |
dclm | 1M | 1 | text | seed-0 | 0.00174 |
dclm | 1M | 1 | text | seed-1 | 0 |
dclm | 1M | 1 | word | seed-0 | 0 |
dclm | 1M | 1 | word | seed-1 | 0 |
dclm | 1M | 1.617 | raw | seed-0 | 0.00879 |
dclm | 1M | 1.617 | raw | seed-1 | 0.005 |
dclm | 1M | 1.617 | text | seed-0 | 0.00174 |
dclm | 1M | 1.617 | text | seed-1 | 0.00174 |
dclm | 1M | 1.617 | word | seed-0 | 0.08594 |
dclm | 1M | 1.617 | word | seed-1 | 0.03906 |
dclm | 1M | 3.228 | raw | seed-0 | 0.02222 |
dclm | 1M | 3.228 | raw | seed-1 | 0.00793 |
dclm | 1M | 3.228 | text | seed-0 | 0.1224 |
dclm | 1M | 3.228 | text | seed-1 | 0.00521 |
dclm | 1M | 3.228 | word | seed-0 | 0.48438 |
dclm | 1M | 3.228 | word | seed-1 | 0.02083 |
dclm | 1M | 6.449 | raw | seed-0 | 0.03564 |
dclm | 1M | 6.449 | raw | seed-1 | 0.03687 |
dclm | 1M | 6.449 | text | seed-0 | 0.17101 |
dclm | 1M | 6.449 | text | seed-1 | 0.12674 |
dclm | 1M | 6.449 | word | seed-0 | 0.48958 |
dclm | 1M | 6.449 | word | seed-1 | 0.48698 |
dclm | 1M | 12.891 | raw | seed-0 | 0.03052 |
dclm | 1M | 12.891 | raw | seed-1 | 0.0603 |
dclm | 1M | 12.891 | text | seed-0 | 0.11806 |
dclm | 1M | 12.891 | text | seed-1 | 0.14236 |
dclm | 1M | 12.891 | word | seed-0 | 0.52344 |
dclm | 1M | 12.891 | word | seed-1 | 0.52604 |
dclm | 1M | 17.723 | raw | seed-0 | 0.02332 |
dclm | 1M | 17.723 | raw | seed-1 | 0.05896 |
dclm | 1M | 17.723 | text | seed-0 | 0.17101 |
dclm | 1M | 17.723 | text | seed-1 | 0.12934 |
dclm | 1M | 17.723 | word | seed-0 | 0.50781 |
dclm | 1M | 17.723 | word | seed-1 | 0.52344 |
dclm | 24M | 0.1 | raw | seed-0 | 0.01038 |
dclm | 24M | 0.1 | text | seed-0 | 0 |
dclm | 24M | 0.1 | word | seed-0 | 0 |
dclm | 24M | 0.25 | raw | seed-0 | 0.01184 |
dclm | 24M | 0.25 | text | seed-0 | 0 |
dclm | 24M | 0.25 | word | seed-0 | 0.0026 |
dclm | 24M | 0.5 | raw | seed-0 | 0.02124 |
dclm | 24M | 0.5 | text | seed-0 | 0.19271 |
dclm | 24M | 0.5 | word | seed-0 | 0.56771 |
dclm | 24M | 1 | raw | seed-0 | 0.05139 |
dclm | 24M | 1 | text | seed-0 | 0.21615 |
dclm | 24M | 1 | word | seed-0 | 0.64583 |
dclm | 24M | 1.617 | raw | seed-0 | 0.06946 |
dclm | 24M | 1.617 | text | seed-0 | 0.21615 |
dclm | 24M | 1.617 | word | seed-0 | 0.57812 |
dclm | 24M | 3.228 | raw | seed-0 | 0.05298 |
dclm | 24M | 3.228 | text | seed-0 | 0.18576 |
dclm | 24M | 3.228 | word | seed-0 | 0.63802 |
dclm | 24M | 6.449 | raw | seed-0 | 0.05615 |
Self-play vs. text pretraining: ICL results
These are the full evaluation results from comparing the in-context learning (ICL) of:
- the self-play learners from Self-Play Pretraining with Zero Data, and
- same-size models trained on ordinary web text for the same number of tokens.
Code and write-up: github.com/mihir-s-05/icl-selfplay-vs-text. Text-model checkpoints: rihim/icl-selfplay-vs-text-checkpoints.
What was scored
Model families (the arm column):
| arm | What | Seeds | Checkpoints |
|---|---|---|---|
sp |
the paper's self-play learners | 4 | rounds 0, 256, 512, 1024, 2048, 2816, 8191 |
up |
the paper's universal-prior baseline (random programs) | 4 | same rounds, sizes up to 6M |
dclm |
same architecture trained on DCLM text | 1–2 | 0.1B → 17.7B tokens |
sp2dclm |
self-play checkpoint (round 8191), then DCLM | 1 | 0.05B → 0.5B tokens |
Sizes: 100k, 500k, 1M, 3M, 6M and 24M parameters.
Task suites (every task shows m worked examples, then a query; the score is greedy exact match on the next byte or bytes):
- raw: the paper's own harness, byte for byte: random bytes 1–255 with a zero byte between examples. Tasks are reverse string, stack, associative recall, sum, max, min, first, last, plus the paper's extra and control cells.
- printable: the same task families written as text, like
\nqd=q(first/last/max/min of two letters),\n37+85=122,\n+a+b-b-→a(stack), letter key→value recall, and a letter cipher. - words: real English words in 10 categories, like
\ncat=t\nmouth=b\nrabbit=→t. Covers classification with 2 or 4 arbitrary labels where the query word is unseen in the examples, lookup of a seen word, and reversing a 3-word phrase.
Multi-byte answers count only if every byte is right. The raw reverse task uses per-byte accuracy, as the paper does.
Layout
results/icl/<arm>/<size>/<checkpoint>.json raw + printable suites (all arms)
results/icl_words/<arm>/<size>/<checkpoint>.json word suite (all arms)
results/icl_s1/dclm/<size>/<checkpoint>.json second DCLM seed (1M, 3M, 6M), all suites
results/report/ auto-generated report and figures
training_logs/{dclm,sp2dclm}/<size>/seed-*/ log.jsonl + run.json for every training run
training_logs/lrsweep/lrsweep-<lr>/<size>/seed-0/ the short LR-sweep runs
lr.json LR chosen per size
tables/*.csv tidy tables (also shown in the viewer above)
For self-play and universal-prior checkpoints, <checkpoint> is the self-play round. For
DCLM runs it is the token count in millions (e.g. 17723M).
Each result JSON looks like this:
{"meta": {"arm": "dclm", "size": "6M", "tokens": 17723031552, "checkpoints": [...], "bf16": true, ...},
"results": {"<section>": {"<task|m|...>": {"n_trials": 128, "chance": 0.038,
"per_seed": {"seed-0/learner_17723M.pth": {"acc": 0.93, "p_correct": 0.71, ...}},
"ensemble": {...}}}}}
ensemble averages the seeds' probabilities before taking the argmax, which is the paper's
convention. It is present when a file scores more than one seed.
The tables:
headline.csvhas one row per (arm, size, tokens, suite, seed), holding the suite's headline mean. The word headline excludes word-order reversal, which every model scores 0 on.tasks.csvhas the per-task cells behind those means.mcurves.csvhas accuracy against the number of examples m, for every task.valbpb.csvhas validation bits/byte over training for the DCLM and warm-start runs.sp_bpb.csvhas the paper's own zero-shot DCLM bits/byte for the self-play ladder.
How these were made
The raw-suite prompts were checked to be byte-identical to the paper's scripts. Scoring the paper's 24M ensemble reproduces its Fig. 4 to within a mean absolute difference of 0.002. Evaluations ran in bf16 on an A100. The paper's scripts used fp32; the difference was noise-level when checked.
- Downloads last month
- 478