Dataset Viewer
Auto-converted to Parquet Duplicate
arm
stringclasses
4 values
size
stringclasses
6 values
tokens
float64
0.01
68.7
suite
stringclasses
3 values
seed
stringclasses
6 values
value
float64
0
0.77
dclm
100k
0.1
raw
seed-0
0.00952
dclm
100k
0.1
text
seed-0
0
dclm
100k
0.1
word
seed-0
0
dclm
100k
0.25
raw
seed-0
0.00403
dclm
100k
0.25
text
seed-0
0
dclm
100k
0.25
word
seed-0
0
dclm
100k
0.5
raw
seed-0
0.01013
dclm
100k
0.5
text
seed-0
0.0026
dclm
100k
0.5
word
seed-0
0
dclm
100k
1
raw
seed-0
0.01135
dclm
100k
1
text
seed-0
0.00087
dclm
100k
1
word
seed-0
0
dclm
100k
1.617
raw
seed-0
0.00549
dclm
100k
1.617
text
seed-0
0.00694
dclm
100k
1.617
word
seed-0
0
dclm
100k
3.228
raw
seed-0
0.01318
dclm
100k
3.228
text
seed-0
0.00434
dclm
100k
3.228
word
seed-0
0
dclm
100k
6.449
raw
seed-0
0.00928
dclm
100k
6.449
text
seed-0
0
dclm
100k
6.449
word
seed-0
0
dclm
100k
12.891
raw
seed-0
0.00952
dclm
100k
12.891
text
seed-0
0.00347
dclm
100k
12.891
word
seed-0
0
dclm
100k
17.723
raw
seed-0
0.0083
dclm
100k
17.723
text
seed-0
0.00347
dclm
100k
17.723
word
seed-0
0
dclm
1M
0.1
raw
seed-0
0.00513
dclm
1M
0.1
raw
seed-1
0.01086
dclm
1M
0.1
text
seed-0
0
dclm
1M
0.1
text
seed-1
0
dclm
1M
0.1
word
seed-0
0.00521
dclm
1M
0.1
word
seed-1
0
dclm
1M
0.25
raw
seed-0
0.00806
dclm
1M
0.25
raw
seed-1
0.01453
dclm
1M
0.25
text
seed-0
0
dclm
1M
0.25
text
seed-1
0
dclm
1M
0.25
word
seed-0
0
dclm
1M
0.25
word
seed-1
0
dclm
1M
0.5
raw
seed-0
0.00793
dclm
1M
0.5
raw
seed-1
0.00757
dclm
1M
0.5
text
seed-0
0.00087
dclm
1M
0.5
text
seed-1
0
dclm
1M
0.5
word
seed-0
0
dclm
1M
0.5
word
seed-1
0
dclm
1M
1
raw
seed-0
0.00952
dclm
1M
1
raw
seed-1
0.00732
dclm
1M
1
text
seed-0
0.00174
dclm
1M
1
text
seed-1
0
dclm
1M
1
word
seed-0
0
dclm
1M
1
word
seed-1
0
dclm
1M
1.617
raw
seed-0
0.00879
dclm
1M
1.617
raw
seed-1
0.005
dclm
1M
1.617
text
seed-0
0.00174
dclm
1M
1.617
text
seed-1
0.00174
dclm
1M
1.617
word
seed-0
0.08594
dclm
1M
1.617
word
seed-1
0.03906
dclm
1M
3.228
raw
seed-0
0.02222
dclm
1M
3.228
raw
seed-1
0.00793
dclm
1M
3.228
text
seed-0
0.1224
dclm
1M
3.228
text
seed-1
0.00521
dclm
1M
3.228
word
seed-0
0.48438
dclm
1M
3.228
word
seed-1
0.02083
dclm
1M
6.449
raw
seed-0
0.03564
dclm
1M
6.449
raw
seed-1
0.03687
dclm
1M
6.449
text
seed-0
0.17101
dclm
1M
6.449
text
seed-1
0.12674
dclm
1M
6.449
word
seed-0
0.48958
dclm
1M
6.449
word
seed-1
0.48698
dclm
1M
12.891
raw
seed-0
0.03052
dclm
1M
12.891
raw
seed-1
0.0603
dclm
1M
12.891
text
seed-0
0.11806
dclm
1M
12.891
text
seed-1
0.14236
dclm
1M
12.891
word
seed-0
0.52344
dclm
1M
12.891
word
seed-1
0.52604
dclm
1M
17.723
raw
seed-0
0.02332
dclm
1M
17.723
raw
seed-1
0.05896
dclm
1M
17.723
text
seed-0
0.17101
dclm
1M
17.723
text
seed-1
0.12934
dclm
1M
17.723
word
seed-0
0.50781
dclm
1M
17.723
word
seed-1
0.52344
dclm
24M
0.1
raw
seed-0
0.01038
dclm
24M
0.1
text
seed-0
0
dclm
24M
0.1
word
seed-0
0
dclm
24M
0.25
raw
seed-0
0.01184
dclm
24M
0.25
text
seed-0
0
dclm
24M
0.25
word
seed-0
0.0026
dclm
24M
0.5
raw
seed-0
0.02124
dclm
24M
0.5
text
seed-0
0.19271
dclm
24M
0.5
word
seed-0
0.56771
dclm
24M
1
raw
seed-0
0.05139
dclm
24M
1
text
seed-0
0.21615
dclm
24M
1
word
seed-0
0.64583
dclm
24M
1.617
raw
seed-0
0.06946
dclm
24M
1.617
text
seed-0
0.21615
dclm
24M
1.617
word
seed-0
0.57812
dclm
24M
3.228
raw
seed-0
0.05298
dclm
24M
3.228
text
seed-0
0.18576
dclm
24M
3.228
word
seed-0
0.63802
dclm
24M
6.449
raw
seed-0
0.05615
End of preview. Expand in Data Studio

Self-play vs. text pretraining: ICL results

These are the full evaluation results from comparing the in-context learning (ICL) of:

Code and write-up: github.com/mihir-s-05/icl-selfplay-vs-text. Text-model checkpoints: rihim/icl-selfplay-vs-text-checkpoints.

What was scored

Model families (the arm column):

arm What Seeds Checkpoints
sp the paper's self-play learners 4 rounds 0, 256, 512, 1024, 2048, 2816, 8191
up the paper's universal-prior baseline (random programs) 4 same rounds, sizes up to 6M
dclm same architecture trained on DCLM text 1–2 0.1B → 17.7B tokens
sp2dclm self-play checkpoint (round 8191), then DCLM 1 0.05B → 0.5B tokens

Sizes: 100k, 500k, 1M, 3M, 6M and 24M parameters.

Task suites (every task shows m worked examples, then a query; the score is greedy exact match on the next byte or bytes):

  • raw: the paper's own harness, byte for byte: random bytes 1–255 with a zero byte between examples. Tasks are reverse string, stack, associative recall, sum, max, min, first, last, plus the paper's extra and control cells.
  • printable: the same task families written as text, like \nqd=q (first/last/max/min of two letters), \n37+85=122, \n+a+b-b- → a (stack), letter key→value recall, and a letter cipher.
  • words: real English words in 10 categories, like \ncat=t\nmouth=b\nrabbit= → t. Covers classification with 2 or 4 arbitrary labels where the query word is unseen in the examples, lookup of a seen word, and reversing a 3-word phrase.

Multi-byte answers count only if every byte is right. The raw reverse task uses per-byte accuracy, as the paper does.

Layout

results/icl/<arm>/<size>/<checkpoint>.json      raw + printable suites (all arms)
results/icl_words/<arm>/<size>/<checkpoint>.json  word suite (all arms)
results/icl_s1/dclm/<size>/<checkpoint>.json    second DCLM seed (1M, 3M, 6M), all suites
results/report/                                 auto-generated report and figures
training_logs/{dclm,sp2dclm}/<size>/seed-*/     log.jsonl + run.json for every training run
training_logs/lrsweep/lrsweep-<lr>/<size>/seed-0/  the short LR-sweep runs
lr.json                                         LR chosen per size
tables/*.csv                                    tidy tables (also shown in the viewer above)

For self-play and universal-prior checkpoints, <checkpoint> is the self-play round. For DCLM runs it is the token count in millions (e.g. 17723M).

Each result JSON looks like this:

{"meta": {"arm": "dclm", "size": "6M", "tokens": 17723031552, "checkpoints": [...], "bf16": true, ...},
 "results": {"<section>": {"<task|m|...>": {"n_trials": 128, "chance": 0.038,
             "per_seed": {"seed-0/learner_17723M.pth": {"acc": 0.93, "p_correct": 0.71, ...}},
             "ensemble": {...}}}}}

ensemble averages the seeds' probabilities before taking the argmax, which is the paper's convention. It is present when a file scores more than one seed.

The tables:

  • headline.csv has one row per (arm, size, tokens, suite, seed), holding the suite's headline mean. The word headline excludes word-order reversal, which every model scores 0 on.
  • tasks.csv has the per-task cells behind those means.
  • mcurves.csv has accuracy against the number of examples m, for every task.
  • valbpb.csv has validation bits/byte over training for the DCLM and warm-start runs.
  • sp_bpb.csv has the paper's own zero-shot DCLM bits/byte for the self-play ladder.

How these were made

The raw-suite prompts were checked to be byte-identical to the paper's scripts. Scoring the paper's 24M ensemble reproduces its Fig. 4 to within a mean absolute difference of 0.002. Evaluations ran in bf16 on an A100. The paper's scripts used fp32; the difference was noise-level when checked.

Downloads last month
478

Paper for rihim/icl-selfplay-vs-text-results