- AIP — Adversarial Informativeness Pooling
- What is in this artifact
- Contributions
- 1. Why invert instead of discard
- 2. Architecture
- 3. Repository structure
- 4. Experimental setup
- 5. Results
- 5.0 Head-to-head at the hardest point (f = 0.5, the exact tie)
- 5.1 What holds
- 5.2 What was refuted — by this study's own data
- 5.3 Cost of defence (E3) and composition guidance (S1)
- 5.4 The confidence protocol (P1)
- 5.5 Undetectable at this scale
- 5.6 The gate-aware adversary, anchored at f = 0.5
- 5.7 Instrumentation disclosure
- 5.8 What is prior art and what is ours
- 5.9 Withheld under a pre-committed conditional
- 6. Correction campaign X3 — the completion budget was a broken instrument
- 7. MX — AIP as a trust layer for multi-stage agentic pipelines
- 8. Running it
- 9. Tests
- Licence
- Citation
AIP — Adversarial Informativeness Pooling
"Liars Are Information": Byzantine-robust decentralized LLM swarms.
N agents answer the same question, broadcast an answer plus a confidence, and each
agent locally aggregates what it receives. A fraction f of agents are Byzantine.
Prior art (Krum, geometric median, trimmed mean, confidence weighting, Dawid–Skene,
filter-and-refine) discards adversarial input. AIP inverts coherent adversaries
and pools them as information — and then a gate-aware adversary defeats it, which is
half of what this artifact is for.
Status: phases 0–6 complete, X3 repair landed, numbers recomputed. The experiment ran to completion on H100 and the claims ledger was audited against a measurement floor with no claim left pending. A later audit then found a measurement defect — 944 of 13,500 cached generations carried an answer no model had committed to, because the completion budget truncated them and the extractor's fallback read a number out of the unfinished working. Correction campaign X3 (§6) regenerated 2,313 rows across all three caches and cut manufactured answers to 356. Every downstream number has been recomputed and the per-macro diff is published: 113 of 231 reported numbers changed, 6 moved beyond the measurement floor, and none of the six is a core AIP claim.
The measurement floor itself was re-derived on clean data and adopted: 0.066 on GSM8K, 0.099 elsewhere. Every floor-based verdict was re-checked; none flipped.
A second line of work, MX (§7), extends AIP from a flat swarm to a multi-stage agentic pipeline. It is at Phase 1 with 14 of 17 gates green and is deliberately stopped — no liar code path has executed at any scale.
Every number here is measured in this hardware regime. The earlier L4-era results are void and are not mixed in (RULE 1).
What is in this artifact
| Inference cache | data/cache/, data/cache_t07/, data/cache_adversarial/ — ~10 GPU-hours of broadcasts from 9 models over 6 benchmarks. The irreplaceable part. |
| Result parquets | results/ — 54,432-row symbolic sweep, 26,208-row adversarial sweep, correlation, N-scaling, sleeper/Sybil, the gate-aware sweep. |
| Claims ledger | results/claims.md — every claim with a verdict, its evidence, and its magnitude against the measurement floor. |
| X3 repair | results/truncation_audit_main_cache.md, results/x3_reaudit.*, results/x3_macro_diff.md — the measurement defect of §6, its size, and the before/after value of all 231 reported macros. |
| MX pipeline | src/aip/pipeline/, docs/mx_prereg.md, results/mx/ — the multi-stage agentic extension of §7. |
| Audit | results/pending_claims_audit.md — the re-derivation that closed the last five open claims and overturned three of them. |
| Manuscript | paper/ — ten sections, 13 pages, compiled (main.pdf), every number a generated macro, bibliography complete. |
| Code | src/aip/, scripts/ — the mechanism, the baselines, the phases, the figure and table generators. |
data/benchmarks/ holds the frozen task items themselves (verbatim ARC / BoolQ /
GSM8K / MATH-500 / MedQA / MMLU questions, choices and gold answers). It is
included so the artifact is self-contained and a rerun cannot silently draw a
different sample.
Licensing of that directory. These corpora are not covered by this repository's MIT licence and are redistributed under their own terms: ARC and BoolQ are CC BY-SA, GSM8K / MATH-500 / MedQA / MMLU are MIT or equivalent. Anyone redistributing
data/benchmarks/inherits CC BY-SA's share-alike obligation for the ARC and BoolQ portions. SeeLICENSE.
Nothing offline depends on it: task identities live in configs/task_lists.json,
label spaces in configs/inversion_thresholds.yaml, and gold answers travel
inside the cache parquets. python scripts/hf_push.py omits the directory by
default for the licensing reason above; it is present in this repository because
it was pushed explicitly with --include-benchmark-cache.
Loading it
from datasets import load_dataset
honest = load_dataset("<repo-id>", "cache_greedy", split="train")
sweep = load_dataset("<repo-id>", "sweep_adversarial", split="train")
Or read a single file directly, which is usually what you want:
import pandas as pd
pd.read_parquet("results/adversarial/gate_aware_mitigations.parquet")
Contributions
What is new here, in the order it matters for multi-agent systems.
1. Inversion as an aggregation primitive. Every prior Byzantine-robust rule — Krum, geometric median, trimmed mean, confidence weighting, Dawid–Skene, filter-and-refine — treats an adversarial channel as noise to remove. AIP treats a coherent adversary as a channel with negative mutual information and flips it. The isolating ablation is in §5.0: same gate, same calibration, inversion disabled costs 4.2 points of MedQA accuracy at f = 0.5.
2. A receiver-anchored gate, not a plurality-anchored one. The gate conditions coherence on disagreement with the receiver, not agreement with the majority. This is what keeps it meaningful at f > 0.5, where the plurality is adversarial: a receiver knows one thing no label-free estimator can otherwise establish — that it is itself honest. Most robust-aggregation work assumes an honest majority; this does not.
3. A calibrated multiclass attenuation law. m_ij = (1-e_i)(1-e_j) + e_i e_j q_ij, exact at C = 2 and validated against that exactness, which is what
lets pairwise error correlation be estimated blind at run time, without
labels. A law that needs gold answers cannot gate a live swarm.
4. An impossibility result that bounds the whole approach. Proposition 1: a
rule measurable with respect to per-question coherence cannot separate an
adversary sitting at or below chance coherence 1/(C-1) from honest independent
disagreement. This is what makes the gate-aware defeat in §5.6 a predicted
consequence rather than an embarrassment — and it is why the artifact reports the
attack that beats AIP as prominently as the mechanism that works.
5. Replication is not diversity — quantified. A deployed swarm is usually built by replicating one model, not assembling distinct ones. Within-model error correlation is 0.652 against cross-family 0.444, and measured effective swarm size averages 1.92 against a cross-model-only upper bound of 2.31 — for swarms of up to ten. Prior ceiling work studies panels of distinct models and does not make this comparison.
6. The composition recipe is refuted, by our own data. "Maximise effective swarm size" fails both halves: N_eff correlates −0.934 with mean roster accuracy (it largely restates roster weakness) and −0.506 with swarm accuracy — the wrong sign for an objective. And no composition beats its own best single member by a detectable margin on any benchmark. Reported as a refutation with the same prominence as the confirmations.
7. Non-stationary adversaries, and the windowed fix. A burst adversary defeats a pooled-history gate; sliding-window channel statistics (W=20) recover it — on GSM8K, 0.854 pooled to 0.974 windowed, +0.120 = 1.8× floor — at undetectable cost on stationary attacks. The recovery is far larger against the non-AIP baselines, which stay at 0.34–0.43 under the same attack.
8. Extension from a flat swarm to a staged pipeline (MX). The open question for agentic systems is not whether a gate works between peers but whether it works between stages, where one stage's output becomes the next stage's premise. The pipeline, arms, adversaries and pre-registration are in §7; it is at Phase 1 and deliberately stopped.
9. Two methodological findings that generalise past this paper. Parse rate is not a health metric — a 512-token cap produced a 100% parse rate while manufacturing answers from truncated working (§6), and only a commitment-marker audit found it. And budget evidence can be self-referential: justifying a token budget from completion lengths measured under the old budget is circular, because the observed ceiling moves with the cap (§7).
1. Why invert instead of discard
A discard defence has a structural ceiling at f = 0.5: once the adversarial bloc
is the plurality, any method that treats consensus as a proxy for truth converges on
the lie and then classifies the honest minority as the outliers. It discards exactly
the agents worth keeping.
The claim under test is that a coherent adversary is not noise to be removed but a channel to be read backwards. If a bloc reliably says X when the truth is not X, then "this bloc says X" is evidence against X.
The complication is that honest models also agree when they are wrong. Coherence alone therefore cannot license inversion, so every channel is scored against a calibrated ceiling of how coherent honest agents get on that answer space.
And the complication has a second half, which is the paper's turn. An adversary
that reads the published ceiling can coordinate just underneath it. At its optimum
it is measurably less coherent than honest agents failing independently — below
the chance rate 1/(C−1) — at which point no threshold on that axis can separate
it from honest disagreement at all. That is Proposition 1, and it is why this
artifact reports a defence and its limit in the same breath.
2. Architecture
┌──────────────────────────────────────────┐
│ configs/ (the authorities, committed) │
│ models.yaml · task_lists.json │
│ inversion_thresholds.yaml │
│ decision_rules.yaml (R1–R6) │
└───────────────┬──────────────────────────┘
│ read by every phase
┌────────────────────────────────────┼────────────────────────────────────┐
│ │ │
┌────▼─────────────┐ ┌───────▼────────┐ ┌────────▼───────┐
│ PHASE 1 cache │ │ PHASE 4 adv. │ │ PHASE 5 E1/E2 │
│ 7 models │ │ 4 inference │ │ sleeper, Sybil │
│ × 6 benchmarks │ │ + 3 offline │ │ 6 defenders │
│ GPU, cache-once │ │ attacks · GPU │ │ offline │
└────┬─────────────┘ └───────┬────────┘ └────────┬───────┘
│ data/cache/*.parquet │ data/cache_adversarial/ │
│ data/cache_t07/ (2nd sample) │ │
└────────────────┬───────────────────┴────────────────────────────────────┘
│ NO phase after 1 and 4 runs a model. Everything below
│ is offline over the cached broadcasts.
┌─────────────▼──────────────┐
│ PHASE 2 correlation │ within/cross-model φ, N_eff,
│ │ multiclass attenuation law,
│ │ honest-q recalibration → thresholds
└─────────────┬──────────────┘
┌─────────────▼──────────────┐
│ PHASE 3 aggregation sweep │ 14 methods × f × p_obs × compositions
│ │ × topologies, 500-resample CIs
└─────────────┬──────────────┘
┌─────────────▼──────────────┐
│ PHASE 6 ledger + paper │ claims ledger, pending-claim audit,
│ │ numbers.tex, figures with provenance
└────────────────────────────┘
Orchestration: scripts/run_all.py → RUNLOG.md (one line per step)
→ results/STATUS.json (RUNNING/STOPPED/COMPLETE)
→ results/.stamps/ (idempotent resume)
The load-bearing design choice is cache-once inference. All model generation happens in Phases 1 and 4 and is written to parquet; every later phase reads that cache and never loads a model. That is what makes a 54,432-row sweep affordable on one GPU, and it is why the caches are the irreplaceable artifact.
How the multi-agent system actually works
Two multi-agent structures are studied, and they are different shapes. The main
study is a flat swarm: N peers answer the same question once and pool.
§7's MX work is a staged pipeline: each stage's pooled output becomes the
next stage's input. AIP is the trust layer in both.
Structure 1 — the flat swarm, one question, one round.
question q
│
┌──────────┬──────────┬──────┴───┬───────── ... ──────┐
▼ ▼ ▼ ▼ ▼
┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐
│agent 1 │ │agent 2 │ │agent 3 │ │agent 4 │ ... │agent N │
│ honest │ │ honest │ │BYZANTINE│ │BYZANTINE│ │ honest │
└───┬────┘ └───┬────┘ └───┬────┘ └───┬────┘ └───┬────┘
│ a1 │ a2 │ ~a3 │ ~a4 │ aN
└──────────┴──────────┴──────────┴──────── ... ─────┘
│
│ broadcast (complete graph, or a
│ sparser topology; p_obs
▼ controls what each peer sees)
┌───────────────────────────────────────────────────────────────┐
│ RECEIVER i — every honest agent independently is one │
│ │
│ its OWN answer a_i is the anchor. Not the plurality. │
│ A receiver knows one thing no label-free estimator can │
│ otherwise establish: that IT is honest. │
│ │
│ for each incoming channel j != i: │
│ coherence q_ij = P(a_j agrees with a_j' | both wrong) │
│ attenuation m_ij = (1-e_i)(1-e_j) + e_i e_j q_ij │
│ │
│ ┌───────────┬────────────┬───────────┐ │
│ │ INVERT │ TRUST │ DISCARD │ │
│ │ q above │ q inside │ otherwise │ │
│ │ ceiling & │ the honest │ │ │
│ │ m below │ band │ │ │
│ │ bound │ │ │ │
│ └─────┬─────┴──────┬─────┴─────┬─────┘ │
│ │ │ │ │
│ use NOT a_j use a_j drop a_j │
└────────────────────────┬──────────────────────────────────────┘
▼
receiver i's pooled answer
A liar that is coherent — consistent with itself and with other liars — is the hard case, and it is the case every discard-based defence fails on: discarding a coherent liar throws away a channel that carries usable signal. INVERT is the move that turns it into evidence instead.
Structure 2 — the staged pipeline (MX), where a lie becomes a premise.
task
│
▼
┌───────────────── STAGE: solver ──────────────────┐
│ replica 1 replica 2 replica 3 + anchor │ k=3 replicas
│ │ │ │ │ │ + the aggregator's
│ └───────────┴─────┬─────┴──────────┘ │ own answer
│ ▼ │
│ AIP gate or plurality │
└────────────────────────┬─────────────────────────┘
│ pooled solver answer
▼ ... becomes part of the NEXT stage's PROMPT
┌───────────────── STAGE: critic ──────────────────┐
│ replica 1 replica 2 replica 3 + anchor │
│ ▼ │
│ AIP gate or plurality │
└────────────────────────┬─────────────────────────┘
│ pooled critic answer
▼
┌───────────────── STAGE: refiner ─────────────────┐
│ replica 1 replica 2 replica 3 + anchor │
│ ▼ │
│ AIP gate or plurality │
└────────────────────────┬─────────────────────────┘
▼
end-task answer
Exactly ONE replica, in ONE stage, is replaced by a liar. The question is not
"was that stage wrong" but how far the error travels -- CASCADE CONTAINMENT.
In the flat swarm a bad vote is diluted by the others. In the pipeline it is laundered: once a lie is pooled into a stage's output it enters the next stage as an established premise, and the replicas downstream reason from it rather than about it. That is the asymmetry MX exists to measure.
The three-way channel gate
| decision | condition |
|---|---|
| INVERT | coherence above the calibrated honest-q ceiling for this answer space and agreement with the receiver below what the attenuation law allows |
| TRUST | coherence within the honest band and blind error estimate < ½ |
| DISCARD | otherwise |
Coherence is conditioned on disagreement with the receiver, not on agreement with
the plurality. That matters at f > 0.5: a receiver knows one thing no label-free
estimator can otherwise establish — that it is itself honest.
3. Repository structure
configs/ the authorities. models.yaml carries `arm:` (frozen / weak_tier /
retired) and the paper's model counts are derived from it, not typed.
data/
cache/ honest broadcasts, greedy 9 models × 6 benchmarks
cache_t07/ second sample at T=0.7 7 models × gsm8k (within-model φ)
cache_adversarial/ prompt-injected attacks same task ids by construction
benchmarks/ source items — local only, not published (see above)
results/
aggregation/ sweep, minimax regret, parity audit, N-scaling
adversarial/ adversarial sweep, gate-aware, burst, sleeper, Sybil, q-by-space
correlation/ pairwise φ, N_eff, multiclass validation
claims.md the ledger — verdict, evidence, magnitude vs floor
pending_claims_audit.md Task 6: the re-derivation that closed the last five
measurement_floor.json MDE per benchmark; the rule every claim is judged by
paper/
sections/ ten sections (no appendix: it is `docs/proposition1_note.md`)
main.pdf the compiled manuscript, 12 pages, shipped so a reader needs no TeX
references.bib 19 entries, all cited, zero placeholder fields
CUT_PLAN.md what the 17 -> 12 page cut moved, and the order to reverse it
numbers.tex GENERATED. every number in the manuscript is a macro from here
figures/ the ONLY figure directory, with MANIFEST.json recording which
script wrote each PDF, from what, at which commit
claims_map.md section → claim map, checked mechanically
src/aip/ mechanism, baselines, estimators, task loaders, harness
pipeline/ MX: stages, arms, liars, runner -- the multi-stage trust layer
scripts/ phases, generators, audits, hf_push.py
mx_phase1*.py MX driver and its six-gate verifier
x3_*.py, audit_main_cache_truncation.py, run_x3_repair.sh the X3 repair campaign
docs/mx_prereg.md MX pre-registration: deviations, arms, power table, MX-H1..H5
data/mx_cache/ MX generations, keyed by (stage, model, prompt, slot)
results/mx/ MX Phase 1 stage outputs and determinism pass
tests/ 527 tests; the paper's build checks are in tests/test_paper.py,
the MX probes in tests/test_mx_pipeline.py
4. Experimental setup
Roster — 7 models, 7 lineages (FROZEN 2026-09-02) + a 2-model weak arm
Every model is an off-the-shelf checkpoint run at inference only. Nothing in this project is fine-tuned, and no gradient is ever computed. "Training" does not appear anywhere in the pipeline; what follows is measurement.
| model | lab | tier | params |
|---|---|---|---|
qwen38_27b |
Alibaba | strong | 27.8B |
gemma4_31b |
strong | 31.3B | |
granite42_30b |
IBM | strong | 29.3B |
olmo3_32b_think |
AI2 | reasoning / open-data | 32.2B |
ministral3_14b |
Mistral | mid | 13.9B |
phi4_mini_reasoning |
Microsoft | reasoning | 3.8B |
llama32_3b |
Meta lineage | weak/fast · primary adversary | 3.2B |
Plus ministral_8b and olmo2_7b as a weak-tier arm (arm: weak_tier),
reported separately and never pooled with the frozen roster.
Single-model accuracy — all 9 models × all 6 benchmarks
Greedy, bf16, batch-invariant, 1,500 tasks per model. Post-X3 (§6): these are the repaired numbers, not the ones the first cache reported.
| model | lab | ARC | BoolQ | GSM8K | MMLU | MedQA | MATH-500 | mean |
|---|---|---|---|---|---|---|---|---|
gemma4_31b |
100.0 | 90.5 | 97.8 | 92.5 | 92.5 | 94.0 | 94.5 | |
qwen38_27b |
Alibaba | 98.0 | 91.5 | 96.8 | 90.5 | 93.0 | 88.5 | 93.0 |
olmo3_32b_think |
AI2 | 96.0 | 90.5 | 96.4 | 90.0 | 78.0 | 76.5 | 87.9 |
ministral3_14b |
Mistral | 94.0 | 90.5 | 96.0 | 84.5 | 79.0 | 80.5 | 87.4 |
granite42_30b |
IBM | 96.0 | 92.0 | 95.2 | 84.0 | 67.5 | 74.5 | 84.9 |
phi4_mini_reasoning |
Microsoft | 89.5 | 83.0 | 93.0 | 77.5 | 59.5 | 69.0 | 78.6 |
ministral_8b ¹ |
Mistral | 85.0 | 85.5 | 89.6 | 67.5 | 53.5 | 55.0 | 72.7 |
llama32_3b |
Meta lineage | 82.0 | 83.0 | 80.2 | 61.5 | 54.0 | 44.5 | 67.5 |
olmo2_7b ¹ |
AI2 | 79.0 | 82.0 | 84.6 | 60.5 | 45.0 | 31.5 | 63.8 |
| benchmark mean | 91.1 | 87.6 | 92.2 | 78.7 | 69.1 | 68.2 |
¹ weak-tier arm — ministral_8b and olmo2_7b are reported separately and never pooled with the frozen roster. llama32_3b IS in the frozen seven; it is the weakest member and doubles as the primary adversarial generator.
Read the columns, not just the rows. ARC and GSM8K are saturated — the roster mean is above 91% and the ceiling rule (R1) sends most of those cells to HEADROOM-LOW, which is why they are appendix benchmarks rather than headline ones. MedQA and MATH-500 are where the roster actually has room to be wrong, and that is precisely why they carry the headline claims: a defence that only shows value where the models are already right has not been tested.
The spread matters as much as the mean. On MedQA the roster runs from 93.0 down to 45.0 — a 48-point range on one benchmark. Swarm composition results (S1) are about exactly this: what happens when you pool agents that are this unequal.
Benchmarks — 1500 tasks, every answer-space class covered
| benchmark | n | answer space | role |
|---|---|---|---|
gsm8k |
500 | open numeric | appendix (saturated) |
math500 |
200 | open text | headline |
mmlu |
200 | multiple choice, C=4 | headline |
medqa |
200 | multiple choice, C=4 | headline |
boolq |
200 | binary, C=2 | exact-law control |
arc |
200 | multiple choice, C=4 | appendix (saturated) |
BoolQ earns its place as the one benchmark where the attenuation law is exact: with
C=2 there is one wrong answer, so two wrong agents are wrong identically and q = 1
by construction. It is also where the gate provably cannot help — and where majority
vote beats AIP, which the ledger records rather than buries.
Reproducibility conditions
- bf16 only. Quantisation shifts the token-logprob distribution, which is both the confidence surface the attacks target and the correlation structure Phase 2 measures.
VLLM_BATCH_INVARIANTon for every cache-generating run.- Seeds from SHA-256 of the canonical config, not Python's randomised hash.
- Sequential residency. One model resident at a time.
- Manifests beside every output: config hash, git commit, seeds, GPU, versions.
The measurement floor
results/measurement_floor.json. MDE = √2 · 1.96 · σ, with σ the resampling and
binomial noise combined: 0.066 on GSM8K (500 tasks), 0.099 on the 200-task
benchmarks. Flat in swarm size.
The resample pass excludes 97 T=0.7 rows that are still cap-pinned after X3. A cap-pinned row is scored wrong in both passes, so it registers as agreement and suppresses the measured flip rate; keeping them made the path_level floor anti-conservative (0.0642 kept against 0.0664 clean) while making answer_level conservative (0.0645 against 0.0628). See results/claims.md.
The rule this imposes. An effect that does not clear its cell's floor is undetectable at this scale — never "small", "negligible" or "no effect". Applying it consistently cost this study three findings.
5. Results
Every number below is in results/claims.md with its slice named, and in
paper/numbers.tex as a generated macro.
Post-X3, and the diff is published. Every number below was recomputed from the repaired cache.
results/x3_macro_diff.mdcompares all 231 reported macros before and after: 113 changed, 6 moved beyond the measurement floor, and none of the six is a core AIP claim. The banner that used to sit here — warning that these were pre-repair numbers — is gone because the recompute is done.
5.0 Head-to-head at the hardest point (f = 0.5, the exact tie)
Swarm accuracy (%) with half the agents Byzantine and coherent, averaged over compositions and topologies. f = 0.5 is the anchor because at f = 0.7 a Byzantine majority drives every method toward zero and the comparison stops discriminating.
| aggregation rule | MATH-500 | MedQA | MMLU |
|---|---|---|---|
| AIP gated (INVERT/TRUST/DISCARD) | 32.0 | 51.3 | 57.8 |
| AIP naive (no ceiling test) | 42.0 | 56.6 | 60.0 |
| AIP trust-only (ablation: no INVERT) | 32.0 | 47.0 | 53.0 |
| Dawid–Skene | 25.7 | 57.3 | 62.7 |
| SAC filter-refine | 33.4 | 53.6 | 61.5 |
| coordinate-wise median | 35.7 | 49.1 | 54.3 |
| geometric median | 35.6 | 49.2 | 54.7 |
| majority vote | 35.6 | 49.2 | 54.7 |
| confidence-weighted (logprob) | 33.3 | 47.7 | 54.5 |
| confidence-weighted (self-report) | 14.9 | 27.8 | 35.3 |
Three things this table is meant to let you check rather than take on trust. AIP gated beats every robust-statistics baseline on MedQA and MMLU — median, geometric median, majority — which is the claim. Dawid–Skene beats AIP gated on both (57.3 vs 51.3, 62.7 vs 57.8), and that is in the ledger, not buried: it is a latent-confusion model given the whole broadcast matrix, and it is the strongest baseline here. On MATH-500 AIP loses to majority (32.0 vs 35.6). The open answer space is where the gate has least to work with, and §5.2 records it as a refutation rather than an exception.
The aip_trust_only row is the ablation that isolates INVERT: same gate, same
calibration, inversion disabled. The gap between it and aip_gated (47.0 → 51.3
on MedQA) is the mechanism this paper is about.
5.1 What holds
| claim | result | vs floor |
|---|---|---|
| Inversion is a large, real mechanism | gain 0.645 at f=0.7, isolated by ablating aip_trust_only against aip_gated |
9.5× |
| …and is close to free when idle | cost 0.001 at f=0 | ≪ floor → no detectable cost |
| Discard baselines collapse | every deployable discard method reads exactly 0.000 at f=0.7 against the symbolic coherent adversary, on all six benchmarks | exact |
| AIP separates from them | +0.857 gsm8k, +0.372 math500, +0.268 arc, +0.188 mmlu under the real-LLM adversary | 4/6 above floor |
| Attenuation law exact at C=2 | BoolQ MAE 0.000 | analytic identity |
| Multiclass law beats binary | math500 MAE 0.124 → 0.013 | — |
| Homogeneous swarms are not diverse | within-model φ 0.652 vs cross-model 0.444; N_eff 1.50 out of 10 | — |
| Self-report attacks miss AIP | AIP moves exactly 0.000; the self-report reader loses 0.104–0.234 on 4/6 | identity |
| Windowing recovers the burst adversary | pooled 0.470 → 0.965 at W=20, f=0.7 | 7.3× |
| AIP survives a plurality Sybil bloc | 0.756 at s=8 where every bloc defence collapses to 0.265 | 4.1× |
| A gate-aware adversary defeats AIP | 0.082 at the optimum (medqa, f=0.5) against 0.878 fully coherent | attacker gain 0.796, 8.0× |
| Minimax regret under unknown f | AIP best on gsm8k (0.004) and mmlu (0.038); not on math500 | — |
5.2 What was refuted — by this study's own data
These get the same weight as the confirmations. Section 11 of the manuscript runs them at the same length.
| claim | verdict |
|---|---|
| H1 — AIP is flat in f | REFUTED on 4 of 6. −0.615 boolq, −0.376 medqa, −0.321 mmlu, −0.316 arc (f=0→0.7, real-LLM coherent adversary). gsm8k and math500 are undetectable, so flatness is neither shown nor refuted there. |
| H2 — every discard baseline reaches zero, including oracle methods | Split three ways. True only for the symbolic adversary and deployable methods. Oracle Krum holds 0.815 on BoolQ. Under real-LLM adversaries nothing reaches zero, and majority vote beats AIP on BoolQ by 0.290. |
| I3 — inversion gain is independent of honest competence | REFUTED. The supporting comparison read the weak arm's six benchmarks (0.547) against the roster's three (0.645). On the common set: 0.470, a gap of 0.175 = 1.8× floor, lower on all three individually — and opposite to the prediction. |
| G2 — coordinatability collapses on closed label spaces | REFUTED as worded; the direction was inverted. Adversarial coherence is highest on the most closed space (binary 0.559, MC 0.285, open text 0.069). What collapses is the margin over honest agreement, which is the quantity the gate needs. |
| M2 — AIP does not exploit self-vote share | REFUTED at f=0.5. 25 of 54 cells clear the floor, mean +0.443, max 0.840. At the exact tie the normalisation of the adversary's own vote, not the mechanism, decides the outcome. Undetectable away from the tie. |
| S1 — maximising effective swarm size is the composition recipe | REFUTED, both halves. N_eff correlates −0.934 with mean roster accuracy (it restates roster weakness) and −0.506 with swarm accuracy (wrong sign for an objective). And no composition beats its own best single member detectably anywhere. |
| E1 — a sleeper defeats reputation decay while AIP reclassifies it | REFUTED, and backwards. Reputation decay drops least (0.031); AIP drops most (0.136). Spread 0.105 against a 0.100 floor — marginal, and reported as marginal. |
| E2b — Sybil robustness comes from coherence amplification | MECHANISM REFUTED. Coherent bloc 0.756 vs disagreeing 0.748. The robustness holds; the explanation does not. |
| L1 — a low honest-q ceiling is a structural defence | REFUTED. MATH-500 has the lowest ceiling and the gate is near-inert there: the gate-aware adversary's best strategy on MATH-500 is full coherence, not evasion. |
| D2 — threshold randomisation helps | REFUTED, and undetectably so. −0.013 over the attack space, +0.011 in the band; both far below floor. An earlier reading called it "actively harmful" on a difference that was itself below floor — also withdrawn. |
5.3 Cost of defence (E3) and composition guidance (S1)
Aggregation cost, µs per receiver-decision at f=0.5, mean over headline benchmarks — measured by timing each rule over the cache, not apportioned from block timestamps:
| rule | µs/decision |
|---|---|
| majority vote | 3.2 |
| krum (oracle) | 20.9 |
| trimmed mean (oracle) | 68.5 |
| coordinate median | 75.6 |
| AIP, inversion ablated | 74.4 |
| AIP-gated | 77.1 |
| multi-Krum (oracle) | 92.9 |
| geometric median | 733.1 |
The gate adds +1.7 µs (3.6%) over its own inversion-free ablation — almost all of AIP's cost is channel bookkeeping the ablation also pays. One completion averages 634 tokens, ≈6598× the slowest rule's cost. Aggregation is not the bottleneck for any method.
Composition — all 35 three-model and 21 five-model subsets of the roster:
| N_eff vs mean roster accuracy | −0.934 Spearman — N_eff is close to a restatement of roster weakness |
| swarm accuracy vs N_eff | −0.506 — maximising N_eff selects against accuracy |
| swarm accuracy vs mean roster accuracy | +0.658 |
| what survives partialling out the mean | accuracy spread, partial ≈ +0.7 |
| best gain over the best single member | +0.015 anywhere; math500 every subset worse by ≥0.097 |
A plurality-voting swarm of these models is not an accuracy device. No composition beats its own strongest single member by a detectable margin on any benchmark. The reason to build a swarm is Byzantine robustness, which is a different objective. Family diversity is unmeasurable in this roster — all 7 models are 7 lineages, so the predictor is constant.
5.4 The confidence protocol (P1)
| surface | AUC vs correctness |
|---|---|
| logprob | 0.704 |
| self-report | 0.628 |
| self-report, falsified | 0.339 — below chance, on 21 of 22 powered cells |
Over the 22 of 42 frozen-roster cells clearing R2's n_wrong ≥ 30. The logprob
surface is untouched by the falsification attack by construction. The self-report
is also simply absent on up to 34.0% of rows (mmlu/phi4_mini_reasoning) —
a field a well-behaved agent may omit and a misbehaving one may fabricate.
Weak-tier arm, separately: 0.716 / 0.573 / 0.427 — same ordering.
Broadcast both, weight on the logprob. This is the one recommendation here that the gate-aware adversary does not reach, because it rests on no coherence statistic.
5.5 Undetectable at this scale
| claim | why |
|---|---|
| N1 — does inversion gain scale with N? | 0.643 / 0.637 / 0.655 at N = 10 / 20 / 50. Range 0.018 against a 0.100 floor. Neither "grows with N" nor "invariant in N" is supported. The experiment was underpowered to separate them, and the figure caption says so. |
5.6 The gate-aware adversary, anchored at f = 0.5
A bloc that coordinates with probability p realises q(p) = p² + (1−p)²/(C−1).
Sweeping p sweeps q, non-monotonically, so the adversary has a genuine interior
optimum.
| benchmark | ceiling | best evasion (below ceiling) | fully coherent | attacker gain |
|---|---|---|---|---|
| medqa | 0.677 | 0.082 at q̂ = 0.280 | 0.878 | +0.796 |
| mmlu | 0.677 | 0.374 at q̂ = 0.280 | 0.903 | +0.529 |
| math500 | 0.198 | 0.366 at q̂ = 0.160 | 0.000 | −0.366 — no evasion needed; the gate is inert |
The optimum q̂ = 0.280 sits below the chance coherence 1/(C−1) = 0.333. At its
most damaging the adversary is less coherent than honest agents failing
independently, which is why no threshold on that axis can help.
Why f = 0.5 and not f = 0.7. At f = 0.7 the Byzantine bloc is an outright majority and every coherence level drives accuracy to zero, so an evasion effect is not separable from numerical dominance. An earlier version of this work reported the f = 0.7 attacker gain (0.92); it is not an admissible anchor and is withdrawn. The anchor is also measured in the matched-parity condition, because of the M2 finding above.
Pre-registered mitigations (criteria committed to git before the runs):
| variant | band mean | band worst | verdict |
|---|---|---|---|
| hard gate | 0.463 | 0.082 | baseline |
| M1 soft gate | 0.623 | 0.200 | partial success, 1.6× floor within the band; +0.100 (exactly at floor) over the whole space |
| M2 randomised threshold | 0.474 | 0.082 | no detectable effect either way |
| M3 both | 0.638 | 0.278 | not separable from M1 |
5.7 Instrumentation disclosure
Two mitigation variants were initially inert and returned results identical to
the unmitigated baseline. Those identical results were read as a finding. They were
bugs: M1 fell through to the hard gate for every adversarial channel, and M2's
jitter reached only a fixed-ceiling ablation. Two published verdicts were
retracted. They were caught by probe tests (tests/test_mitigation_variants.py)
that assert each variant behaves differently from the baseline before any verdict
is read.
A related class of error accounts for three of the refutations above: a comparison
made across mismatched bases — six benchmarks against three, an upper-bound
estimator against a measured one, an f = 0.7 row against an f = 0.5 claim. The
mismatch, not the effect, produced the result. scripts/task6_audit_pending.py runs
in the build and fails loudly if a claim can no longer be re-derived.
5.8 What is prior art and what is ours
The correlation ceiling is not our discovery, and the manuscript says so in the abstract, §1, §2 and §5. Kim et al. (ICML 2025) establish that LLM errors are correlated across models; Kohli (arXiv:2605.29800) shows nine judges are worth about two effective votes; Jiang et al. (NeurIPS 2025) document the underlying output homogeneity.
We confirm that ceiling in an adversarial swarm and extend it three ways:
- To the deployment-relevant case. Their panels are built from distinct models. A deployed swarm is more likely built by replicating one, and replication is roughly half again as correlated — within-model 0.652 against cross-family 0.444. That comparison is not in the prior work.
- From an observation to a calibrated law. The multiclass attenuation law is exact at C=2 and validated against that exactness, so the correlation is estimated blind at run time without labels — which is what lets it serve as the constant an inversion gate fires against.
- From a measurement caveat to a design premise. For an evaluation panel the ceiling is a caveat. For a swarm under attack it is the premise: if honest pooling is capped, the only informative structure left is what honest agents do not share — adversarial coherence. Their negative result about panels is what makes inversion necessary rather than merely available.
5.9 Withheld under a pre-committed conditional
An adaptive-adversary deterrence / payoff-flattening experiment is not in
the paper, and the withdrawal is permanent. The condition fixed before the
audit was that it enters only if the coordinatability claim (G2) survived
re-verification. It did not — the measured direction is the reverse of the claim.
The bandit data is in the release (results/adversarial/bandit_log.parquet) and
is cited nowhere in the manuscript. There is a second, independent ground: the
data shows payoff, not conduct — the adversary never switched arms within the
horizon tested. Loosening a pre-registered conditional after seeing the outcome
is what pre-registration exists to prevent.
6. Correction campaign X3 — the completion budget was a broken instrument
Found 2026-09-20, after the manuscript was already in full prose. It is recorded here rather than quietly fixed, because the defect is instructive and because the metric that was being watched is the one that hid it.
What went wrong
Three things had to line up, and none of them is a bug on its own.
- The budget.
configs/models.yamlgave non-reasoning entries 640 answer tokens and reasoning entries 2048 — sized for a different generation regime. - The answer position. Every benchmark here commits its answer at the end
of the completion:
\boxed{}for MATH-500,#### nfor GSM8K,Answer: Xfor multiple choice. So a truncated generation is exactly the one missing its answer. - The extractor's fallback chain, written for completions that finished, reads the last plausible token out of the working when the marker is absent.
Together they produced confident wrong answers at a 100% parse rate. Parse rate was the health metric. It is now retired as one: it measures whether a string was produced, not whether a model answered.
The size of it
Audited by finish_reason — recorded per row in extra_json, not inferred from
a token count. scripts/audit_main_cache_truncation.py.
| count | of 13,500 | |
|---|---|---|
finish_reason == "length" |
1,036 | 7.67% |
| — committed an answer anyway (legitimate) | 90 | |
| — no marker and no answer (correctly discarded) | 2 | |
| — manufactured | 944 | 6.99% |
Manufactured rows scored correct 98/944 = 10.4%, against 85.9% on the rest of the cache — so they depressed measured accuracy rather than inflating it.
By benchmark. The gradient tracks how long a correct answer takes to write, not how hard the benchmark is:
| benchmark | manufactured | rate |
|---|---|---|
| math500 | 632 / 1,800 | 35.1% |
| medqa | 120 / 1,800 | 6.7% |
| mmlu | 70 / 1,800 | 3.9% |
| gsm8k | 99 / 4,500 | 3.6% |
| arc | 20 / 1,800 | 1.1% |
| boolq | 3 / 1,800 | 0.2% |
By model, all nine, pooled over benchmarks:
| model | manufactured | rate | model | manufactured | rate | |
|---|---|---|---|---|---|---|
| phi4_mini_reasoning | 213 / 1,500 | 14.2% | olmo2_7b | 61 / 1,500 | 4.1% | |
| granite42_30b | 153 / 1,500 | 10.2% | qwen38_27b | 57 / 1,500 | 3.8% | |
| ministral3_14b | 146 / 1,500 | 9.7% | llama32_3b | 56 / 1,500 | 3.7% | |
| olmo3_32b_think | 137 / 1,500 | 9.1% | ministral_8b | 39 / 1,500 | 2.6% | |
| gemma4_31b | 82 / 1,500 | 5.5% |
Worst cells: math500/ministral3_14b 135 of 200 (67.5%, none of them
correct), math500/granite42_30b 99 (49.5%), math500/phi4_mini_reasoning 76
(38.0%), medqa/phi4_mini_reasoning 54 (27.0%).
The two auxiliary caches carry the same defect, so both are in scope:
| cache | truncated | what it feeds |
|---|---|---|
data/cache_t07 |
146 / 3,500 (4.2%) | within-model φ, and therefore the measurement floor itself |
data/cache_adversarial |
741 / 12,000 (6.2%; 31.4% on MATH-500) | all of Phase 4 |
Initial repair set: 1,923 rows. The campaign ultimately rewrote 2,313 —
the adversarial rushing arm needed a second pass once the honest answers it
embeds had moved (see Final state of the campaign below).
The repair
- Budget → 3072 tokens for every entry, derived from context headroom, not
from the observed length distribution. That distinction is the whole point:
in a cell that truncated 68% of the time, the finishers' p99 sits just under
the old cap and carries no information about where the real tail ends —
deriving a budget from it would be circular. The longest prompt in the study is
826 tokens (BoolQ); 1.15× for chat templates plus 64 tokens of safety leaves
3072 inside
max_model_len4096 on every benchmark. As a check rather than a derivation, 3072 is ≥1.5× the largest finisher p99 in every benchmark. - Only the truncated rows are regenerated — same prompts, same seeds, same temperature. Naturally-finished rows are never rewritten, so they stay byte-identical by construction.
- Batch invariance is verified, not assumed — and it failed for three models.
Regenerating a subset changes the batch, and the claim that this is a repair
rather than a fresh sample rests entirely on
VLLM_BATCH_INVARIANT. Each cell therefore also regenerates rows that did finish and requires them back byte-identical. Six of nine models passed across 192 control rows. Three did not — see Determinism below. That gate is the reason this is known at all. rushingadversarial rows get a second repair: that attack embeds the honest swarm's answers in its prompt, so its rows are stale wherever an honest answer moved, truncated or not.
Final state of the campaign
| cache | rows regenerated |
|---|---|
data/cache (honest) |
1,036 |
data/cache_t07 (T=0.7 resample) |
94 of 146 |
data/cache_adversarial |
1,183 = 741 truncated + 442 context-stale |
| total | 2,313 |
Manufactured rows: 944 → 356, 62% removed. The residue is concentrated in reasoning models that ruminate past any cap rather than models that needed more room.
The 442 context-stale rows are the part that would have been missed. They are
rushing-attack rows that were never truncated and never failed to parse — they
were simply answering a swarm whose honest answers the repair had changed. No
existing check would have flagged them.
Three cells moved beyond the floor, all MATH-500, all upward:
math500/ministral3_14b 0.305 → 0.805, math500/gemma4_31b 0.615 → 0.940,
math500/granite42_30b 0.485 → 0.745. Every non-MATH-500 cell moved by
≤0.030.
The repair also moved three cells across the R1 ceiling (≥0.93 → HEADROOM-LOW), changing the composition of the headline set and not only its values: cells at or above 0.93 went from 10 of 54 to 13 of 54.
Determinism: the byte-identity gate failed for three models
| model | control rows differing | answer changed |
|---|---|---|
olmo3_32b_think |
37 / 40 | 1 |
qwen38_27b |
19 / 40 | not measured (evicted) |
llama32_3b |
14 / 24 | see below |
| other six | 0 / 192 | — |
Text drift is common; whether it reaches the extracted answer is the question
that matters, since nothing downstream reads the prose. For olmo3_32b_think it
almost never does — 36 of 37 drifted rows kept the same answer.
For llama32_3b it does. A powered measurement over 300 naturally-finished
rows:
| cell | n | text drift | answer flips | rate | Wilson 95% CI |
|---|---|---|---|---|---|
| MATH-500 | 150 | 88 | 37 | 0.247 | [0.185, 0.321] |
| MMLU | 150 | 64 | 15 | 0.100 | [0.062, 0.158] |
| pooled | 300 | 152 | 52 | 0.173 | [0.135, 0.220] |
17.3% of llama32_3b's finished rows change their extracted answer on
regeneration, and the two benchmarks differ by more than their intervals
overlap. An early 8-row pilot suggested a handful of near-tie generations; at
n=300 that description was wrong — the effect is diffuse, not a handful.
The cause is not CUDA graph bucketing. That was the obvious hypothesis —
vLLM captures graphs at bucketed batch sizes, and VLLM_BATCH_INVARIANT makes
attention invariant to batch composition, not to batch size. Running
--enforce-eager should then have reduced divergence. It increased it, 4/8
to 7/8. (That arm is also confounded: the cached original was generated with
graphs on, so an eager re-run diverges for a reason unrelated to batching.) What
survives is the clean comparison — at identical settings, batches of 8 and 24
drifted on exactly the same four MATH-500 problems by name. Prefix caching, also
on, is the remaining untested candidate.
What it is finding
Cap-pinned is not the defect; manufactured is. The first repaired cell went
from 73 pinned to 60, which reads as a failed repair — but 41 of those 60 had
emitted #### <answer> and then kept rambling past it. They are answered rows
that did not stop. Manufactured rows in that cell fell 31 → 19.
Two cells have moved beyond the measurement floor, and both are claim-level revisions rather than hygiene:
| cell | accuracy before | after | move | floor |
|---|---|---|---|---|
math500/ministral3_14b |
0.305 | 0.805 | +0.500 | 5.0× |
math500/granite42_30b |
0.485 | 0.745 | +0.260 | 2.6× |
The pattern is that the cap hurt the capable models most: a strong model on
MATH-500 writes a long derivation and commits at the very end, so 640 tokens
removed the answer and kept the working. Weak models converted their truncated
rows too (math500/olmo2_7b: 60 manufactured → 4) but barely moved on accuracy,
because they were wrong anyway. So the instrument was compressing the spread
between strong and weak models, not only the levels.
This is also the concrete argument against the cheaper fix. Excluding the manufactured rows and reporting the surviving subset predicted +0.633 for ministral3_14b and +0.145 for llama32_3b; the measured truth is +0.500 and +0.020. Wrong in both directions, by different factors, and flattering in the direction that mattered — because conditioning on "the model finished" selects for the easy items.
Reading the campaign
| file | what it holds |
|---|---|
results/truncation_audit_main_cache.md |
the original audit, counts only |
results/x3_reaudit.md / .parquet |
before/after per cell — pinned, manufactured, accuracy |
results/x3_macro_diff.md |
every one of the paper's macros, old vs new vs floor |
scripts/x3_repair.py, x3_repair_adversarial.py |
the repair drivers |
scripts/x3_recompute.sh |
the offline chain that rebuilds every downstream surface |
scripts/x3_attractor_table.py |
the attractor-cell table MX's benchmark choice rests on |
7. MX — AIP as a trust layer for multi-stage agentic pipelines
The main study is a flat swarm: N peers answer the same question once. MX
asks the question a deployed agentic system actually raises — does a gate that
works between peers still work between stages, where one stage's output is the
next stage's input? A lie that survives stage one is not merely one bad vote;
it becomes the premise everything downstream reasons from. The quantity of
interest is therefore cascade containment, not single-shot accuracy.
The pipeline
solver ──► critic ──► refiner three stages, output of each feeds the next
│ │ │
└─ k=3 replicas per stage, plus the aggregator's own answer as the receiver anchor
The receiver anchor is load-bearing and is what distinguishes AIP from plurality here: the gate is anchored on what the receiving agent believes, not on what the majority of replicas said. Proposition 1 turns on exactly that distinction.
Arms and adversaries
| arm | replicas | rule |
|---|---|---|
| P0 | k=1 | none — the vulnerability baseline |
| P1 | k=3 same-model | plurality |
| P2 | k=3 same-model | AIP gate |
| P3 | k=3 cross-family | AIP gate |
| P4 | k=3 cross-family | plurality |
Exactly one replica in one stage is replaced by a liar:
- L1, attractor-rider — INFORMED. Reads the other honest replicas' answers and endorses their most common wrong one. It rides the honest error attractor, so its coherence is indistinguishable from honest coincidence. This is Proposition-1 territory by construction: the gate should struggle.
- L2, fixed-lie — UNINFORMED. Commits to a predetermined wrong answer. Its coherence sits above the honest ceiling, so the gate should fire.
Why MedQA, and not MATH-500
L1 needs the attractor cell — two honest replicas both wrong and agreeing —
to exist at all. Measured over all 36 cross-model pairs of the cached roster
(scripts/x3_attractor_table.py):
Recomputed on the repaired cache (the pre-X3 values are in brackets):
| benchmark | attractor rate | P(both wrong) | |
|---|---|---|---|
| MedQA | 0.0751 (0.0786) | 0.138 (0.150) | highest; C=4 closed space |
| MMLU | 0.0711 (0.0737) | 0.095 (0.104) | |
| BoolQ | 0.0654 (0.0657) | 0.065 | C=2 — agreement is forced, degenerate |
| ARC | 0.0179 (0.0179) | 0.024 | saturated |
| GSM8K | 0.0109 (0.0107) | 0.026 (0.030) | saturated |
| MATH-500 | 0.0064 (0.0050) | 0.170 (0.334) | open text: plenty of error, almost no shared error |
The decision criterion was pre-set: MedQA stays primary unless the repaired MATH-500 attractor rate comes within 4× of MedQA's. The gap was 16×; it is now 11.7×. D-2 holds.
The way it held is the interesting part. MATH-500's attractor rate rose 28% while its P(both wrong) halved, 0.334 → 0.170. The manufactured rows were inflating honest error enormously while contributing almost no shared error — they were near-random disagreement. Removing them cut the error rate in half and slightly raised the coordinated fraction of what survived. The contrast the benchmark was chosen to illustrate is now a real property of an open answer space rather than partly an artifact of truncation.
MATH-500 has the most honest error and the least coordinated honest error — an open answer space means two wrong models pick two different wrong answers. At 0.0050 a full run yields ≈5 exploitable cells, so MX-H1 and MX-H4 would be untestable there. This reverses the benchmark choice I first recommended on headroom grounds, and it is the same G2′ effect the paper already reports, arriving as an experimental-design constraint. This table is recomputed after X3 before Phase 3 runs; the pre-registered criterion is that MedQA stays primary unless cleaned MATH-500 comes within 4× of it (current gap: 16×).
Hypotheses (pre-registered, docs/mx_prereg.md)
| id | claim | pre-registered expectation |
|---|---|---|
| MX-H1 | P2/P3 contain more cascades than P1 under L1 | reportable only if ≥ 0.100 |
| MX-H2 | cross-family (P3) beats same-model (P2) | predicted undetectable — order 0.03, below floor; registered so the null is interpretable |
| MX-H3 | gate overhead < 5% wall-clock per stage | grounded in E3: +3.6% at single-question scale, and generation is ~6598× the aggregation cost |
| MX-H4 | L2 easier to contain than L1 | positive in every AIP arm |
| MX-H5 | the MX-H4 gap widens on MedQA relative to MATH-500 | direction corrected by measurement — the spec predicted the opposite |
Two of the five are registered as expected-undetectable. That is the point of pre-registering them.
Verification — six permanent gates
Phase 1 is a 20-task probe whose only job is to prove the harness before any
liar runs. scripts/mx_phase1_verify.py:
| gate | checks |
|---|---|
| G1 | per-stage outputs exist and are well-formed |
| G2 | every gate decision is logged, not inferred |
| G3 | liar telemetry — which slot, which variant, whether it fired |
| G4 | determinism across an identical re-run |
| G5 | truncation audit — added because of the X3 defect |
| G6 | external consistency: MX's own accuracy vs the main-study cache, per model |
Plus an S5 invariant and a standing ritual of eyeballing five raw prompt/completion pairs per model × stage at a stated seed. Neither costs GPU time.
Phase 1: 14 of 17 gates green, then a hard stop
The first Phase 1 run failed on four defects. Three are fixed and confirmed by the gates that caught them; the fourth is the reason the chain is stopped.
fit_stagewas never called.AIPAggregatordeclaresneeds_fit; an unfitted gate has no calibrated honest-coincidence ceiling, logged 0 of 120 decisions, and emitted a constantA— P2/P3 at 0.10 on a four-way space, below chance. Fixed: the stage loop is now collect-then-fit-then-aggregate, and decisions went 0/120 → 120/120 with the decision vocabulary now the gate's own.- G6 was miscalibrated, and that was my error. It compared MX's 20-task
subsample against the cached 200-task population, so at n=20 the binomial
standard error (≈0.11) exceeded the floor by construction and flagged ordinary
sampling noise as a wiring fault. Restricted to matched task ids,
ministral3_14breproduces its cached accuracy to +0.000. Fixed. A gate that fires on its own noise trains you to ignore it. - 63 of 706 generations unparsed, 62 of them phi4 discards — the same defect as (4). Now 27 of 690.
phi4_mini_reasoningdoes not terminate. Still open, and it is what stopped the chain.
Current gate status. Green: G1, G1b, G2, G2b, G3, NC-P2, NC-zero-INVERT, stage-row extraction, G4, G4b, and G6 on all three models. Red:
| gate | first run | after the fix |
|---|---|---|
NC P3 == P4 (no liar) |
48/60 differ | 4/60 |
| every generation parsed | 63/706 | 27/690 |
| G5 discards ≤ 1% per cell | phi4 25–46% | phi4 15–28% |
All three are one root cause. phi4 hits its cap, the truncation guard correctly refuses to manufacture an answer, the row is discarded, and two arms that must be identical diverge because one member has no answer. The four NC-differing rows are exactly the two tasks carrying a phi4 discard — the guard behaving correctly, not a wiring fault.
The finding that stopped the chain: budget evidence can be self-referential
Amendment A2 raised phi4 to 3072 tokens, justified by its finished-completion lengths measured at a 2048 cap: p50/p90/max = 1003/1538/2013, read as "a real ceiling around 2000".
At 3072, the same statistics are 1051/1981/2935.
The ceiling moved with the cap. The tail was never a property of the model — it was the cap truncating the very distribution used to justify the cap. Raising the budget again chases the same moving target, so it was not done. This is rumination-class non-termination, not budget starvation, and it is being written up as a deployment-suitability finding: a model that will not stop is disqualified from staged-pipeline participation regardless of its accuracy.
Status
STOPPED. Phase 2 not started. Phase 3 not started. No liar code path has
executed at any scale. Tagged mx-phase1-red; state in
results/mx_phase1_state.md.
A roster amendment (A3) replacing phi4 is under review and not yet written — its precondition check found that the proposed substitute is clean on MedQA (the benchmark MX uses) but carries 4 manufactured MATH-500 rows, and that the swap would raise mean MedQA attractor participation from 0.0917 to 0.1067. Both facts are on the table before any text is committed, because a roster change made after seeing results is not a roster change anyone should trust.
Code: src/aip/pipeline/ (stages.py, arms.py, liars.py, runner.py),
drivers scripts/mx_phase1.py and scripts/mx_phase1_verify.py, pre-registration
docs/mx_prereg.md, 38 probe tests in tests/test_mx_pipeline.py.
8. Running it
uv venv --python 3.11 && uv pip install -e ".[dev]" # offline phases + tests
uv pip install -e ".[gpu]" # vllm 0.19.1, Phases 1 and 4
uv pip install -e ".[hub]" # only to publish to the Hub
python scripts/preflight.py --all # GPU + Hub check; never skip
python scripts/run_all.py --dry-run # ordered plan, runs nothing
python scripts/run_all.py --smoke # tiny inputs, proves the wiring
python scripts/run_all.py # the real chain
Progress is in RUNLOG.md, state in results/STATUS.json, completion stamps
in results/.stamps/. Every phase is resume-safe and skips finished work by artifact,
not by reloading a model to discover there is nothing to do.
Offline reproduction needs no GPU and no benchmark download. With the cache in
place, scripts/task6_audit_pending.py, scripts/make_numbers.py, the figure
generators and the whole analysis chain run on CPU.
Publishing to the Hub
python scripts/hf_push.py --dry-run # print the plan, upload nothing
python scripts/hf_push.py --repo-id <user>/<name> # dataset repo
The script refuses to push if the artifact is incomplete, and checks the token's shape before contacting the Hub — a previous attempt failed on a credential that was not a Hugging Face token at all, and the resulting 401 said nothing about that.
Pinned versions, and why
vllm==0.19.1— the only release satisfying both constraints: Gemma 4 support first appears in 0.19.0, and 0.20.0+ resolves a CUDA 13 torch that cannot initialise on this node's 12.8 driver.transformers>=5.5.1— Gemma 4's tokenizer schema is unparseable by 4.x; vLLM 0.19.1 excludes 5.0–5.5.0.
Both were forced by Gemma 4 support, not chosen, and both are recorded in every run manifest.
Building the manuscript
# tectonic is a single binary and needs no TeX Live installation
curl -sSL -o tectonic.tar.gz https://github.com/tectonic-typesetting/tectonic/releases/download/tectonic%400.15.0/tectonic-0.15.0-x86_64-unknown-linux-musl.tar.gz
tar xzf tectonic.tar.gz
cd paper && ../tectonic -X compile main.tex
Compiles clean: 12 pages (IEEE two-column, references included, no appendix),
0 undefined references, 0 undefined citations, 0 overfull boxes, 0 non-font LaTeX
warnings. paper/main.pdf is committed so a reader does not need the toolchain.
One \scaffold marker remains by design — the archive DOI, pending
de-anonymisation.
9. Tests
pytest # 377 fast tests
pytest -m slow # + 47 dataset downloads and real-gold round-trips
Beyond ordinary unit tests, the suite enforces the artifact's integrity rules: a hand-typed decimal anywhere in the manuscript fails the build; a claim cited in the text must exist in the ledger; no claim may sit unaudited against the measurement floor; the evasion anchor must be f = 0.5; and every figure the manuscript cites must carry a provenance entry naming the script that produced it. That last check exists because seven figures spent this entire regeneration as L4-era leftovers, in a second top-level figures directory LaTeX never read.
Licence
MIT for the code, the derived results and the manuscript.
The benchmark corpora under data/benchmarks/ are redistributed under their own
licences, not under MIT: ARC and BoolQ are CC BY-SA (share-alike), the rest
MIT or equivalent. Model outputs in data/cache*/ are subject to the licences of
the checkpoints that produced them — see configs/models.yaml for the exact
model revisions.
Citation
@misc{liars_are_information,
title = {Liars Are Information: Threshold-Gated Inversion of Coherent
Adversaries in Decentralized LLM Swarms},
note = {Artifact: inference cache, result parquets, claims ledger and manuscript},
year = {2026}
}