This page catalogues the eleven foundational platform experiments
(P1–P11, single-arena, Paper 0) and the v0.9 / v1.0 governance
D-series (D1–D5, single-arena → world-substrate, Papers 1 / 2),
organised by research question. Each entry covers the hypothesis,
configuration, key result, and figure reference.
RQ1: The Fragility Threshold
P1 — Aggressive Fraction Sweep
| Attribute |
Detail |
| Hypothesis |
A single aggressive agent creates a discontinuous step-collapse |
| Sweep |
0–6 aggressive agents in 6-agent pool (\(r = 0.12\), \(K = 150\), 20–30 runs/cell); rest conservative |
| Result |
Gate pass drops 100% → 0% with 1 aggressor. Gini peaks at intermediate fractions (1–2 agents). Game length recovers slightly in uniform pools. |
| Figure |
aggressive_fraction__light_paper.png |
P2 — Regen Rate Sweep
| Attribute |
Detail |
| Hypothesis |
Sufficiently high \(r\) might offset aggressive extraction |
| Sweep |
\(r \in [0.04, 0.32]\), 1 aggressive + 2 conservative + 1 adaptive |
| Result |
Gate pass 0% throughout. Game length rises with \(r\) (23→41 turns) but depletion is only delayed, never prevented. Gini rises with \(r\) — longer games compound extraction advantage. |
| Figure |
regen_rate_sweep__light_paper.png |
P3 — Pool Size Scaling
| Attribute |
Detail |
| Hypothesis |
Larger pools dilute aggressor impact (1/N hypothesis) |
| Sweep |
1 aggressive + \(N-1\) conservative, \(N \in [2, 12]\) |
| Result |
Gate passes only at \(N = 2\); 0% for all \(N \geq 4\). Game length decreases with \(N\) — more conservative agents sum to greater total harvest. Rules out dilution hypothesis. |
| Figure |
pool_size_scaling__light_paper.png |
P4 — Pool Size Fine Sweep
| Attribute |
Detail |
| Hypothesis |
Any transition to gate passage in \(N = 8\)–\(20\) range? |
| Sweep |
\(N \in [8, 20]\), 1 aggressive + rest conservative |
| Result |
Gate pass 0% throughout. Game length declines monotonically 20→12 turns. In large pools, aggregate extraction — not individual aggression — drives collapse. |
| Figure |
pool_size_fine__light_paper.png |
P5 — 2D Parametric Map
| Attribute |
Detail |
| Hypothesis |
Gate pass frontier exists in \(r \times \text{aggressive fraction}\) space |
| Sweep |
\(r \in \{0.06, 0.10, 0.14, 0.18, 0.24, 0.32\} \times\) aggressive agents \(\in \{0, 1, 2, 4, 6\}\), 8-agent pool |
| Result |
The only sustainable region is zero aggressive agents, regardless of \(r\). Welfare rises with \(r\) even in collapse — divergence between individual and collective outcomes. |
| Figure |
heatmap_regen_x_aggr__light_paper.png |
RQ1–RQ2: Structural Asymmetries and Institutional Limits
P6 — Initial Endowment Asymmetry
| Attribute |
Detail |
| Hypothesis |
Wealthier starting position compounds into structural advantage |
| Sweep |
Aggressor initial endowment 10–120; 3 others at 30 |
| Result |
Zero-sum under collapse: aggressor wealth rises, others' falls in perfect mirror symmetry. Game length decreases, Gini rises 0.30→0.44, gate never passes. |
| Figure |
asymmetry_sweep__light_paper.png |
P7 — Scheduler Comparison
| Attribute |
Detail |
| Hypothesis |
Turn order rules significantly alter outcomes |
| Sweep |
1 aggressive + 3 conservative across 3 schedulers |
| Result |
All three schedulers produce indistinguishable outcomes. Gate pass 0%, game length ~27–28, Gini ~0.20–0.21. Permutation test (B5): welfare \(p = 0.675\), game length \(p = 1.000\). Strategy composition dominates institutional turn-order rules. |
| Figure |
scheduler_comparison__light_paper.png |
RQ2–RQ3: Adaptation and Observability
P8 — Adaptive vs. Conservative Effectiveness
| Attribute |
Detail |
| Hypothesis |
Adaptive agents (trend-reactive) outperform conservative (fixed moderate) |
| Sweep |
1 aggressive + 3 adaptive/conservative/mixed, \(r \in \{0.10, 0.14, 0.20\}\) |
| Result |
Conservative pools yield higher welfare and lower Gini at all \(r\). The reactive < conservative effect: reactive agents create an exploitable vacuum that the aggressor fills. |
| Figure |
adaptive_effectiveness__light_paper.png |
P9 — Observability Noise
| Attribute |
Detail |
| Hypothesis |
Intermediate noise is welfare-optimal |
| Sweep |
\(\sigma/\mu \in \{0.0, 0.05, 0.1, 0.2, 0.35, 0.5\}\) + hidden; 1 aggressive + 2 conservative + 1 adaptive |
| Result |
Welfare-optimal at 35–50% noise. Mechanism: noise dampens aggressive extraction (stock looks smaller) without affecting conservatives. Full occlusion is worst — Gini 0.51, welfare collapse. |
| Figure |
observability_noise__light_paper.png |
RQ4: Temporal Dynamics
P10 — Perturbation: Regen Shock
| Attribute |
Detail |
| Hypothesis |
Seneca asymmetry: negative shocks hurt more than positive shocks help |
| Sweep |
Shock factor \(f \in \{0.10, 0.25, 0.50, 0.75, 1.50, 2.00\}\) at turn 20; 1 aggressive + 2 conservative + 1 adaptive |
| Result |
Asymmetry modest but visible. Negative shocks shorten game by ≤2 turns (plateau — collapse already inevitable). Positive shocks provide larger gains (+5 turns at ×2.0). Gate pass 0% for all. B4 isolates the pure resource component. |
| Figure |
perturbation_regen_shock__light_paper.png |
P11 — Mule: Strategy Override
| Attribute |
Detail |
| Hypothesis |
Cooperative equilibria fragile to late defection |
| Sweep |
Conservative agent switched to aggressive at \(T \in \{5, 10, 15, 20, 25, 30, 35, 40, 45\}\); 3 conservative + 1 mule |
| Result |
Gate pass 0% at all defection timings (including \(T = 45\)). No amount of prior cooperation banks resilience against betrayal in the presence of a co-present aggressor. Welfare rises with later defection (shorter aggressive window); Gini falls, but the no-defection baseline is unreachable. |
| Figure |
mule_strategy_override__light_paper.png |
RQ5: Governance — pacts, cheap talk, sanctioning (Paper 1, v0.9)
D1 — Cliff under active cheap talk
| Attribute |
Detail |
| Hypothesis |
Pure cheap talk solves the cliff (Farrell–Rabin positive case) |
| Sweep |
5 conservative + 1 aggressive LLM agents, K=150, \(r\)=0.12, 30 turns, N=5 seeds, cell A (no cheap talk) vs B (full surface) |
| Result |
Negative. A preserves 0/5, B preserves 2/5 — cheap talk helps but does not solve. lie_score ≈ 0 (announcements honoured); silent_defection is the robust diagnostic at 0.38 ± 0.07 in B |
| Script |
scripts/pilot_d1.mjs |
D2 — Graduated sanction ladder under Mule
| Attribute |
Detail |
| Hypothesis |
A graduated ladder dominates a flat sanction |
| Sweep |
5-agent Mule scenario, ladder [1, 3, 10] vs constant 3; built-in agents |
| Result |
Ladder dominates: 28 vs 21 turns survival, −50% violations, mule wealth −40%. Mechanism: wealth ablation, not pure deterrence |
| Script |
scripts/pilot_d2.mjs |
D3 — Voluntary sanctioner ladder sweep
| Attribute |
Detail |
| Hypothesis |
Pareto trade-off between graduated and deterrent ladders |
| Sweep |
N=5 × 3 ladders: soft [1,2,4], canonical [1,3,10], hard [1,5,25] |
| Result |
Flat-5 preserves 1/5; graduated ladders preserve ≥ 4/5. Soft: 4/5 + mule wealth 77.8; hard: 5/5 + mule wealth 21.6 (−72% deterrence). LLM N=5 replication: B preserves the reduction; gate-by-repeat (C_voluntary50_r2) backfires (16.8 violations) |
| Script |
scripts/pilot_v10_d3_*.mjs |
RQ6: World substrate — links, treaties, exclusion (Paper 2, v1.0)
D1-v10 — Cliff confinement
| Attribute |
Detail |
| Hypothesis |
The cliff stays local in a federated world |
| Sweep |
3×3 world: 1 focal Mule arena + 8 conservative arenas |
| Result |
Cliff confined: 5/5 collapses in arena0, 9/9 survivors elsewhere |
| Script |
scripts/pilot_v10_d1_*.mjs |
D4 — Treaty enforcer
| Attribute |
Detail |
| Hypothesis |
A treaty enforcer at world level deters arena-collective breaches |
| Sweep |
3-arena world (5p + 3p + 3p), treaty harvest_cap_per_arena_per_round = 8.0, TreatyEnforcerAgent with sanctionFactor=0.9. Built-in baseline + LLM N=5 |
| Result |
Built-in: enforcer fires 1.0× / round, shock 18.1. LLM N=5: enforcer fires 2.4 ± 0.9, shock 38.2 ± 13.6 — LLMs breach the cap 2.4× more often |
| Script |
scripts/pilot_v10_d4_llm_n5.mjs |
D5 — Exclusion + junk arena (behavioural contagion)
| Attribute |
Detail |
| Hypothesis |
Migrating violators to a junk arena cleans the mains without externalising defection |
| Sweep |
4-arena world: 3 mains under voluntary_sanctioner + 1 shared junk. exclude perturbation routes repeat violators. Built-in + LLM N=5 |
| Result |
Built-in junkPurity 1.00 (junk stays inert). LLM N=5: junkPurity 0.63 ± 0.34 — migrated LLM violators continue defecting in the junk. Behavioural contagion finding |
| Script |
scripts/pilot_v10_d5_llm_n5.mjs |
F2.1 — Seed-robustness sweep
| Attribute |
Detail |
| Hypothesis |
v1.0 mini-pilots are seed-robust at headline-metric level |
| Sweep |
5 pilots × PILOT_SEED ∈ {17, 23, 29, 31, 37} × observability.commonsStock.noise = 0.03 |
| Result |
Institutional events (collapse, sanction counts, exclude counts) flat 5/5. Depth-of-cliff metrics (final stock when stock is near threshold) sensitive, cv 20-60%. Full ledger in Paper 2 §A1 |
| Script |
scripts/sweep_v10_seed.mjs |
See Also