Skip to content

Platform Sweeps (P1–P11 + D-series)

This page catalogues the eleven foundational platform experiments (P1–P11, single-arena, Paper 0) and the v0.9 / v1.0 governance D-series (D1–D5, single-arena → world-substrate, Papers 1 / 2), organised by research question. Each entry covers the hypothesis, configuration, key result, and figure reference.


RQ1: The Fragility Threshold

P1 — Aggressive Fraction Sweep

Attribute Detail
Hypothesis A single aggressive agent creates a discontinuous step-collapse
Sweep 0–6 aggressive agents in 6-agent pool (\(r = 0.12\), \(K = 150\), 20–30 runs/cell); rest conservative
Result Gate pass drops 100% → 0% with 1 aggressor. Gini peaks at intermediate fractions (1–2 agents). Game length recovers slightly in uniform pools.
Figure aggressive_fraction__light_paper.png

P2 — Regen Rate Sweep

Attribute Detail
Hypothesis Sufficiently high \(r\) might offset aggressive extraction
Sweep \(r \in [0.04, 0.32]\), 1 aggressive + 2 conservative + 1 adaptive
Result Gate pass 0% throughout. Game length rises with \(r\) (23→41 turns) but depletion is only delayed, never prevented. Gini rises with \(r\) — longer games compound extraction advantage.
Figure regen_rate_sweep__light_paper.png

P3 — Pool Size Scaling

Attribute Detail
Hypothesis Larger pools dilute aggressor impact (1/N hypothesis)
Sweep 1 aggressive + \(N-1\) conservative, \(N \in [2, 12]\)
Result Gate passes only at \(N = 2\); 0% for all \(N \geq 4\). Game length decreases with \(N\) — more conservative agents sum to greater total harvest. Rules out dilution hypothesis.
Figure pool_size_scaling__light_paper.png

P4 — Pool Size Fine Sweep

Attribute Detail
Hypothesis Any transition to gate passage in \(N = 8\)\(20\) range?
Sweep \(N \in [8, 20]\), 1 aggressive + rest conservative
Result Gate pass 0% throughout. Game length declines monotonically 20→12 turns. In large pools, aggregate extraction — not individual aggression — drives collapse.
Figure pool_size_fine__light_paper.png

P5 — 2D Parametric Map

Attribute Detail
Hypothesis Gate pass frontier exists in \(r \times \text{aggressive fraction}\) space
Sweep \(r \in \{0.06, 0.10, 0.14, 0.18, 0.24, 0.32\} \times\) aggressive agents \(\in \{0, 1, 2, 4, 6\}\), 8-agent pool
Result The only sustainable region is zero aggressive agents, regardless of \(r\). Welfare rises with \(r\) even in collapse — divergence between individual and collective outcomes.
Figure heatmap_regen_x_aggr__light_paper.png

RQ1–RQ2: Structural Asymmetries and Institutional Limits

P6 — Initial Endowment Asymmetry

Attribute Detail
Hypothesis Wealthier starting position compounds into structural advantage
Sweep Aggressor initial endowment 10–120; 3 others at 30
Result Zero-sum under collapse: aggressor wealth rises, others' falls in perfect mirror symmetry. Game length decreases, Gini rises 0.30→0.44, gate never passes.
Figure asymmetry_sweep__light_paper.png

P7 — Scheduler Comparison

Attribute Detail
Hypothesis Turn order rules significantly alter outcomes
Sweep 1 aggressive + 3 conservative across 3 schedulers
Result All three schedulers produce indistinguishable outcomes. Gate pass 0%, game length ~27–28, Gini ~0.20–0.21. Permutation test (B5): welfare \(p = 0.675\), game length \(p = 1.000\). Strategy composition dominates institutional turn-order rules.
Figure scheduler_comparison__light_paper.png

RQ2–RQ3: Adaptation and Observability

P8 — Adaptive vs. Conservative Effectiveness

Attribute Detail
Hypothesis Adaptive agents (trend-reactive) outperform conservative (fixed moderate)
Sweep 1 aggressive + 3 adaptive/conservative/mixed, \(r \in \{0.10, 0.14, 0.20\}\)
Result Conservative pools yield higher welfare and lower Gini at all \(r\). The reactive < conservative effect: reactive agents create an exploitable vacuum that the aggressor fills.
Figure adaptive_effectiveness__light_paper.png

P9 — Observability Noise

Attribute Detail
Hypothesis Intermediate noise is welfare-optimal
Sweep \(\sigma/\mu \in \{0.0, 0.05, 0.1, 0.2, 0.35, 0.5\}\) + hidden; 1 aggressive + 2 conservative + 1 adaptive
Result Welfare-optimal at 35–50% noise. Mechanism: noise dampens aggressive extraction (stock looks smaller) without affecting conservatives. Full occlusion is worst — Gini 0.51, welfare collapse.
Figure observability_noise__light_paper.png

RQ4: Temporal Dynamics

P10 — Perturbation: Regen Shock

Attribute Detail
Hypothesis Seneca asymmetry: negative shocks hurt more than positive shocks help
Sweep Shock factor \(f \in \{0.10, 0.25, 0.50, 0.75, 1.50, 2.00\}\) at turn 20; 1 aggressive + 2 conservative + 1 adaptive
Result Asymmetry modest but visible. Negative shocks shorten game by ≤2 turns (plateau — collapse already inevitable). Positive shocks provide larger gains (+5 turns at ×2.0). Gate pass 0% for all. B4 isolates the pure resource component.
Figure perturbation_regen_shock__light_paper.png

P11 — Mule: Strategy Override

Attribute Detail
Hypothesis Cooperative equilibria fragile to late defection
Sweep Conservative agent switched to aggressive at \(T \in \{5, 10, 15, 20, 25, 30, 35, 40, 45\}\); 3 conservative + 1 mule
Result Gate pass 0% at all defection timings (including \(T = 45\)). No amount of prior cooperation banks resilience against betrayal in the presence of a co-present aggressor. Welfare rises with later defection (shorter aggressive window); Gini falls, but the no-defection baseline is unreachable.
Figure mule_strategy_override__light_paper.png

RQ5: Governance — pacts, cheap talk, sanctioning (Paper 1, v0.9)

D1 — Cliff under active cheap talk

Attribute Detail
Hypothesis Pure cheap talk solves the cliff (Farrell–Rabin positive case)
Sweep 5 conservative + 1 aggressive LLM agents, K=150, \(r\)=0.12, 30 turns, N=5 seeds, cell A (no cheap talk) vs B (full surface)
Result Negative. A preserves 0/5, B preserves 2/5 — cheap talk helps but does not solve. lie_score ≈ 0 (announcements honoured); silent_defection is the robust diagnostic at 0.38 ± 0.07 in B
Script scripts/pilot_d1.mjs

D2 — Graduated sanction ladder under Mule

Attribute Detail
Hypothesis A graduated ladder dominates a flat sanction
Sweep 5-agent Mule scenario, ladder [1, 3, 10] vs constant 3; built-in agents
Result Ladder dominates: 28 vs 21 turns survival, −50% violations, mule wealth −40%. Mechanism: wealth ablation, not pure deterrence
Script scripts/pilot_d2.mjs

D3 — Voluntary sanctioner ladder sweep

Attribute Detail
Hypothesis Pareto trade-off between graduated and deterrent ladders
Sweep N=5 × 3 ladders: soft [1,2,4], canonical [1,3,10], hard [1,5,25]
Result Flat-5 preserves 1/5; graduated ladders preserve ≥ 4/5. Soft: 4/5 + mule wealth 77.8; hard: 5/5 + mule wealth 21.6 (−72% deterrence). LLM N=5 replication: B preserves the reduction; gate-by-repeat (C_voluntary50_r2) backfires (16.8 violations)
Script scripts/pilot_v10_d3_*.mjs

D1-v10 — Cliff confinement

Attribute Detail
Hypothesis The cliff stays local in a federated world
Sweep 3×3 world: 1 focal Mule arena + 8 conservative arenas
Result Cliff confined: 5/5 collapses in arena0, 9/9 survivors elsewhere
Script scripts/pilot_v10_d1_*.mjs

D4 — Treaty enforcer

Attribute Detail
Hypothesis A treaty enforcer at world level deters arena-collective breaches
Sweep 3-arena world (5p + 3p + 3p), treaty harvest_cap_per_arena_per_round = 8.0, TreatyEnforcerAgent with sanctionFactor=0.9. Built-in baseline + LLM N=5
Result Built-in: enforcer fires 1.0× / round, shock 18.1. LLM N=5: enforcer fires 2.4 ± 0.9, shock 38.2 ± 13.6 — LLMs breach the cap 2.4× more often
Script scripts/pilot_v10_d4_llm_n5.mjs

D5 — Exclusion + junk arena (behavioural contagion)

Attribute Detail
Hypothesis Migrating violators to a junk arena cleans the mains without externalising defection
Sweep 4-arena world: 3 mains under voluntary_sanctioner + 1 shared junk. exclude perturbation routes repeat violators. Built-in + LLM N=5
Result Built-in junkPurity 1.00 (junk stays inert). LLM N=5: junkPurity 0.63 ± 0.34 — migrated LLM violators continue defecting in the junk. Behavioural contagion finding
Script scripts/pilot_v10_d5_llm_n5.mjs

F2.1 — Seed-robustness sweep

Attribute Detail
Hypothesis v1.0 mini-pilots are seed-robust at headline-metric level
Sweep 5 pilots × PILOT_SEED ∈ {17, 23, 29, 31, 37} × observability.commonsStock.noise = 0.03
Result Institutional events (collapse, sanction counts, exclude counts) flat 5/5. Depth-of-cliff metrics (final stock when stock is near threshold) sensitive, cv 20-60%. Full ledger in Paper 2 §A1
Script scripts/sweep_v10_seed.mjs

See Also