Appendix C — Supplement Experiments and Statistical Inference

This appendix documents the five statistical supplements (B1–B5) introduced during analysis of the initial experimental results to address five specific questions arising from the original P-sweep findings: the parametrisation-dependence of the discontinuous threshold (B1), the generality of the reactive < conservative effect (B2), the claim that no level of regeneration offsets a single aggressor (B3), the conflation of strategic depletion with pure Seneca dynamics (B4), and the absence of formal statistical backing for the reported differences (B5). All simulation code and JSON definitions are available in the BDPD repository under experiments/definitions/.

C.1 B1 — Aggressor Intensity: Cliff-Edge Localisation

Motivation. The canonical P1 sweep uses the default aggressive heuristic with \(i = 0.5\) (fraction of stock requested per turn). The analysis raised the question of whether the “discontinuous step” might reflect an extreme parametrisation rather than a structural property. B1 sweeps \(i\) from 0.05 to 1.0 (coarse) and from 0.04 to 0.11 (fine, at step 0.01) to identify the precise intensity threshold below which a single aggressor no longer dooms the commons. Population: 1 aggressive + 5 conservative agents, \(r = 0.12\), \(K = 150\), 5–30 runs per cell.

Result. The cliff edge lies at \(i \approx 0.05\)\(0.06\): at \(i = 0.05\) the gate passes 100%; at \(i = 0.06\) it falls to 0%. The canonical heuristic (\(i = 0.5\)) sits nearly an order of magnitude above this threshold. The step-collapse is therefore not an artefact of an extreme parametrisation — even a minimally extractive aggressor (6% of stock per turn) dooms the commons within the canonical regime. Bootstrap CI on the cliff edge: [0.04, 0.05] (10000 iterations; quantised by the 0.01 fine-grid spacing). The coarse and fine sweeps are shown side-by-side in Figure C.1 (panels Figure C.1 (a) and Figure C.1 (b)).

(a)
(b)
Figure C.1: Left: B1 canonical sweep (intensity 0.05–1.00, step 0.05). Gate pass rate (top), game length and Gini (centre), welfare (bottom). Only \(i = 0.05\) passes; all higher values produce collapse. The cliff line annotation marks \(i = 0.10\) (first grid point with gate = 0%). Right: B1 fine sweep (intensity 0.04–0.11, step 0.01). Resolves the transition at unit precision. Gate passes at \(i = 0.04\) and \(i = 0.05\); collapses from \(i = 0.06\) onward. Cliff edge annotation: \(i \approx 0.05\)\(0.06\) (bootstrap CI [0.04, 0.05]).

Key finding — B1. The step-collapse threshold is a structural property of agent-type composition, not a consequence of extreme parametrisation. The canonical aggressor (\(i = 0.5\)) is representative of any strategy that requests a meaningful fraction of available stock; the safe zone (\(i \lesssim 0.05\)) corresponds to extraction so minimal it would not be recognised as aggressive play in practice.

C.2 B2 — Reactive Reduction Factor: Generalising the Effect

Motivation. The canonical P8 result shows that reactive agents underperform conservative ones. A further question arising from P8 was whether this holds across different strengths of the reactive response, or only for the specific heuristic hardcoded in the platform. B2 introduces a continuous parameter \(f \in [0, 1.5]\) — the reductionFactor — that scales how strongly the reactive agent cuts harvest in response to a falling stock signal (\(f = 0\): no reaction; \(f = 1\): canonical P8 heuristic; \(f > 1\): over-reactive). Population: 1 aggressive (default intensity) + 3 reactive agents, swept at three regen rates (\(r \in \{0.10, 0.14, 0.20\}\)), 20 runs per cell.

Result. Welfare declines monotonically with \(f\) at all tested regen rates; Gini rises monotonically (Figure C.2). The paradox is not a threshold phenomenon but a graded strategic externality: every unit of reactive response incurs a proportional welfare cost. The slope steepens with \(r\) (higher regen provides more harvest space to concede). All 95% bootstrap CIs on the slope exclude zero.

Table C.1: Welfare slope estimates from B2. Each CI computed from 10000 bootstrap iterations resampling cells (not replicates).
Regen rate Welfare slope per unit \(f\) 95% bootstrap CI
\(r = 0.10\) -14.62 [-16.39, -13.79]
\(r = 0.14\) -16.17 [-19.16, -14.26]
\(r = 0.20\) -20.61 [-22.91, -19.67]
Figure C.2: B2: Welfare score (left) and Gini coefficient (right) vs. reductionFactor \(f\), grouped by regen rate. Bars are coloured from dark (f = 0, no adaptation) to yellow (f = 1.5, over-reactive). Welfare is maximised at \(f = 0\) — when the reactive agent behaves as a fixed-fraction conservative — and declines monotonically. Gini rises correspondingly: the more the reactive agents yield, the more the aggressor captures.

Key finding — B2. The reactive < conservative effect is not an heuristic artefact: it holds continuously across the entire range of reactivity, with magnitude proportional to \(f\) and steeper at higher regen rates. Partial reactivity is worse than no reactivity.

C.3 B3 — Extended Regen Rate: Identifying the Rescue Boundary

Motivation. P2 sweeps \(r \in [0.04, 0.32]\) and finds gate = 0% throughout, leading to the claim “no level of regeneration offsets a single aggressor.” B3 extends this sweep to \(r = 1.50\) (coarse, irregular grid) and \(r \in [0.80, 1.05]\) (fine, step 0.025) to determine whether a rescue point exists and, if so, precisely where. Population: same as P2 (1 aggressive + 2 conservative + 1 reactive), \(K = 150\), 5–20 runs per cell.

Result. A rescue point exists: gate passage first occurs at \(r \approx 0.95\) (coarse scan: between \(r = 0.80\) (0%) and \(r = 1.00\) (100%); fine scan: transition between \(r = 0.925\) (0%) and \(r = 0.950\) (100%)). Bootstrap CI on rescue point: [0.95, 1.025] (10000 iterations). This boundary lies \(\approx 3\times\) above the canonical CPR maximum (\(r \leq 0.32\); \(0.95/0.32 \approx 3\)) and \(\approx 8\times\) above the default platform regen rate (\(r = 0.12\)), confirming that the original P2 claim is valid within any empirically realistic regime. The coarse and fine sweeps are shown side-by-side in Figure C.3 (panels Figure C.3 (a) and Figure C.3 (b)).

(a)
(b)
Figure C.3: Left: B3 canonical sweep (\(r \in [0.30, 1.50]\)). Gate pass rate (dashed, right axis) jumps from 0% to 100% at \(r = 1.00\). Annotation marks the rescue point. Game length (solid, left axis) saturates at 60 turns once the gate passes; welfare rises continuously. Right: B3 fine sweep (\(r \in [0.80, 1.05]\), step 0.025). Resolves the transition: gate = 0% at \(r = 0.925\), gate = 100% at \(r = 0.950\). Bootstrap CI on rescue: [0.95, 1.025].

Key finding — B3. The rescue boundary at \(r \approx 0.95\) corresponds to regeneration approximately matching the scale of combined per-turn extraction — the point at which stock replenishment finally outpaces the canonical aggressor. For reference, real common-pool resources – commercial fisheries and slow-growth forests – sit in a low-\(r\), slow-regeneration regime (Li et al. 2026; Scheffer 2009), far below this boundary. The rescue regime is thus physically unreachable for any plausible real CPR, which makes exclusion mechanisms – not environmental generosity – the only viable policy lever.

C.4 B4 — Shock Response without an Aggressor: Isolating Resource Dynamics

Motivation. P10 finds that negative regen shocks produce only modest, bounded damage while positive shocks provide larger gains — the reverse of the classical Seneca prediction. This inverted pattern could reflect either (a) an intrinsic property of the logistic commons, or (b) the aggressive agent having already pre-committed the system to its depletion trajectory before the shock arrives, so that negative shocks are marginal and positive shocks merely extend a doomed trajectory. B4 removes the aggressive agent entirely — 4 conservative agents, same shock protocol — to isolate component (a). Population: 4 conservative agents, \(r = 0.12\), shock at turn 20, \(f \in \{0.10, 0.25, 0.50, 0.75, 1.00, 1.50, 2.00\}\), 20 runs per cell.

Result. Without strategic pre-depletion, the logistic commons responds proportionally to shocks: negative shocks require severe magnitude (\(f < 0.25\)) to trigger collapse at all (welfare loss \(\approx\) 27.8 points at \(f = 0.10\) due to collapse), and moderate negative shocks (\(f = 0.50\), \(f = 0.75\)) produce moderate, real welfare losses (-21 and -11 points from a baseline of 112.8). Positive shocks provide comparable or larger gains (+23 at \(f = 1.50\), +46 at \(f = 2.00\)). The asymmetry indices confirm that the logistic model has no intrinsic Seneca-direction asymmetry (negative hurting more than positive helps) at any tested magnitude: indices are near zero at moderate magnitudes and positive at extremes — meaning positive shocks provide somewhat more relief than matched negative shocks cause damage, not the reverse.

Table C.2: Asymmetry indices from B4, defined as (welfare gain from positive shock) - (welfare loss from matched negative shock). A positive index means positive shocks help more than negative shocks hurt (anti-Seneca direction); a negative index means the reverse (classical Seneca direction). Baseline welfare = 112.8.
Paired shocks Asymmetry index Interpretation
\(f = 0.50\) vs \(f = 1.50\) +2.0 Near-symmetric; slight anti-Seneca lean
\(f = 0.25\) vs \(f = 1.50\) -0.7 Near-symmetric (negligible magnitude)
\(f = 0.10\) vs \(f = 2.00\) +18.2 Positive shocks help substantially more at extremes

This contrasts sharply with P10 (aggressive agent present), where the same moderate negative shock (e.g. \(f = 0.50\)) acts on a system already committed to collapse, making its marginal damage negligible. The contrast between the two experiments is therefore not about the resource model — it is about the strategic state of the system at the moment the shock arrives. The full shock-factor response and the near-symmetric \(\Delta\)-panel are shown in Figure C.4.

Figure C.4: B4: Game length by shock factor (left, red = negative, green = positive) and \(\Delta\) welfare and \(\Delta\) turns from no-shock baseline (right). Unlike P10, negative shocks require factor \(< 0.25\) to trigger collapse. The \(\Delta\)-panel shows near-symmetric response at moderate magnitudes, with no intrinsic Seneca-direction asymmetry.

Key finding — B4. The logistic commons has no intrinsic asymmetric shock response in the Seneca direction: without a co-present aggressive agent, matched negative and positive shocks produce near-symmetric welfare outcomes at moderate magnitudes. The inverted pattern observed in P10 — where negative shocks are bounded and positive shocks provide larger gains — is entirely an artefact of the aggressive agent having pre-depleted the commons before the shock arrives.

C.5 B5 — Statistical Inference: Cross-Sweep Analysis

Motivation. The analysis of the first set of experimental results suggested the need for formal statistical backing to support the reported differences. B5 applies retrospective inference to all 17 loaded sweep results (167 cells in total): logistic regression of gate pass rate on canonical covariates, bootstrap CIs for threshold locations, and permutation tests for stochastic sweeps.

Variance audit. All platform sweeps except two are deterministic (zero intra-cell variance), reflecting the deterministic nature of the built-in heuristics under a fixed scheduler. Only observability_noise (max \(\sigma = 0.85\)) and scheduler_comparison (max \(\sigma = 0.66\)) exhibit stochastic variance. For deterministic sweeps, inference is performed at the between-cell level via bootstrap (resampling cells). This audit applies to platform sweeps only; card-tournament sweeps (CT1–CT5) are inherently stochastic (Forest Die rolls, deck shuffles, mandatory-defection draws) and are sampled at \(n = 200\) games per cell to control this variance — see §3.3.

Table C.3: Cross-sweep logistic regression coefficients (log-odds scale). “Robust” = 95% bootstrap CI excludes zero. Model fit: 167 cells, 17 sweeps. The positive coefficient for pool size reflects a composition confound — larger pools in these experiments tend to have a smaller fraction of aggressive agents; controlling for aggressive_fraction would likely absorb this effect.
Predictor Coefficient 95% Bootstrap CI Robust?
aggressor intensity -67.4 [-130.4, +0.8] Borderline
regen rate +40.3 [+23.8, +77.3] \(\checkmark\)
aggressive count -32.9 [-63.6, -19.7] \(\checkmark\)
reactive reductionFactor -6.4 [-36.3, +2.9] \(\times\)
pool size (\(n\) agents) +1.8 [+0.9, +4.0] \(\checkmark\)

C.5.1 Cross-Sweep Logistic Regression

A binomial GLM was fitted to the 167-cell pooled dataset (outcome: gate successes out of \(n\) replicates; predictors: aggressor intensity, regen rate, aggressive count, reactive reductionFactor, pool size). Bootstrap CIs (10000 iterations resampling cells) are reported alongside model estimates; for deterministic sweeps, the bootstrap CI is more reliable than the model standard error. The same coefficients with their bootstrap intervals are visualised in Figure C.5.

Figure C.5: B5 forest plot. Cross-sweep logistic regression coefficients with 95% bootstrap CIs. Green = CI excludes zero on the positive side (regen rate, pool size); red = CI excludes zero on the negative side (aggressive count); grey = CI crosses zero. aggressor_intensity is borderline (CI upper bound = +0.8, direction correct). The two green predictors reflect structural features of the system: more regeneration and larger dilution of the single aggressive agent both shift the equilibrium towards gate survival.

C.5.2 Sweep-Specific Bootstrap CIs

Table C.4: Threshold bootstrap CIs (10000 iterations). CIs are quantised by the fine-grid spacing (0.01 for B1, 0.025 for B3); finer grids would narrow them further.
Experiment Observed threshold 95% Bootstrap CI Source sweep
B1 cliff edge (last \(i\) with gate > 50%) 0.050 [0.040, 0.050] aggressive_intensity_sweep_fine
B3 rescue point (first \(r\) with gate > 50%) 0.950 [0.950, 1.025] regen_rate_extended_fine

B2 welfare slopes (10000 iterations, resampling cells within each regen stratum):

Regen rate Slope (welfare per unit \(f\)) 95% CI
\(r = 0.10\) -14.62 [-16.39, -13.79]
\(r = 0.14\) -16.17 [-19.16, -14.26]
\(r = 0.20\) -20.61 [-22.91, -19.67]

All slopes exclude zero, confirming the monotonic strategic externality of the reactive response across all tested regen rates.

C.5.3 P7 Scheduler Permutation Test

Table C.5: P7 permutation test results (10000 permutations). Scheduler labels: simultaneous / sequential_random / wealth_weighted. None of the three metrics differs significantly across schedulers. The observed spread of under 1 welfare point falls entirely within chance variation.
Metric Scheduler means Observed statistic \(p\)-value Conclusion
Gate pass rate all 0% 0.000 1.000 Not distinguishable
Game length (turns) 27.0 / 27.5 / 27.8 0.332 1.000 Not distinguishable
Welfare score 69.1 / 68.2 / 68.1 0.544 1.000 Not distinguishable

Key finding — B5. Formal inference confirms three claims from the main text: (1) aggressive count and regen rate are the only covariates with bootstrap CIs robustly excluding zero in the cross-sweep regression; (2) the cliff edge (\(i \approx 0.05\), CI [0.04, 0.05]) and rescue boundary (\(r \approx 0.95\), CI [0.95, 1.025]) are quantified at the fine-grid precision of the corresponding sweep; (3) scheduler differences in welfare and game length are statistically indistinguishable (\(p > 0.05\) by permutation), confirming that strategy composition dominates turn-order institutional rules.

The positive pool size coefficient (+1.8, CI [0.9, 4.0]) should be interpreted as a dilution proxy, not a direct pool-size effect: in these experiments a single aggressive agent was held fixed while pool size varied (P3–P4), so larger pools mechanically reduce the aggressive fraction from 0.50 (N=2) to 0.05 (N=20). The coefficient captures aggressor dilution, not any independent benefit of group size. A model substituting aggressive_fraction for the separate aggressive_count and pool_size terms would likely absorb this effect and potentially reverse the sign on pool size. The three substantive conclusions above are unaffected by this confound.