4  Discussion

The four pilots deliver a composite picture: the transition from heuristic to LLM agents is the dominant driver of cooperation; communication redistributes wealth but does not preserve the commons; graduated sanctioning is robustly effective — and the shape of the escalation, not its magnitude, carries the effect.

4.1 Architecture vs communication

The D1 decomposition (Section 3.2) is the methodological headline. The original two-cell design (Cell A: built-in/no-talk vs Cell B: LLM/with-talk) confounded agent architecture with the communication channel. Cell C (LLM/no-talk) resolves the confound. Three findings emerge from the \(2 \times 2\) factorial:

  1. LLM architecture is the primary effect. The A \(\to\) C transition (built-in to LLM, no talk in either) raises cooperation_index from \(0.28\) to \(0.50\) and extends the game from \(24\) to \(\sim 30\) turns. LLM agents cooperate more than heuristic agents before any communication is possible. The contrast — fixed-policy heuristics versus signal-conditioned reasoning — has been central to bounded-rationality theory since its original formulation (Simon 1955; Gigerenzer and Gaissmaier 2011): which side of that divide outperforms the other is task-dependent, and on the BDPD cliff it is the signal-conditioned side.
  2. Cheap-talk is orthogonal to preservation. Cell C vs Cell B (both LLM, talk off vs on) shows \(\Delta\bar{x} \approx 5\) units on a pooled \(\sigma \approx 26\) (Cohen’s \(d = 0.17\), \(n \approx 540\) per cell required for \(80\%\) power). The talk channel adds no detectable preservation benefit.
  3. Talk redistributes, it does not preserve. The aggressor’s wealth drops \(\sim 33\%\) when talk is enabled (Cell C \(\to\) B: \(177 \to 119\)), while conservators gain \(\sim 17\%\) (\(46.6 \to 54.4\)). The cooperation_index is marginally lower in Cell B (\(0.43\)) than Cell C (\(0.50\)), ruling out a fairness-priming explanation: talk enables coordination against the defector, not greater cooperation overall.

The architecture finding is the paper’s baseline, not its conclusion: LLM agents exhibit a cooperation predisposition that heuristic agents lack, but governance structures determine whether that predisposition translates into stable, maximal preservation. The \(4/5\) preservation rate of Cell C — high but not perfect — is the starting point from which the remaining experiments ask: what institutional design completes the remaining gap?

4.2 Relationship to Farrell-Rabin

The classical Farrell-Rabin prediction — that cheap talk selects the Pareto-superior equilibrium in coordination games with multiple rankable equilibria (Farrell and Rabin 1996) — does not straightforwardly apply. The BDPD cliff has no rival equilibrium to be selected by talk: below threshold, individual short-horizon optimisation strictly prefers over-extraction. The Cell C result confirms this: LLM agents cooperate more than heuristics without talk, suggesting the cooperation lift is architectural (richer state representation, longer planning horizon) rather than communicative.

4.3 Implications for governance design

The C1 mini-pilot demonstrates that a one-line sanction perturbation suffices to make pact violations costly and deterministic, echoing the broader finding from laboratory public-goods experiments that peer punishment transforms cooperation (Fehr and Gächter 2000). The D2 mini-pilot at \(N = 5\) refines the picture: not all sanction schedules are equivalent. A graduated ladder \([1, 3, 10]\) preserves the commons in \(5/5\) seeds; a flat amount \(3\) preserves it in \(3/5\).

Reading D1 and D2 jointly, the governance interventions span a preservation table (Table 4.1) that accumulates cooperation sources one row at a time. Each row adds one intervention layer; the preservation column records whether that layer alone shifts the outcome. The two dominant steps are easy to spot: the architecture switch (row 1 to row 2) and the shift from flat to graduated sanctions (rows 4–5 vs rows 6–8):

Table 4.1: Preservation table across the four pilots, \(N = 5\) each. The first two rows isolate the architecture effect; the remaining rows hold the LLM baseline and accumulate governance layers. The Config column makes the substrate change between D1 and D2/D3 explicit (see Chapter 3 for rationale): the table therefore reports results from distinct experimental configurations, not a within-experiment controlled comparison. The two flat-sanction rows (flat-3 at \(3/5\), flat-5 at \(1/5\)) come from different pilots (D2 and D3), so their ordering is not a controlled flat-vs-flat contrast; the load-bearing evidence is the within-D3 comparison of flat-5 against the graduated ladders.
Intervention Agent type Config Commons preserved at \(T = 30\)
no talk, no sanction (Cell A, D1) built-in \(K=150\), 6p \(0 / 5\)
LLM, no talk, no sanction (Cell C, D1) LLM \(K=150\), 6p \(4 / 5\)
LLM \(+\) talk, no sanction (Cell B, D1) LLM \(K=150\), 6p \(2 / 5\)
substrate change below (D2/D3)
LLM \(+\) talk \(+\) flat sanction \(3\) (Cell B, D2) LLM \(K=100\), 5p \(3 / 5\)
LLM \(+\) talk \(+\) flat sanction \(5\) (Cell C, D3) LLM \(K=100\), 5p \(1 / 5\)
LLM \(+\) talk \(+\) graduated \([1, 2, 4]\) (Cell A, D3) LLM \(K=100\), 5p \(4 / 5\)
LLM \(+\) talk \(+\) graduated \([1, 3, 10]\) (Cell A, D2) LLM \(K=100\), 5p \(5 / 5\)
LLM \(+\) talk \(+\) graduated \([1, 5, 25]\) (Cell B, D3) LLM \(K=100\), 5p \(5 / 5\)

The table reveals two orthogonal effects. First, switching from built-in to LLM agents (\(0/5 \to 4/5\)) is the single largest preservation lift, without any governance intervention. Second, within LLM agents, cheap-talk alone does not improve preservation; only graduated sanctioning brings a clean positive. The apparent non-monotonicity from Cell C (\(4/5\)) to Cell B (\(2/5\)) reflects sample variance at \(N = 5\), not a systematic talk penalty: Cohen’s \(d = 0.17\) between the two cells confirms no significant difference in commons stock. The architecture-to-LLM transition is the dominant step; talk adds variance without shifting the mean. The D3 ladder-geometry sweep then rules out the natural confound — “maybe a high-enough flat amount would have done the same” — by showing that a flat-\(5\) schedule (amount roughly halfway between the canonical ladder’s extremes, \(1\) and \(10\)) preserves the commons in only \(1/5\) seeds. All three graduated ladders preserve it in \(\geq 4/5\), including the soft \([1, 2, 4]\) whose terminal value (\(4\)) is below the failing flat-\(5\). The load-bearing feature is therefore the escalation shape, not the amount.

The three graduated cells trace a Pareto frontier between final commons stock and violator-wealth retention — soft \([1, 2, 4]\) at (stock \(22\), wealth \(\sim 78\)), \([1, 3, 10]\) at (\(26\), \(\sim 55\)), hard \([1, 5, 25]\) at (\(38\), \(\sim 22\)) — a genuine trade-off rather than a single dominated point. On the binary preservation rate the frontier would degenerate (\([1, 3, 10]\) and \([1, 5, 25]\) both score \(5/5\)), so the continuous stock is the discriminating axis. Choosing the ladder is a policy choice about how punitive the regime is, holding preservation roughly constant.

On a fragile commons in the BDPD cliff regime, LLM agent architecture alone produces a large cooperation lift over heuristic baselines. Communication adds no marginal preservation benefit, but graduated-sanctioning plugins dominate the LLM baseline cleanly. The marginal value of grading the enforcement is the difference between \(1/5\) (flat-\(5\)) and \(5/5\) (graduated \([1, 3, 10]\) or \([1, 5, 25]\)). Within graduated regimes, the choice of ladder shape sets the punitive-vs-preservative trade-off rather than overall efficacy. Talk’s contribution is redistributive: it shifts wealth from defectors to conservators without changing the commons trajectory.

Table 4.2 maps the results to the three Ostrom design principles exercised in the present design (Ostrom 1990):

Table 4.2: Mapping for the principles exercised by this paper’s mini-pilots. The remaining five Ostrom principles (boundaries, rules-fit-context, conflict resolution, recognition of rights, nested enterprises) are not operationalised here: boundaries and nested enterprises are taken up in BDPD\(^2\) (Brunelli 2026b), the others remain open. The three principles shown compose (adding #5 to #4 moves preservation from \(3/5\) to \(5/5\)) and do not substitute (talk without enforcement does not preserve).
Ostrom principle BDPD lever Result
#3 Collective-choice arrangements Cheap-talk channel (broadcast, announce, pact) Redistribution, not preservation (D1 Cell B vs C)
#4 Monitoring Sanction mechanism (detect pact violations) Necessary condition for #5 (C1)
#5 Graduated sanctions Escalating ladder \([1, 3, 10]\) vs flat \([5, 5, 5]\) Graduation \(\geq 4/5\); flat \(1/5\) (D3)

Ostrom’s specific insistence on graduated — rather than fixed — sanctions is mechanically load-bearing at this scale.

4.4 Limitations

  • Mini-pilot \(N\). All four pilots (C1, D1, D2, D3) are \(N = 5\) seeds per cell. At \(N = 5\), quantitative effect sizes are unreliable; the load-bearing claims are qualitative — which cells preserve vs which collapse — following the logic of theoretical replication (Yin 2009) rather than statistical generalisation. Bayesian credible intervals (Beta(1,1) prior) confirm that the key contrasts (0/5 vs 4/5, 1/5 vs 5/5) are credible even at this \(N\), while finer comparisons (3/5 vs 5/5) are not; see BDPD\(^2\) (Brunelli 2026b, Appendix A4, §Bayesian credible intervals). A campaign-scale rerun (\(N \geq 30\) per cell, denser ladder grid) is deferred to future work.
  • D1 architecture–talk decomposition. The \(2 \times 2\) factorial resolves the original confound but leaves an empty cell (built-in + talk) that is structurally impossible (built-in agents cannot use language). The talk-vs-no-talk comparison (Cell B vs Cell C) yields Cohen’s \(d = 0.17\) on commons stock; detecting a real effect this small would require \(\sim 540\) replicates per cell. At \(N = 5\), talk’s preservation contribution is indistinguishable from zero.
  • D2 individual-level claims do not survive at \(n = 5\). Defector wealth, cumulative-loss curves, and seed-to-seed deterrence patterns overlap fully between graduated and constant cells; only the binary commons-survival outcome (5/5 vs 3/5) is robust at this sample size. Single-seed narratives (“learning event at \(T_{21}\)”, “social transmission across agents”) are an artefact of seed selection and require \(N \geq 20\) to test.
  • Single LLM family. All results come from deepseek-v4-flash without thinking. Model-family diversification (Claude, Qwen, GLM) is a campaign-scale extension; in particular, thinking-enabled runs may surface different awareness-without-restraint patterns. The architectural cooperation lift (Cell C) may be specific to this model family’s training distribution. The near-zero lie_score observed in D1 (Cell B) similarly may reflect the archetype-prompt nudge (“you tend toward conservation” / “you tend toward aggressive extraction”) rather than a general announcement- honouring property of LLMs; a prompt-free control would be needed to separate the two. More broadly, LLM-based replications of human-subject experiments, while qualitatively faithful, are known to exhibit systematic distortions such as the hyper-accuracy effect (Aher et al. 2023); the \(4/5\) preservation rate in Cell C should therefore be read as an upper bound rather than a point estimate. Reproducibility is further bounded by the fact that deepseek-v4-flash is a cloud-served checkpoint with no user-pinnable version: provider updates between the run date of this paper and any replication attempt may shift the underlying model; a campaign-scale follow-up should record the API checkpoint identifier or move to a locally hosted model.
  • D2/D3 sanction cells all include talk. We lack an “LLM + no-talk + graduated sanctions” condition, so we cannot isolate whether the graduated effect is independent of talk or interacts with it. Given that talk introduces redistributive dynamics without preservation benefit (D1), the graduated-sanction effect may differ in a no-talk setting; a campaign-scale factorial crossing talk \(\times\) sanction regime would resolve this.
  • Signal-adaptivity vs systematic noise. We cannot exclude that part of the observed signal-adaptivity reflects LLM responses to spurious regularities in the environment rather than to the intended institutional signals. However, the directionality of the effect — LLM agents consistently cooperate more under strong institutions and defect more under weak ones across the four mini-pilots (C1, D1, D2, D3) — argues against pure noise. A formal null-signal control (scrambled or absent institutional cues) would require a dedicated experimental design and is deferred to future work.
  • No governance agent in these pilots. The meta-agent role (non-playing observer applying arena-level sanctions) was scaffolded but not exercised in the mini-pilots reported above (Yamagishi 1986). The companion paper BDPD\(^2\) (Brunelli 2026b) exercises both arena-level and world-level meta-agents across four vignettes (including a voluntary sanctioner and a cross-arena treaty enforcer), with two further substrate-validation vignettes in the appendix.

4.5 Outlook: from mini-pilot to campaign

A campaign-scale follow-up will run the D1 cell at \(n \geq 30\) replicates per cell with model-family diversification, extend the D3 sweep to a denser ladder grid (variable length and growth rate) at \(N \geq 30\) to characterise the Pareto frontier identified here, and extend the cheap-talk surface with a vote primitive (pledge is already implemented as a unilateral harvest_cap shorthand). The companion paper BDPD\(^2\) (Brunelli 2026b) extends this governance layer into a polycentric one — nested arenas, cross-arena resource links, world-level treaties, and meta-agents at both arena and world scope — and is the substrate against which the campaign-scale runs will be measured.

Methodological Coda: LLMs as Synthetic Collaborators

Two distinct roles of LLMs in this work should not be conflated. Role A (epistemic/generative): LLMs participated in conceptual development, implementation, and manuscript preparation — this is the methodological contribution acknowledged in the transparency statement. Role B (experimental): LLM agents (DeepSeek-v4-flash) serve as generative subjects in every non-baseline cell of D1–D3 — this is the empirical content reported above. The workflow follows the human-AI collaborative methodology described in BDPD\(^0\) (Brunelli 2026a, §Methodological Coda): multiple frontier LLMs with different training distributions (Anthropic’s Claude, DeepSeek, Alibaba’s Qwen, Z.ai’s GLM-5.1) were employed for cross-validation of analysis, framing, and manuscript revisions, under the author’s direction.

Data and Code Availability

The BDPD simulation platform, experiment definitions, and analysis scripts are available at gitlab.com/bdpd/bdpd. The raw pilot artefacts referenced in Chapter 3data/pilot/c1/, data/pilot/d1_n5/, data/pilot/d1_cell_c/, data/pilot/d2_llm_n5/, data/pilot/d2_ladder_sweep/ — are intentionally git-ignored as runtime outputs and regenerate when the corresponding scripts/pilot_*.mjs are executed with DEEPSEEK_API_KEY set. All LLM runs use deepseek-v4-flash accessed through the public DeepSeek API at the run date of this paper; the cloud-served checkpoint is not user-pinnable and provider updates may shift the underlying model between the run date and any replication attempt.