1  Introduction

This paper maps the governance surfaces of a nested-commons platform through four vignettes, each isolating one research question at \(N = 5\) seeds (independent random initialisations of observation noise). The platform simulates common-pool resource (CPR) dilemmas where autonomous agents — both heuristic and LLM-driven — harvest from logistic commons that collapse irreversibly once extraction crosses a threshold (the cliff configuration established in BDPD\(^0\), the foundational paper of this series (Brunelli 2026a; Hardin 1968)).

1.1 Scope and format

This paper surveys the governance design space through targeted vignettes — each a minimal experiment isolating one institutional mechanism — rather than a statistical campaign. Within the LLM-driven social-simulation taxonomy of Mou et al. (2026), the work sits in the Scenario Simulation register — multiple LLM agents acting within a structured institutional context — bordering on Society Simulation where the World layer carries meta-agents (Park et al. 2023; Gao et al. 2023). BDPD\(^1\) (Brunelli 2026b; Balliet 2010) describes cheap-talk channels, pacts, and arena-level sanctioners within a single arena. The substrate described here wraps a World object around those mechanisms: multiple arenas, resource links between them, treaties spanning arenas, and meta-agents that observe across arena boundaries (Heikkila et al. 2018). The vignettes map which governance surface does what, and where the open empirical gaps are.

1.2 Bridge to BDPD\(^0\)

This paper sits between the substrate-level questions of (Brunelli 2026a) — engine dynamics, the Seneca asymmetry, the collapse-threshold geometry of logistic commons — and the campaign-scale follow-ups that the nested governance surface invites. BDPD\(^0\) closed with an outstanding promise: to test whether the collapse threshold sharpens under true capital–pollution feedback once strategic agents are added to the Seneca engine — a three-variable dynamics (resource \(R\), capital \(C\), hidden pollution \(P\)) in which capital grows first and pollution arrives as a lagging consequence, creating an asymmetric Seneca cliff studied in depth in BDPD\(^3\) (Brunelli 2026c). That promise frames Road A in our conclusions, which proposes a polycentric Seneca campaign as one of three next candidates. The remaining roads stay on logistic dynamics where BDPD\(^0\)’s calibration is settled, so the governance findings here can be read as governance under known engine dynamics rather than as a joint test of engine and institution.

1.3 The four vignettes

We selected four governance mechanisms — world composition, deterrence, treaty enforcement, and coercive exclusion — because each isolates a distinct institutional level (substrate, arena, world, cross-arena) while sharing the same logistic engine and agent pool. The progression from V1 to V4 climbs the institutional ladder from passive isolation to active removal, and each vignette compares a rule-bound heuristic baseline against a signal-adaptive LLM replication.

Vignette Research question Governance surface Agents
V1 Does the World abstraction leak across arenas? World composition built-in
V2 Does graduated sanctioning still pay when the sanctioner sees less? Arena-level meta-agent built-in + LLM
V3 Does a world-level enforcer absorb the cost of arena-level free-riding? World-level meta-agent built-in + LLM
V4 Does coercive exclusion cluster aggressors, or does it spill over? exclude perturbation built-in + LLM
V5 (appendix) Does the cliff finding migrate to a tabletop card-game substrate? Card-game replication tabletop players
V6 (appendix) What does seed-and-noise calibration of the built-in baselines look like? Seed-robustness ledger built-in

V1 establishes that the World substrate does not introduce spurious cross-arena coupling — a prerequisite for trusting V2–V4. V2–V4 each compare a built-in heuristic baseline against an LLM replication, because the qualitative direction of the governance findings turns out to depend on which class of agents one runs. Appendix A3 validates the platform’s cliff finding against the Forest of Humbaba card-game substrate; Appendix A4 collects the methodological scaffolding (seed-robustness, cost calibration).

1.4 What is and is not in scope

This is a survey at \(N = 5\) seeds — a deliberate choice in line with the empirical-PG critique that the literature largely fails to address evolution over time (Baldwin et al. 2024; Morrison et al. 2019): the substrate enables campaign-scale longitudinal follow-ups that no single-arena framework affords. It is not powered to detect small effects, and we do not report inferential statistics. Every result below should be read as a plausibility check of a finding the substrate makes natural to study, with a forward pointer to the campaign-scale follow-up that would settle it. We list those follow-ups in the conclusions and frame the choice between them as the next research question for the platform.

We isolate a single LLM (DeepSeek-flash without chain-of-thought, temperature 0.4, platform default prompt) to control for model-specific priors. Our focus is on the class-level difference between heuristic and LLM agents, not on inter-model variance; cross-model validation (Claude, Qwen, GLM) is deferred to a campaign-scale follow-up.