4 Signal-adaptive vs rule-bound
The four vignettes share a pattern that none of them states in full: the qualitative direction of every governance finding depends on whether the agents under the institution are rule-bound (heuristic) or signal-adaptive (LLM). The closest methodological precedent used a spatial common-pool resource experiment to study how communication and punishment — alone, combined, and in different orders — interact in commons management among human subjects (Janssen et al. 2010); the present paper extends this experimental logic by varying the decision-making substrate itself rather than the institutional levers within it. Table 4.1 collects the evidence for V2–V4. V1 is excluded from it: as the isolation check (Section 3.2) it runs only the rule-bound baseline, so there is no second architecture against which a headline could invert — it establishes the substrate is clean, not that governance behaves one way or another.
| Vignette | Governance lever | Rule-bound (built-in) | Signal-adaptive (LLM) |
|---|---|---|---|
| V2 | Sanctioner visibility | Ladder \(1.00 \to 0.93 \to 0.86\), flat across seeds | Ladder inverts: visible deterrence suffices; delayed deterrence triggers cascade (\(21 \pm 9\) turns in Cell C) |
| V3 | Treaty enforcement | 1.0 firing, defector cap-bound after shock | 2.4 firings, defector probes boundary repeatedly; enforcement cost roughly doubles |
| V4 | Coercive exclusion | Junk purity 1.00 (type-pure sorting) | Junk purity \(0.6 \pm 0.3\) (conformists co-breach the pact; passive vs active mechanism untested at this \(N\)) |
Figure 4.1 summarises the three measurable inversions. The mechanism is straightforward, and the dichotomy is canonical: the contrast between fixed-policy heuristics and signal-conditioned reasoning has been central to bounded-rationality theory since Simon’s original formulation (Simon 1955; Gigerenzer and Gaissmaier 2011), and the present vignettes can be read as asking which institutional designs survive crossing that divide. Rule-bound agents execute a fixed strategy — harvest at the cap, ignore sanctions after absorbing them, never update on social signals. They are indifferent to institutional design: the same strategy runs regardless of whether a sanctioner is reflexive or voluntary, whether a treaty enforcer is attached, or whether exclusion is in force. Seed variation has no behavioural channel through which to enter, so built-in results are flat by construction. When human subjects are sorted into homogeneous groups by cooperative type, free-rider groups rapidly decay to zero contribution while cooperative and reciprocating groups sustain high, stable contributions (Burlando and Guala 2005) — a within-human parallel to the heuristic-vs-LLM divide observed here, with the critical difference that the variation studied in this paper operates at the architecture level rather than the preference-type level.
Signal-adaptive agents decode the institutional context and modulate their behaviour accordingly. Under strong, visible deterrence (V2 cells A–B) they self-moderate beyond what heuristics achieve — 3 violations in 30 turns vs 16 in 21. Under weak or delayed signals (V2 cell C, V3 repeated probing) they escalate, testing boundaries that heuristics never perceive. The same adaptivity that makes them more cooperative under strong institutions makes them more fragile under weak ones. Among humans, roughly half of experimental subjects are conditional cooperators — adjusting their contribution to match the group average — while about a third free-ride consistently (Fischbacher et al. 2001). The LLM agents in V2–V4 exhibit a qualitatively different pattern: they respond to institutional signals (sanctioner visibility, enforcement shocks, exclusion rules) rather than to peer behaviour alone. Even among humans, cooperative types are only moderately stable — about half of subjects retain the same classification across three measurement waves at five-month intervals (Volk et al. 2012) — so architecture-level variation (LLM vs heuristic) operates on a categorically different dimension than within-human type drift.
This asymmetry has a direct implication for governance design on this substrate: institutions calibrated against rule-bound agents will systematically under-specify the signal strength needed for signal-adaptive agents. The efficiency ladder of V2, the one-shot enforcement of V3, and the type-pure sorting of V4 are all artefacts of the rule-bound baseline in the tested institutional design space — not general properties of the governance mechanism. We do not claim that heuristic agents are always indifferent to institutional design; only that in these experimental conditions their fixed strategies cannot respond to the institutional variations tested. Any campaign-scale follow-up must therefore run both agent classes, or explicitly declare which baseline its findings rest on.
4.1 Limitations
Three limitations bound how far the cross-cutting pattern above can be generalised from this survey.
- Mini-pilot \(N\). Each vignette runs \(N = 5\) seeds. The load-bearing claims are qualitative — which cells invert and which do not — following the logic of theoretical replication (Yin 2009) rather than statistical generalisation. Quantitative effect sizes (e.g., enforcer multiplier, junk purity) are unreliable at this sample size and should be read as order-of-magnitude indicators. Bayesian credible intervals
- show strong directional separation for the wide rate gaps the vignettes turn on — for V4 the \(1/5\) (LLM, no exclusion) vs \(5/5\) (LLM, exclusion) contrast has posterior probability of difference \(> 0.95\) — while narrower comparisons (\(3/5\) vs \(5/5\)) are not resolved.
- Single LLM family. All LLM results use DeepSeek-flash without chain-of-thought at temperature 0.4. We isolate a single model to control for model-specific priors; the signal-adaptive pattern may be specific to this family’s training distribution. Model-family diversification (Claude, Qwen, GLM) is a campaign-scale extension. More broadly, LLM-based replications of human-subject experiments are known to exhibit systematic distortions such as the hyper-accuracy effect (Aher et al. 2023); the governance inversion findings reported here should be read with this caveat in mind. Reproducibility is further bounded by the inference process itself: at temperature 0.4 the model samples stochastically, so runs vary across seeds, and
deepseek-v4-flashis a cloud-served checkpoint with no user-pinnable version, so provider updates between this paper’s run date and any replication attempt may shift the underlying model. The inversion findings should therefore be read as directional (the sign and mechanism of the effect, within a cross-seed band), not as point estimates; bit-identical replication would require a locally hosted, version-pinned model or trace-based replay of the recorded responses. - No LLM cell for V1. The isolation vignette runs built-in agents only. We cannot confirm that the World abstraction remains leak-free under signal-adaptive agents, though V2–V4 (which do use LLMs inside a World) show no cross-arena artefacts.