2  Why Communication Wasn’t Enough

We gave our agents a voice. We expected cooperation to follow. What followed instead was a redistribution of wealth — and a small Cohen’s d that quietly told a different story.

Roberto Brunelli · Independent Researcher · June 2026 · BDPD v1.1 · 14 min read

D1 headline: commons stock trajectories across the three implemented cells of the 2×2

The D1 mini-pilot. Three implemented cells, five seeds each (the built-in + talk cell is out of scope — heuristic bots would only mechanically replay a fixed broadcast template, not signal). The vertical separation between the LLM trajectories and the built-in trajectory is the agent architecture; the cheap-talk axis runs between the two LLM cells and barely moves them.

Here is the kind of claim that people in policy circles have repeated for a long time: talk is cheap, but it is better than nothing. Let parties to a commons signal their intentions. Let them propose pacts. Let them say, in advance, what they are about to do. Common sense — and a long tradition in mechanism design — says that even without binding enforcement, a channel for cheap talk should help. People will declare things, and others will trust some declarations, and the system will coordinate better than under silence.

On a fragile commons, with agents that can think, we tested the claim. What we found is more particular, and a little quieter, than the usual framing predicts.

2.1 The Setup, In Plain Language

The substrate is the same fragile commons described in our first paper: a logistic stock with regeneration rate \(r\) in the empirically realistic slow-regeneration band (\(r \le 0.32\), the canonical regime for commercial fisheries and slow-growth forests), populated by a handful of agents who decide each turn how much to extract. In BDPD0 we showed that this commons has a discontinuous threshold: a single aggressive agent suffices to doom it. The cliff is binary, not gradual.

For BDPD1 we added the institutional surface that Garrett Hardin’s diagnosis ignored and Elinor Ostrom’s response insisted on: communication and graduated sanctioning. The platform now ships a pact registry, a six-tool cheap-talk channel (broadcast, private message, announce-intended-harvest, propose-pact, accept-pact, pledge), five governance metrics, and both flat and graduated sanction perturbations. All of it engine-agnostic; the same surface would work over a different commons substrate without modification.

The D1 mini-pilot was the cleanest experiment we could design to ask whether cheap talk, on its own, restores cooperation against the cliff. A 2×2 factorial. Two factors, each with two levels.

NoteD1 — the design

A. Built-in heuristic agents, channel off. Five conservative bots and one aggressive bot. No way to talk. This is the BDPD0 baseline at the canonical cliff configuration.

B. LLM agents, channel on. Six DeepSeek-flash players, exactly the cheap-talk surface above. They can broadcast, announce, propose pacts.

C. LLM agents, channel off. The same six LLM players, but with no cheap-talk surface at all. The control we almost did not run.

D. Built-in agents, channel on. Left out of scope: a heuristic bot would only replay a fixed broadcast template, not signal adaptively, so the cell would conflate mechanical broadcasting with cheap talk.

Every cell, five seeds. The commons either survives all 30 turns or it does not — a binary preservation outcome that the BDPD0 cliff gives us cleanly.

2.2 What We Expected, and What Happened

We expected, walking in, the canonical answer. Built-in agents would crash in 5/5 seeds (cell A, the cliff). LLM agents with the channel would survive in some seeds, maybe 2/5, maybe 3/5 — the talk would help, the architecture would not be doing most of the work. Cell C, the LLM-without-talk control, was almost an afterthought. What could LLMs do without a channel that built-ins could not?

The reality:

  • Cell A — built-in, no talk: 0/5 seeds preserved. The cliff.
  • Cell B — LLM, talk on: 2/5 seeds preserved.
  • Cell C — LLM, no talk: 4/5 seeds preserved.
  • Built-in + talknot implemented: a heuristic bot can only replay a fixed broadcast template, not signal adaptively; the comparison would conflate mechanical broadcasting with cheap talk.

The headline is the second row of that table. The LLM agents preserved the commons in 4 out of 5 seeds without any communication channel at all. Adding the channel did not push the preservation rate from 4/5 to 5/5 or even from 4/5 to something measurably better — the difference between B and C on the continuous commons-stock metric was a Cohen’s \(d\) of \(0.17\). By the rough informal scale that effect-size literature uses, that is well below “small”. It is barely an effect at all.

“Words have no power to impress the mind without the exquisite horror of their reality.” — Edgar Allan Poe

The result is not that talk is useless. The result is that, on this cliff, the dominant axis is architecture, not channel. Whatever conservation behaviour the LLM agents are doing, they are doing it inside their own heads. The talk surface, when added, rearranges who keeps which share of what the cooperators save — but it does not change whether the commons survives.

2.3 What the Channel Did Do

We need to be precise about what cheap talk accomplished, because “it does not preserve more commons” is not the same as “it does nothing”. The channel moved one variable visibly: wealth.

In cells B (with talk) vs C (without), the aggressor’s final earnings dropped by 33%. The cooperators’ earnings rose by +17%. The total wealth across the six players was roughly conserved. What had changed was its distribution.

Read the transcripts and the mechanism is obvious. Cooperators in cell B used the announce-intended-harvest tool to signal moderation. They proposed harvest-cap pacts that bound them and only sometimes bound the aggressor. They sent private messages to each other coordinating restraint. The aggressor still defected, but in the presence of binding pacts and public announcements its defections were detectably above the agreed cap; the cooperators knew, and adjusted their own play, and trimmed the aggressor’s effective space.

The channel, on this commons, with these agents, is an instrument of equity, not preservation. It redistributes the surplus that the cooperators were already going to save. It does not save more surplus.

2.4 Why This Is Disconcerting, and Why It Should Be

The standard intuition for institutional design is that talk reduces uncertainty, that pacts coordinate expectations, that visibility constrains opportunism. There is a serious literature behind every one of those claims. None of them is contradicted by D1. Talk did reduce uncertainty about intended harvests; pacts did coordinate expectations; visibility did constrain the aggressor’s effective take. But the load-bearing variable was somewhere else — in the cognitive architecture of the agent that chose how much to take in the first place.

The uncomfortable corollary, the one that puts this finding in tension with a lot of confident policy talk, is that the choice of agent matters more than the design of the channel. A communication surface added to a population of rule-bound extractors will not, by itself, induce them to coordinate. A communication surface added to a population of signal-adaptive agents will redistribute the welfare they were already going to produce. In neither case does the surface manufacture cooperation out of nothing.

The literature on real commons institutions, from Ostrom onward, has long suspected something like this. Talk is necessary but not sufficient; institutions work when the actors inside them are capable of running the institutions. What D1 contributes is a clean quantitative form of the claim, on a substrate where we can vary the architecture and the channel independently and watch what each of them does.

2.5 Where the Real Lifting Happens

BDPD1 has three more mini-pilots — C1, D2, D3 — and they tell the other half of the story. In D2 and D3 we test graduated sanctioning against flat sanctioning, holding the expected cost constant. The graduated \([1, 3, 10]\) ladder preserves the commons in 5/5 seeds; a flat sanction of equal expected cost preserves it in only \(3/5\). A ladder-geometry sweep narrows the mechanism: the shape of the escalation, not the amount, is the load-bearing parameter. Flat-5 preserves 1/5; every graduated schedule we tried preserves at least 4/5.

Read alongside D1, the BDPD1 picture clarifies. Talk redistributes welfare. Cognitive architecture decides whether the commons is preserved at all in the absence of an explicit enforcement device. Once an enforcement device is added, the geometry of that device — graduated versus flat — is what makes it work.

The lesson maps almost too cleanly onto Ostrom’s design principles. Graduated sanctions, monitoring, the alignment of rules with the capacity of the actors who must run them. None of which is new. What BDPD1 adds is a way to see, in a controlled and replicable setting, where each of these levers does and does not move the system.

2.6 What Comes Next

D1 is a mini-pilot, \(N = 5\) at the canonical cliff configuration. It flags an effect; campaign-scale work will measure it. The directions that follow are obvious. Multi-model sweeps with \(N \gg 5\), across DeepSeek, Claude, Qwen, GLM, at multiple temperatures, are the immediate next step — the architecture-versus-channel decomposition we report here is a single-model snapshot. Whether the same ordering holds across providers, and across deeper reasoning budgets, is an empirical question we cannot answer from one pilot.

Whether the same effect appears with human players around a table is perhaps the most interesting open question of all. The card game, The Forest of Humbaba, makes the strategic structure playable. The next iteration of the project will run human-subject experiments at conferences and in classrooms. We do not know whether human players will reproduce the architecture-dominates-talk pattern. We suspect the result will be substantially more interesting than our LLM pilot — humans bring social norms, faces, reputations, things our agents do not have. The question is not whether talk helps. The question is which of the things people do when they talk are doing the work.

For now, the finding from D1 is small and precise: on the BDPD0 cliff, cheap talk is an instrument of distributional fairness, not of survival. Pacts do not save what the architecture would otherwise lose. They share differently what the architecture has already chosen to save.

Talk is cheap. So, on this particular cliff, is its effect on preservation. The pact is the architecture; the channel is the accounting.