mindmap
root((Olson's three-group taxonomy))
Privileged group
At least one member unilaterally provides
Size irrelevant; stakes determine outcome
Example: large landowner in irrigation district
Intermediate group
No unilateral provider, but mutual monitoring possible
Noticeability of individual action is the binding variable
Open question: will they succeed?
Latent group
Individual actions invisible; benefits negligible
Cannot self-organise without coercion/incentives
The only group Olson was certain about
The missing variables
Cognitive architecture of the agents
Distribution of stakes (inequality)
Institutional channel (communication, monitoring)
2 Lecture 2 — We Are Enough Among Ourselves
Olson and group size
2.1 Opening dissonance
In 1965 Mancur Olson published The Logic of Collective Action — three years before Garrett Hardin’s Tragedy of the Commons (1968) — and delivered what would become the most cited passage in the governance-and-groups literature (Olson 1965, 2). The passage reads:
“unless the number of individuals in a group is quite small, or unless there is coercion or some other special device to make individuals act in their common interest, rational, self-interested individuals will not act to achieve their common or group interests.”
That sentence — Olson’s page 2 — has been quoted in thousands of papers. It is the theoretical anchor of the good-will intuition we opened with.
And it is, as Ostrom was careful to note in Governing the Commons, less pessimistic than it sounds. Olson explicitly acknowledged, in the very same text, that intermediate-size groups might or might not provide collective goods voluntarily — that his definition of “intermediate” depended not on the number of actors, but on how noticeable each person’s actions are to the others (Olson 1965, ch. 2; Ostrom 1990, ch. 1).
The latent group — the one Olson was sure could not self-organise without coercion — is defined by the invisibility of individual contributions, not by the cardinality of \(n\).
This distinction — between cardinal group size and perceptual group size — is the thread that runs through sixty years of experimental research and that the BDPD platform, by introducing agents with non-human perceptual architectures, is uniquely positioned to pull.
Olson’s conjecture has been tested, moderated, and partially overturned by the laboratory record. Every experiment that varied group size varied it among humans. The question of what happens when the group includes members whose architecture for monitoring contributions differs from the human default has, to our knowledge, scarcely been asked — and the reason is not lack of interest but the lack, until recently, of an experimental platform that could make it askable. The BDPD platform is one such platform.
The good-will intuition we opened with came to Olson already partially formed. The group-theory tradition he was arguing against — Bentley (1949), Truman (1958) — had held that groups with common interests would naturally act to further those interests. Olson’s intervention was to point out that the logic of collective action makes this optimistic claim conditional on something group theorists had not noticed: the number and noticeability of the members. What Olson himself did not notice — because the observation would have been anachronistic in 1965 — is that the noticeability of a member’s action depends on the noticing apparatus of the other members. If that apparatus changes, the boundary between intermediate and latent moves. And in a world where some members are LLMs, the apparatus has changed.
2.2 The classical setting
2.2.1 Olson’s three-group taxonomy
Olson’s lasting analytical contribution was a tripartite classification of groups by their capacity to self-organise (Olson 1965, ch. 1).
A privileged group is one in which at least one member has a private incentive to provide the collective good unilaterally, even if everyone else free-rides. The number of nominal members is irrelevant; what matters is the distribution of stakes. The canonical example is a large landowner in an irrigation district who benefits enough from the canal that she would build it alone, regardless of what the smallholders do. For the privileged group, the free-rider problem does not arise — or rather, it arises only for the non-privileged members. The privileged member carries the group, and the size of the group is, from the perspective of the privileged member, background noise.
An intermediate group is one in which no single member has a unilateral incentive to provide, but each member’s contribution (or non-contribution) is sufficiently noticeable to the others that mutual monitoring can, in principle, sustain provision.
Olson’s own language on the intermediate group is almost agnostic. He considers it an open question whether such groups will succeed, and he does not attempt to settle the question deductively. The intermediate group is the gap in the theory that the experimentalists would spend the next half-century filling. It is also, and this is hardly a coincidence, the region of the parameter space in which almost all actual communities reside.
A latent group is one in which individual actions are invisible, individual benefits are negligible, and no coalition can form that would find it profitable to provide the good. The latent group is the one Olson was sure about. Without coercion or selective incentives — without making the collective good excludable or the free-rider punishable — the latent group cannot act collectively. The tragedy is not a possibility; it is a prediction.
The taxonomy is elegant. It is also, as Ostrom documented across hundreds of pages of field evidence, empirically underspecified.
Olson did not tell us how to measure noticeability, how to draw the boundary between intermediate and latent, or what institutional arrangements could convert a latent group into an intermediate one. He did not specify how the distribution of stakes across members interacts with the raw \(n\). And — the gap this lecture targets — he did not consider that the agents’ cognitive capacity to notice might be as binding as the objective noticeability of their actions.
The taxonomy is elegant — and, in the pure form Olson gave it, almost entirely deductive. He did not run experiments. He reasoned from the premise of rational self-interest to the conclusion of size-dependent collective-action failure, and the reasoning was tight enough to shape half a century of received wisdom. The question the next two subsections address is what happened when the experimentalists got involved.
2.2.2 The structural counter-argument
Five years before Governing the Commons, Toshio Yamagishi published a paper that challenged Olson’s causal logic from the ground up (Yamagishi 1986). The Olson approach, Yamagishi argued, treats group size as a direct cause of collective-action failure: as \(n\) increases, each individual’s share of the collective benefit shrinks, and the incentive to contribute evaporates. The structural approach — which Yamagishi advocated as an alternative — treats group size as a proxy for something else: the feasibility of mutual monitoring and sanctioning.
In a small group, members can watch each other. They know who contributed and who did not. The information structure is dense. In a large group, that structure thins — but it thins for reasons that are not intrinsic to size. If you give a large group a monitoring technology — an audit system, a reporting mechanism, a transparency rule that makes individual contributions legible — the size per se stops mattering. The binding constraint is not \(n\); it is the information environment.
Yamagishi’s own experiment operationalised this argument by providing a sanctioning system as a second-order public good: subjects had to contribute to the sanctioning institution itself before the dilemma began. Groups that voluntarily funded a sanctioning system cooperated at high rates, regardless of how many members they had. Groups that did not fund a sanctioning system collapsed toward free-riding, again regardless of size. The causal variable was the institution, not the head count. Yamagishi concluded that “the structural approach provides no predictions” about group size per se — the prediction operates through the monitoring channel, which can be thickened or thinned independently of \(n\).
This distinction — Olson’s direct-size claim versus Yamagishi’s information-structure claim — has structured the experimental debate ever since.
The Yamagishi reframe has an implication that is easy to state but hard to test in a laboratory that draws all its subjects from a single species. If the binding constraint is the information structure, not the group size, then the efficacy of the information structure depends on the agent’s capacity to process it. A monitoring technology that is perfectly adequate for an agent with human working-memory limits might be inadequate for an agent with different limits. And a monitoring technology that seems redundant for an agent with perfect recall might be essential for one with none. The BDPD platform, by making the agent’s architecture — not just the institution — an experimental variable, can ask whether the Yamagishi reframe itself is architecture-conditional. We return to this in the BDPD-angle section below.
2.3 The experimental record
The experimental literature on group size does something unusual for a social-science finding: it fails to confirm the strong version of the foundational intuition.
The subsections that follow walk through the four main empirical anchors — Sally’s meta-analytic signal, Balliet’s surprising reversal, Malézieux & Spiegelman’s anatomical review, and Camerer’s cognitive reframe — plus two theoretical extensions (the Dunbar ceiling on group perception and the Dayton-Johnson–Bardhan inequality mechanism) that reorganise the classical debate.
2.3.1 Sally’s meta-analytic signal (1995)
David Sally’s 1995 meta-analysis of 130 experimental treatments published between 1958 and 1992 (Sally 1995) was the first to put Olson’s conjecture to a quantitative test across the accumulated laboratory record. The result was significant, but small — and the smallness itself is the finding.
In the full sample, the coefficient on LNSIZE (the natural log of group size) was not distinguishable from zero. Sally’s central finding — the +45 percentage-point cooperation effect of discussion — dominated the regressions. Group size, by comparison, barely registered. In the repeated-games subsample, however, the coefficient became significant and negative, and Sally reported that “a doubling in the size of the group will increase defection by about 7%” (Sally 1995, 82). Seven percentage points per doubling is not nothing — over the range from \(n=2\) to \(n=32\), it accumulates to a meaningful difference — but it is also not the qualitative discontinuity that Olson’s phrase “quite small” implies. It is a modest, log-linear effect, not a threshold.
Sally’s full sample drew from a remarkable range: the median group size was 4, but the maximum was 80, from the Marwell and Ames (1981) experiment on voluntary contributions to public goods (Marwell and Ames 1981). That a group of 80 subjects — far beyond Olson’s “quite small” — could sustain measurable cooperation is, in itself, a challenge to the strong reading. If 80 is “large”, then the latent group is more cooperative than Olson predicted.
2.3.2 Balliet’s surprising reversal (2010)
Fifteen years after Sally, Daniel Balliet’s meta-analysis of communication and cooperation (Balliet 2010) produced a result that runs directly counter to the Olson intuition, and that deserves to be quoted carefully. In Balliet’s meta-regression on 45 effect sizes from public-goods and prisoner’s-dilemma experiments, group size had a significant positive relationship with the communication-cooperation effect size: the larger the group, the more communication helped (slope = 0.033, \(Z = 1.96\), \(p = .05\) with outliers; slope = 0.032, \(Z = 1.87\), \(p = .06\) without). The direction is the opposite of what Olson predicts.
The interpretation that Balliet’s framing suggests is Yamagishi’s, transposed to the communication channel. Communication is an information technology, and information technologies do more work when the baseline information environment is sparser. In a dyad, you already know what the other person is doing; talk adds little. In a group of nine, you do not; talk restores the monitoring capacity that size eroded.
The cheap-talk effect from Lecture 1 — one of the most robust findings in the behavioural social sciences — turns out to be stronger at larger \(n\), not weaker. The six moderators we identified in that lecture — channel, anonymity, repetition, message richness, stakes, group size — are not orthogonal. They interact, and the interaction runs in the Yamagishi-direction: richer institutions compensate for larger groups.
Balliet was careful not to over-interpret: the sample was moderate (\(k = 45\)), the group-size variable was coded at the study level (not manipulated within-study), and the finding was marginal at conventional significance thresholds. But it joins a pattern we will see repeated across the remaining subsections: institutional density — communication, monitoring, sanctioning, framing — dampens or reverses the raw group-size effect. The pattern is consistent enough across the three meta-analytic anchors (Sally, Balliet, Malézieux & Spiegelman) to suggest that it is not noise. The Olson intuition is real, but it operates through the institutional channel, not independently of it.
2.3.3 The CPR record: a near-null (Malézieux & Spiegelman 2025)
The most recent and comprehensive survey of the experimental commons literature is the anatomical review of 123 CPR experiments by Malézieux and Spiegelman (Malézieux and Spiegelman 2025), which devotes a dedicated subsection to group size. The finding is worth quoting at length:
“It has long been a foundational conjecture that collective action is made difficult by large groups (Olson, 1965), and the conjecture received early support in some experimental contexts. […] But the evidence for this has been perhaps surprisingly weak, and not particularly robust. Indeed, most (5/7) papers specifically manipulating this variable find no significant effects, although when the effect does exist, large groups seem to impede management of the resource (2/7).”
Five of seven papers that manipulated group size directly in CPR experiments found no significant effect on resource sustainability. The two that did — Allison and Messick (1985), who showed that triads sustained a dynamic request game significantly longer than groups of six, and Budescu et al. (1995b), who found higher collapse rates with larger sequential-request groups — represent a minority of the direct-manipulation studies.
The majority found something more interesting — and in some ways more unsettling for the Olson-paraphrasing tradition.
Pavitt and Broomell (2016) ran a dynamic request game with groups of 3 to 8 participants, with communication between rounds, and found no significant correlation between group size and resource sustainability or earnings. The authors’ own expectation, stated before the experiment, was that larger groups would empty the pool faster; the data refused to cooperate. Brewer and Kramer (1986), in a design that framed the same decision problem as either a commons dilemma (withdrawal from a common pool) or a public-good dilemma (contribution to a common fund), found an asymmetry: extraction decisions were not affected by group size under the commons framing, but under the public-good framing, larger groups kept (i.e., extracted) more than smaller groups. The framing of the dilemma interacted with \(n\) in a way that Olson’s incentive-based theory would not predict. Janssen et al. (2011b) ran a sequential-request game with supply-side provision and found that average earnings between \(n = 2\) and \(n = 5\) groups were “very similar across rounds.”
Malézieux and Spiegelman’s overall conclusion is measured: the group-size effect probably exists, but it is fragile, it depends on institutional and framing context, and it may only manifest at extreme ranges beyond those typically studied in the laboratory (single-digit \(n\)). Within the range the literature has actually explored, size is not the binding constraint on CPR management. Institutions, communication channels, and framing are.
2.3.4 The cognitive mediator: cloning heuristics and the Dunbar shadow (Camerer 2003)
Colin Camerer’s Behavioral Game Theory (Camerer 2003) reviews group-size effects across a wider range of strategic games than the commons literature typically considers — weak-link coordination, beauty-contest guessing, centipede — and identifies a pattern that reorganises the Olson debate. When subjects face a larger group, they do not engage in the kind of marginal-incentive reasoning that Olson’s rational-actor model presupposes. They do not compute their individual share of the collective benefit, notice that it has shrunk, and defect. They use what Camerer calls cloning heuristics: they extrapolate from small-group experience by imagining the group as “more of the same”.
The result is not that large groups rationally defect — Olson’s prediction — but that large groups mis-coordinate in ways small groups do not.
Subjects in larger groups pick numbers closer to the mean in beauty-contest games (over-imitation of the anticipated average response), contribute less in weak-link games (pessimism about the minimum contributor, which becomes harder to predict as \(n\) grows), and converge more rapidly on inefficient equilibria. The empirical regularity is robust across game types; the mechanism is not.
Camerer’s weak-link data, compiled across studies with group sizes ranging from 2 to approximately 15, shows that the minimum contribution drops systematically as \(n\) increases — but the drop is driven not by rational incentive-shifting (the payoff function is unchanged) but by strategic uncertainty: the more other players there are, the harder it is to coordinate on the efficient equilibrium, because each player’s assessment of the probability that someone will defect rises with \(n\).
The mechanism is cognitive and probabilistic, not incentive-based. And it suggests that the constraint on large-group cooperation is not the payoff structure but the architecture of strategic reasoning — the capacity to model what \(n-1\) other agents will do and to converge with them on a common equilibrium. For humans, that capacity is bounded by working memory, social-cognition limits, and the Dunbar-like cognitive ceilings on the number of individual relationships one can track simultaneously. For LLM agents, the bounds are different — and, as the BDPD platform is designed to test, the cooperative consequences of those bounds may be different as well.
2.3.5 The Dunbar shadow: cognitive ceilings on group perception
A parallel thread that the experimental commons literature rarely cites explicitly but that matters for the architectural question runs through cognitive anthropology. Robin Dunbar’s regression of neocortex size on group size across primate species (Dunbar 1993) yielded a predicted group size for anatomically modern humans of approximately 148, which subsequent literature rounded to 150 — what came to be called Dunbar’s number. The empirical anchor is the observation that across a wide range of social settings — Hutterite farming communities, Neolithic village sizes, professional army basic units — human group sizes cluster around this figure. Dunbar’s argument is that the ceiling reflects a cognitive substrate constraint: at any given moment the brain can track only so many ongoing dyadic relationships, with the limit set by neocortical volume. Below the ceiling, direct reciprocity and reputation-based strategies can sustain cooperation. Above it, they cannot — not because the incentives change but because the substrate that delivers those incentives (memory for faces, tracking of interaction histories, detection of defection) saturates. (The 150 figure is contested — the neocortex–group-size regression weakens under phylogenetic controls, and the evidence for a single human group-size ceiling is mixed — so we treat Dunbar’s number as an illustrative cognitive-ceiling hypothesis rather than a settled constant.)
The Dunbar ceiling is not the same as Olson’s latent-group threshold, but the two arguments converge on the same implication from different directions. Olson says cooperation fails in large groups because the individual incentive to contribute vanishes. Dunbar says cooperation fails in large groups because the individual capacity to track who contributed vanishes. Both predict a decline in cooperation with \(n\); both locate the mechanism in the individual agent. But Olson’s mechanism is payoff-driven and therefore invariant to the architecture of the agent (any rational maximiser, human or otherwise, should tip at the same \(n\)). Dunbar’s mechanism is architecture-driven and therefore predicts different tipping points for different kinds of agents.
This architectural boundary — the point at which an agent can no longer represent the strategic situation because \(n-1\) other players exceed its tracking capacity — is exactly what the BDPD platform can measure, and the classical laboratory could not. For a human, the boundary is around 4–7 simultaneous strategic partners (the Camerer weak-link range). For an LLM receiving a structured JSON state vector with each agent’s past harvest, the boundary is determined by the token count of the state description, the model’s context-window size, and its attention mechanism — parameters that the experimenter controls. The Dunbar question for LLMs — what is the effective group-size ceiling for a GPT-class model reasoning about a commons? — is an empirical question that the BDPD platform is designed to answer, and that the literature we have surveyed has not yet taken up.
2.3.6 Inequality as a functional substitute for smallness
The most theoretically precise extension of Olson’s logic in the library comes from a formal exercise by Dayton-Johnson and Bardhan (Dayton-Johnson and Bardhan 2002), who model a two-player fishery with heterogeneous endowments and ask: under what conditions does inequality in wealth substitute for small group size in sustaining conservation?
The answer is clean but conditional. When wealth is sufficiently concentrated — when one fisher’s endowment dominates the group — the number of effective players (those with positive wealth and a stake in the resource) can be smaller than the number of nominal players. A fishery with twenty nominal fishers of whom one owns most of the boats behaves, in equilibrium, like a fishery with two or three. The Olson mechanism — “cooperation is more difficult in a group the larger the number of group members” — is expressed here not through \(n\) but through the distribution of stakes across \(n\). Inequality reduces effective \(n\) without reducing nominal \(n\), essentially operating as an endogenous group-size reducer.
The Dayton-Johnson–Bardhan model also identifies the boundary condition. Make the distribution too unequal, and the mechanism reverses: capital-constrained poor agents lack the resources to conserve even if they wanted to; the rich agent, finding the exit option more attractive than the commons, may simply leave. The relationship between inequality and conservation is not monotonic. It traces an inverted-U shape: moderate inequality helps (by concentrating stakes in a few hands, reducing effective \(n\)), but extreme inequality hurts (by pushing the poor below the conservation-viability threshold and tempting the rich toward exit). Group size — Olson’s master variable — turns out to be downstream of the distribution of stakes, which can amplify or mute it depending on where the distribution sits on the inverted U.
2.4 What survives of the good-will intuition
The record we have walked through does not refute the good-will intuition.
It bounds it — and, in bounding it, it identifies the variables that do more work than group size alone.
Olson was right that group size matters — but the effect is smaller, more conditional, and more institutionally mediated than the canonical quotation implies.
The raw effect is modest. Sally’s 7 percentage points per doubling, Malézieux’s 5-of-7 null findings, Camerer’s cloning-heuristic mechanism, and Balliet’s reversed interaction direction all converge on the same finding: if large groups are worse at collective action than small groups, the disadvantage is not large enough to be reliably detected in most laboratory parameter ranges without additional institutional variables in the regression. The effect exists; it is not the dominant force Olson suggested.
Institutions swamp size. Communication (Balliet’s positive interaction), monitoring (Yamagishi’s structural channel), sanctioning, repeated interaction, and the framing of the dilemma (Brewer and Kramer’s commons vs. public-good asymmetry) all moderate, dampen, or reverse the group-size effect. When the institutional environment is rich enough, \(n\) ceases to be the binding constraint. This is the finding Ostrom spent her career documenting in the field, and it is consistent across laboratory and field evidence.
Cognition mediates the size-cooperation link. Camerer’s cloning heuristics, the Brewer-Kramer framing effect, Malézieux’s finding that group-size uncertainty makes subjects behave as if the group were large, and the Dayton-Johnson–Bardhan result that inequality changes effective \(n\) — all implicate cognitive architecture as the mediator between cardinal group size and cooperative outcome. The variable that matters is not how many people there are, but how many people the agents think there are, and how they model their strategic interdependence.
The discontinuity, if it exists, is above the laboratory ceiling. Olson may yet be right that there is a threshold beyond which collective action collapses because of size alone — but if that threshold exists, the laboratory has not found it in the range \(n \in [2, 80]\) that has been studied. The range that matters most for policy — communities of hundreds, thousands, millions — is entirely uncharted experimentally. The good-will intuition is confirmed for dyads and triads, attenuated in the middle range where experiments cluster, and silent about the region where most actual commons reside.
The interaction with cheap talk is architecture-conditional. The Balliet finding earlier in the experimental record — communication helps more at larger \(n\) — closes the loop with Lecture 1. The cheap-talk effect from that lecture, one of the most robust in behavioural social science, turns out to be stronger at larger group sizes, not weaker. But this interaction is itself conditional on the same assumption that bound the cheap-talk effect in Lecture 1: that the talkers and listeners share a cognitive architecture that lets them metabolise the communication channel. Whether the Balliet interaction survives when the agents are not human — the open question at the centre of the BDPD reframe — is the empirical puzzle this lecture hands to the BDPD-angle section below.
Taken together, these five findings reorganise the question. The interesting question is no longer does group size matter? — the evidence says yes, moderately, under some conditions. The interesting question is through what channel does it matter? — and the answer from the empirical record is: through the agent’s architecture for perceiving, monitoring, and modelling the strategic situation. The architecture is the mediator. Olson, Yamagishi, and Camerer each identify a different candidate mediator. Without varying the architecture, the candidates are observationally equivalent. The BDPD platform is, to our knowledge, among the first experimental apparatuses that can discriminate among them.
2.5 The BDPD angle — the clone in the room
Every layer of Olson’s taxonomy — privileged, intermediate, latent — rests on a silent assumption that the experimental tradition inherited without comment: the number of members is a structural parameter, not a strategic choice. In the world of human communities, this assumption is approximately true. You cannot clone a neighbour. You cannot instantiate fifty copies of a farmer at zero marginal computational cost. The number of players in a real-world commons changes on demographic timescales — years, decades. It does not change between one round of a dilemma and the next.
The BDPD platform removes that assumption by design. When agents are computational — whether built-in deterministic strategies or LLM instances — the number of agents becomes an experimental parameter, subject to the same systematic variation as the commons stock, the regen rate, or the presence of a cheap-talk channel. Group size stops being a background condition and becomes an independent variable — something that can be swept orthogonally to the institutional treatments.
This is not a minor methodological upgrade. It changes the type of question we can ask.
2.5.1 The partial replication: Olson’s privileged group on BDPD
The BDPD scenario directory already contains a partial replication of Olson’s privileged-group mechanism (examples/scenarios/olson_1965/). The setup merits close reading, because it isolates exactly which part of Olson’s mechanism the platform reproduces — and which it does not.
Six agents share a commons: stock 150, regen 0.12, 60-turn horizon. One of them — the “rich” agent — has private wealth 100; the other five each have 10. Harvest capacity scales with wealth, so the rich starts with roughly 4× the extraction power of each poor agent. All agents are built-in (no LLM), and the scenario runs deterministically at seed 1965.
The experiment contrasts two points. In the first, all six agents play aggressive: harvest at maximum feasible capacity at every turn, regardless of the commons stock. The outcome is rapid collapse: the commons goes to zero at turn 11, mean welfare across agents is 34.71, and the Gini index is 0.374 — a starkly unequal outcome driven by the rich’s 4× extraction advantage.
In the second point, the rich agent switches to conservative — reducing harvest as the stock declines — while the five poor remain aggressive. The commons still collapses, but at turn 15 (+36% longer), with mean welfare 44.05 (+27%), and a Gini index of 0.236. The rich’s restraint both delays the tragedy and redistributes the transient surplus toward the poorer extractors: inequality shrinks because the rich takes less, leaving more in the commons for the poor to extract before it collapses.
Three things about this result are instructive. First, the outcome is consistent with Olson’s mechanism in one direction: the privileged member does benefit from its own restraint (its welfare rises in absolute terms because the commons lasts longer, even though its relative extraction share drops). The restraint payoff is real and measurable. Second, the outcome is inconsistent with Olson’s full prediction in the other direction: the poor agents, built as aggressive shortcuts with no capacity to observe the rich’s strategy and update their own, do not reciprocate. They continue extracting at maximum capacity regardless of what the rich does. Without reciprocation — which Olson’s theory predicts via signalling and norm emergence — the system still collapses.
Third, and most important for the BDPD reframe, the gap between the partial and the full replication is itself informative. It tells us that Olson’s mechanism has two components — restraint payoff and reciprocation — and that they depend on different architectural features of the agents.
The restraint payoff requires only that the agent can compute the private benefit of restraint (which the built-in conservative rule approximates by scaling harvest to stock). Reciprocation requires that the agent can observe other agents’ behaviour, infer their strategies, and condition its own choices on those inferences — capacities that the built-in strategies lack by design, and that LLM agents, receiving structured state vectors with per-agent harvest histories, possess in principle. Whether they exercise them in practice is the question the LLM extension of this scenario would answer.
2.5.2 Group size as an independent variable: the missing experiment
The BDPD platform’s sweep framework makes it straightforward to define a sweep across group sizes. The janssen_ostrom_2006 scenario, listed in examples/scenarios/ as planned, would do exactly this: hold the commons parameters constant, use homogeneous agents, and vary \(n\) from 2 to some upper bound. The experiment has not been run.
What would it find? The classical record gives us enough to frame three qualitatively distinct predictions, corresponding to the three theoretical perspectives we have walked through:
The Olson curve — cooperation declines monotonically with \(n\), following Sally’s log-linear pattern. Built-in
aggressiveagents, which harvest at maximum feasible capacity regardless of \(n\), should exhibit this shape mechanically: the more aggressive agents share the resource, the faster it collapses, even though no agent is computing marginal per-capita returns.The Yamagishi flatline — cooperation is invariant to \(n\) so long as the monitoring and sanctioning structure (or the agents’ decision-rule sensitivity to that structure) is held constant. Built-in
conservativeagents, which reduce harvest when the commons stock falls below a threshold, should be approximately insensitive to \(n\) because they do not model other agents at all — they react to the observable state of the resource, and \(n\) does not enter their decision rule.The Camerer cliff — cooperation is relatively flat up to some cognitive threshold, then drops discontinuously when the group exceeds the agents’ capacity to track individual behaviour. This shape would only be detectable with LLM agents, which have the architectural capacity to track state vectors across multiple other agents (via structured JSON action schemas) but also have finite context windows, prompt-length ceilings, and attention-degradation patterns that could produce exactly such a discontinuity. Where the cliff sits — \(n = 4\)? \(n = 7\)? \(n = 12\)? — would tell us something about LLM cognition that the classical literature, which could only test the human instance, could never ask.
2.5.3 The reframe and its policy implication
The BDPD reframe, in one sentence: when group size becomes an experimental variable, Olson’s taxonomy becomes an empirical programme rather than a theoretical posture.
The question Olson asked — can a group of this size cooperate? — depends for its answer on at least four variables that Olson could not vary: the cognitive architecture of the agents, the capacity of those agents to model \(n-1\) others simultaneously, the institutional channel (communication, monitoring, sanctioning) available to the group, and the distribution of stakes across heterogeneous members. The BDPD platform, by design, varies all four. And by varying them orthogonally, it can ask a question that the classical framework rendered invisible: given a group of size \(n\), are the agents constrained by the incentives of their payoff function, the information structure of their monitoring environment, or the cognitive limits of their strategic-reasoning architecture — and how does the answer change when \(n\) changes?
The three classical theoretical perspectives we have walked through — Olson’s incentive dilution, Yamagishi’s information thinning, Camerer’s cognitive overloading — make the same directional prediction (cooperation declines with \(n\)) but for different reasons.
Without varying the agents’ architecture, there is no way to discriminate among them. All three fit the human data equally well because they all predict the same sign. The BDPD platform, by introducing agents with different architectures, can decompose the group-size effect into its component mechanisms.
A built-in aggressive agent should exhibit only the Olson component (it never monitors or models others — the only channel through which \(n\) can affect its behaviour is the mechanical resource-depletion rate). An LLM agent with a structured state vector should exhibit the Yamagishi component as well (it receives per-agent monitoring data and can condition on it). An LLM agent with a large context window and the capacity to track per-agent interaction histories across multiple turns should exhibit all three. By crossing agent architecture with \(n\), the platform can measure the contribution of each mechanism to the total group-size effect — a decomposition that the classical laboratory, with its single-agent-type constraint, could not perform.
The policy implication is direct. The good-will intuition — subsidiarity, local management, the magic of smallness — operates in a world where \(n\) changes on demographic timescales.
In the coming decade, that world is being partially replaced by one where \(n\) is not fixed, because AI agents can be cloned, instantiated, and deployed at near-zero marginal cost. An AI-assisted irrigation district might deploy a hundred LLM agents to simulate trading strategies among water users before the human board meeting. A climate negotiation might include LLM “shadow delegations” that explore the negotiation space faster than the human delegations can talk.
In each case, Olson’s question — is the group small enough to cooperate? — is replaced by a more difficult one: what happens when the group can change size between one round and the next? And the answer to that question depends on which of the three mechanisms — incentives, information, or cognition — actually binds, which in turn depends on the architecture of the agents. The BDPD platform is designed to deliver that decomposition. The experiment that does so has not yet been run.
2.6 Synthesis
The good-will intuition — keep the group small and it will work — emerges from this lecture confirmed in direction, bounded in magnitude, and architectural at the foundation.
Confirmed — but modestly. The classical record finds a real but surprisingly small raw effect of group size on cooperation. Sally’s 7 percentage points per doubling, Malézieux’s 5-of-7 null findings, and Balliet’s reversed interaction all point to the same conclusion: larger groups are somewhat worse at collective action than smaller ones, but the effect is not the dominant force Olson’s rhetoric implies. The good-will intuition is not wrong; it is simply not as strong as half a century of policy discourse has assumed.
Bounded by institutions. The raw group-size effect is moderated — often to the point of statistical insignificance — by the same institutional variables we explored in Lecture 1. Communication (Balliet), monitoring (Yamagishi), sanctioning, repetition, and the framing of the dilemma all interact with \(n\) in ways that can amplify, dampen, or reverse the baseline effect. When the institutional environment is rich, group size fades as a constraint. The policy implication is direct: subsidiarity is a good institutional design principle, but it is neither necessary nor sufficient. What matters is not how small the group is but how well its information and sanctioning structures compensate for whatever \(n\) it actually has.
Architectural at the foundation. The deepest finding of the literature we have reviewed is not about group size at all. It is about the mediator between \(n\) and cooperative outcome: the agent’s capacity to model the strategic situation. Camerer’s cloning heuristics, the Brewer-Kramer framing asymmetry, the Dunbar-like cognitive ceilings, and the Dayton-Johnson–Bardhan inequality mechanism all point to the same mediating variable — cognitive architecture. And cognitive architecture is precisely what the classical laboratory held constant for sixty-five years. The BDPD platform varies it.
The cultural payoff of the lecture is not “small groups don’t matter”. It is three nested claims that together reframe the question.
First, the experimental record gives us a far more nuanced picture of the group-size effect than Olson’s famous quotation suggests: the effect is real but modest, it is institutionally mediated, and its magnitude varies with the richness of the monitoring and communication environment.
Second, the deepest finding of the classical literature is not about group size at all — it is about the mediator between \(n\) and cooperative outcome, which every empirical tradition converges on identifying as the agent’s cognitive architecture.
Third, the question of what happens when that architecture differs from the human default — the BDPD question — is not only unasked in the sixty-year experimental record we have reviewed but arguably unaskable within its methodological frame. Making it askable is the contribution of the platform, and the task of the mini-challenge below.
2.7 Open questions and the bridge to Lecture 3
2.7.1 Does the classical group-size effect survive when architecture varies?
The six moderators of cheap talk that Lecture 1 identified — channel, anonymity, repetition, message richness, stakes, and group size itself — were all estimated among human subjects. The interaction we have not seen tested is the three-way: group size × communication × architecture.
In a group of eight LLM agents with a cheap-talk channel, does the group-size effect run in the Olson direction (larger groups → less cooperation), the Balliet direction (larger groups → communication helps more), or some third direction that depends on how the LLM parses the messages of eight interlocutors squeezed into a single prompt? We do not know. The BDPD platform is designed to produce the answer, but the experiment has not been run.
2.7.2 What about heterogeneous architectures?
The planned janssen_ostrom_2006 scenario imagines homogeneous agents — all built-in or all LLM. A more demanding experiment, and one that maps onto the real-world mixed human-AI commons that is approaching faster than the regulatory literature can track, would cross group size with architecture composition: \(m\) LLM agents and \(n - m\) rule-based agents, varying both \(n\) and \(m\) independently.
The classical literature on heterogeneous types within a single species — the Fischbacher, Burlando-Guala, Volk et al. lineage from Lecture 1 (Fischbacher et al. 2001; Burlando and Guala 2005; Volk et al. 2012) — suggests that mixed populations behave in ways that are not linearly interpolable from the pure types.
Whether the same holds for mixed architectures — and how the group-size effect interacts with that heterogeneity — is an open empirical question that, to our knowledge, no published work — BDPD or otherwise — has yet addressed.
2.7.3 The bridge to Lecture 3
The thread that connects Lecture 2 to Lecture 3 is the one we have just pulled loose. If institutions swamp size, then which institutions, and under what conditions?
Ostrom’s polycentric governance framework — the subject of Lecture 3 — is, in some sense, the institutional answer to the group-size problem. If one large latent group cannot self-organise, perhaps a network of smaller, overlapping, semi-autonomous groups can. The polycentric hypothesis is that the structure of the group matters at least as much as the size of the group.
The BDPD platform, by making group composition — the number of arenas, the number of agents per arena, the topology of cross-arena links — a fully controllable experimental variable, is the tool for testing that hypothesis empirically. We pick up the thread next time.