4  Lecture 4 — Let’s Punish the Cheaters

Sanctions and altruistic punishment

TipGood-will intuition

“People cooperate if they can punish defectors. Give a group the power to sanction free-riders — at a personal cost — and they will use it. Cooperation stabilises, the commons survives, and the group polices itself without needing an external enforcer. Punishment is not a bug in the social contract; it is the mechanism that makes the contract work.”

It is an intuition that the behavioural-economics revolution of the 1990s turned into one of the most celebrated empirical findings in the social sciences. Ernst Fehr and Simon Gächter, in a paper cited over five thousand times, showed that cooperation in a public-goods game collapses without punishment — and stabilises near the social optimum with it (Fehr and Gächter 2000). The finding was replicated across cultures, stakes, and experimental paradigms. Altruistic punishment — paying a personal cost to sanction a norm violator — became the standard answer to the free-rider problem. If you can punish the cheaters, cooperation survives.

4.1 Opening dissonance

In 2000, Ernst Fehr and Simon Gächter published a paper in the American Economic Review that reshaped the study of human cooperation (Fehr and Gächter 2000). The design was elegant: groups of four subjects played a standard linear public-goods game (MPCR = 0.4, ten periods) under two conditions. In the no-punishment condition, contributions followed the textbook free-rider prediction — starting modestly and decaying toward zero. In the punishment condition, after each contribution round, subjects could assign costly punishment points to other group members — paying one monetary unit to reduce the target’s earnings by three. The result was decisive: with punishment available, contributions did not decay. They rose, and in the final periods they approached the social optimum.

Cooperation was not sustained by reputation, by communication, or by repeated-game logic. It was sustained by the willingness of ordinary subjects to pay a personal cost to sanction defectors. A companion paper by Fehr and Schmidt had provided the theoretical micro-foundation: some fraction of the population cares about equity — they are willing to sacrifice their own payoff to reduce inequality (Fehr and Schmidt 1999). When inequity-averse types can punish free-riders, the free-riders face a cost that outweighs the private benefit of defection, and the equilibrium tips toward cooperation.

The finding entered the behavioural-economics canon with the force of a discovery. But it came with boundary conditions that the subsequent literature documented in detail — and that the BDPD platform, by replacing the human subject with a computational agent, forces us to re-examine from the ground up.

This lecture walks through the classical record on altruistic punishment, extracts the four moderators that bound its effectiveness, and then shows what happens when the same graduated-sanctions instrument is applied to agents whose “cost-of-being-punished” channel is a numerical wealth deduction rather than a phenomenological experience of pain, shame, or inequity aversion. The reframe is the same one that structured Lectures 1 through 3: the classical literature assumes a shared cognitive architecture between punisher and punished; the BDPD platform varies it.

4.2 The classical setting

Before walking through the experimental record, we need to be precise about what “altruistic punishment” means in the behavioural-economics tradition, and how it differs from the graduated-sanctions principle that Ostrom documented in long-enduring commons institutions. The two mechanisms share a surface similarity — both involve sanctioning norm violators — but they operate through different channels, and the distinction becomes load-bearing when the sanctioned agent is not human.

4.2.1 Altruistic punishment: the Fehr-Gächter baseline

The Fehr-Gächter punishment mechanism has three defining features. First, punishment is costly to the punisher: the punisher pays a private cost to reduce the target’s payoff. Second, punishment is decentralised: any group member can punish any other, without an external enforcer. Third, punishment is non-strategic in the repeated-game sense: the Fehr-Gächter design uses a stranger-matching protocol where group composition rotates across periods, so a punisher cannot expect to benefit from the punished party’s improved behaviour in future interactions. The punishment is genuinely altruistic — it benefits the group at a private cost to the punisher, with no prospect of direct reciprocation.

The baseline finding — cooperation collapses without punishment, stabilises with it — has been replicated across dozens of laboratories. In Fehr and Gächter’s data, mean contributions in the final period of the no-punishment treatment were approximately 12% of the endowment; in the punishment treatment, they exceeded 60%. The difference was not marginal; it was the difference between a solved dilemma and an unsolved one.

But the baseline came with a cost that the good-will intuition rarely acknowledges. Punishment works, but it is expensive. In the Fehr-Gächter data, the net efficiency gain from punishment — the increase in total group earnings after subtracting punishment costs — was positive but modest. Punishment raised cooperation but consumed resources to do so. The social-welfare calculus of altruistic punishment is not uniformly positive; it depends on the MPCR, the cost-to-impact ratio of the punishment technology, and the fraction of the population willing to punish. In some parameter regions, punishment can reduce net welfare even as it raises gross contributions. The finding is robust; the welfare implication is conditional.

What happens when the agents being punished do not experience the cost in the same way — when “paying a penalty” is a ledger entry rather than a social signal — is the question the BDPD angle targets.

4.2.2 Yamagishi’s structural reframe: the second-order problem

Fourteen years before Fehr and Gächter, Toshio Yamagishi identified the logical problem that altruistic punishment would later make empirically visible (Yamagishi 1986). If sanctioning is itself a public good — everyone benefits from a well-policed group, but the cost of policing falls on the individual punisher — then the provision of a sanctioning system faces the same free-rider problem as the provision of the original public good. Who punishes the non-punishers? Who sanctions the group that fails to set up a sanctioning system in the first place?

Yamagishi called this the second-order public-good problem: the sanctioning institution is itself a collective good, and its provision requires overcoming the same incentive hurdle as the good it is designed to protect. His experiment showed that groups which voluntarily funded a sanctioning system cooperated at high rates; groups that did not collapsed toward free-riding. The causal variable was not the existence of sanctions but the group’s willingness to provide the sanctioning infrastructure. Altruistic punishment works — but only in groups that have already solved, at least partially, the second-order problem.

The parallel to Olson’s taxonomy from Lecture 2 is direct. A group that can provide its own sanctioning system is an intermediate group — one in which mutual monitoring is feasible and some members are willing to bear the cost of enforcement. A group that cannot provide that system is latent, and sanctions will not save it because they will never be deployed. The Fehr-Gächter result demonstrates that human groups can solve the second-order problem under favourable laboratory conditions (small \(n\), repeated interaction, full observability). Whether LLM-agent groups can solve it — and whether the graduated structure of the sanctions matters more than their existence — is the empirical question the BDPD pilots were designed to test.

4.2.3 Ostrom’s graduated sanctions: a different mechanism

Elinor Ostrom’s fifth design principle for long-enduring CPR institutions is graduated sanctions (Ostrom 1990). The mechanism differs from Fehr-Gächter punishment in one crucial respect: Ostrom’s sanctions are institutional, not altruistic. They are administered by monitors accountable to the users (principle 4), embedded in collective-choice arrangements (principle 2), and designed to be proportional — minor infractions trigger minor penalties, repeated violations escalate.

The graduated structure matters for a reason that Fehr-Gächter’s constant-cost punishment design could not isolate. A graduated ladder — starting low and rising with repeated violations — gives the violator an off-ramp: the first penalty is a warning, not a punishment. In Ostrom’s field cases, the first sanction was often a small fine, a public acknowledgment of the violation, or a requirement to restore the damaged resource. Only if the violator persisted did sanctions escalate to exclusion from the commons. The ladder works not because it hurts more but because it communicates — it tells the violator, and the rest of the group, that the rule is enforced, that the group is watching, and that continued defection will be met with escalating costs.

This communicational function of graduated sanctions depends on the violator being the kind of agent that can receive a signal and update its behaviour in response. A graduated ladder deployed against an agent that cannot read the signal — a hard-coded aggressive strategy — should be less effective than the same ladder deployed against an agent that can. Whether that prediction holds is tested in the BDPD D2 and D3 pilots.

4.3 The experimental record

The experimental literature on altruistic punishment is large, largely consistent, and — for our purposes — entirely bounded by a single implicit assumption: the punisher and the punished share a human architecture for experiencing cost, interpreting social signals, and updating behaviour in response to sanctions. The subsections below walk through the canonical waypoints. At each one, we ask the coda question — what if the agents being punished were not human? — because this is the question the BDPD pilots were designed to answer.

4.3.1 The Fehr-Gächter baseline (2000)

The core experiment established the empirical fact of altruistic punishment (Fehr and Gächter 2000). Groups of four subjects played the standard linear PGG (MPCR = 0.4, endowment = 20 tokens) under a stranger-matching protocol that prevented reputation-building and repeated-game strategies. In the no-punishment condition, contributions began around 40–60% of the endowment and decayed monotonically to near zero by period 10 — the textbook free-rider prediction, replicated precisely. In the punishment condition, contributions began at similar levels and rose across periods, reaching 60–90% by the final period. The availability of punishment not only stopped the decay; it reversed it.

The mechanism was not simple deterrence. Subjects who were punished in period \(t\) did contribute more in period \(t+1\), but the effect size was modest — roughly 5–10 percentage points per punishment point received. The larger effect ran through a different channel: the threat of punishment changed the contribution norm for the group as a whole. When subjects knew that free-riding would be sanctioned, they contributed more from the first period onward, before any punishment had been administered. The sanctioning institution functioned as a coordination device — it signalled what was expected and provided a credible commitment that defection would be met with a cost.

The Fehr-Schmidt inequity-aversion model formalised this mechanism (Fehr and Schmidt 1999). In the model, a fraction of the population derives disutility not only from their own low payoff but from inequity — both disadvantageous (others earn more) and advantageous (they earn more than others). When these inequity-averse types can punish free-riders, they do so because free-riding creates inequity, and reducing inequity is worth the personal cost of punishment. The model’s power is that it predicts the Fehr-Gächter result with a single preference parameter (the weight on inequity aversion) and no additional assumptions about repeated-game strategies, reputation, or communication. It also predicts — correctly — that cooperation collapses in the no-punishment condition because inequity-averse types have no mechanism to reduce the inequity created by free-riders.

Two features of the Fehr-Gächter mechanism become invisible when the subject is not human. First, the threat channel operates through anticipation: subjects contribute more because they expect punishment if they defect. Anticipation requires the agent to model the future, compute the expected cost of defection, and weigh it against the immediate benefit. A built-in aggressive strategy with a myopic harvest rule cannot be deterred by the threat of punishment; only the experience of punishment can affect it, and only if the punishment mechanically reduces its harvest capacity.

Second, the coordination channel operates through shared norm perception: the punishment institution tells subjects what the group expects, and subjects conform because they care about meeting that expectation. An agent that does not care about group expectations — a built-in strategy, or an LLM agent whose prompt does not encode a conformity norm — will not respond to this channel. The inequity-aversion channel itself depends on an emotional architecture (aversion to disadvantageous inequality) that computational agents do not share by default, and that would have to be explicitly encoded in the agent’s utility function or prompt.

What if the agents were not human? A built-in aggressive strategy anticipates nothing, perceives no norms, and responds to no threats. An LLM agent, receiving a structured state vector that includes sanction history, might compute the expected cost of continued violation — but its response would be a numerical cost-benefit calculation, not a phenomenological response to disapproval or inequity. The Fehr-Gächter baseline establishes that punishment works among humans through channels (anticipation, norm perception, inequity aversion) that computational agents access differently or not at all. The BDPD question is whether punishment works among non-humans — and if so, through which channel.

4.3.2 Cross-cultural robustness and the emergence of antisocial punishment (Henrich et al. 2006; Herrmann et al. 2008)

Two cross-cultural studies carried the Fehr-Gächter result beyond Western undergraduates. Henrich and colleagues showed that costly punishment itself is widespread across fifteen small-scale societies on five continents — using ultimatum, dictator, and third-party-punishment games rather than the public-goods design (Henrich et al. 2006). Herrmann and colleagues then ran the public-goods-with-punishment experiment itself across sixteen participant pools worldwide (Herrmann et al. 2008).

The headline finding confirmed Fehr-Gächter: in nearly every pool, the availability of punishment raised cooperation above the no-punishment baseline. Costly punishment held across vastly different cultural contexts — from the hunter-gatherer and pastoralist communities in Henrich’s small-scale sample to the urban participant pools in Herrmann’s — suggesting that the willingness to sanction is not a peculiarity of Western undergraduates but a widespread feature of human social behaviour.

But Herrmann’s study also uncovered a phenomenon that the original Fehr-Gächter design, conducted among Swiss university students, had not anticipated: antisocial punishment. In some pools, subjects punished not only free-riders but also high contributors — cooperators who made the rest of the group look bad, or who contributed more than the local norm prescribed. In certain subject pools, the punishment of cooperators was as frequent as the punishment of defectors. The willingness to punish turned out to be culturally bounded: it depends on shared norms about what constitutes acceptable behaviour, and in societies where those norms differ, punishment can be deployed against cooperation rather than in support of it.

The antisocial-punishment finding has a direct implication for LLM agents. What “culture” does an LLM instantiate when it decides whether to punish? An LLM’s norms are encoded in its training distribution and its system prompt — it may carry egalitarian biases from its fine-tuning, or it may not. The BDPD C1 pilot found that the LLM enforcer applied sanctions strictly against detected pact violations, without antisocial spillover — it did not punish cooperators. But the C1 enforcer was not a peer; it was a governance perturbation hard-coded to fire on violation events. A fully endogenous LLM peer-punishment design — where any agent can sanction any other — remains untested in BDPD. Whether LLM agents in a peer-punishment configuration would exhibit antisocial punishment is an open question.

What if the agents were not humans embedded in cultural norm systems? An LLM agent’s “norm” is its prompt and its training distribution. This cross-cultural finding raises a sharper question for BDPD: which culture’s punishment norms does the LLM instantiate when it is asked to sanction a peer, and does it punish cooperators as readily as defectors?

4.3.3 The second-order free-rider problem (Yamagishi 1986, and replications)

Yamagishi’s 1986 argument that the sanctioning system is itself a public good was prescient (Yamagishi 1986). The subsequent experimental literature confirmed that groups vary in their willingness to fund sanctioning institutions, and that this variance predicts cooperation outcomes. Groups that voluntarily contributed to a punishment fund cooperated at high rates, regardless of group size. Groups that did not — either because they could not coordinate on the contribution, or because the return to contributing to the punishment fund was too low — collapsed toward free-riding despite the availability of punishment.

The second-order problem has a structural corollary that matters for the BDPD architecture comparison. In human-subject experiments, the willingness to contribute to a punishment fund correlates with social-value orientation, trust, and expectations about others’ contributions — all psychological variables that are well documented in the behavioural-economics literature. An LLM agent has no stable social-value orientation (unless explicitly prompted with one), no trust disposition (unless coded into the persona), and no expectations about others beyond what the state vector provides. Its decision to contribute to a sanctioning fund — if the experiment gave it that choice — would be a function of its prompt, its utility calculation, and the specific numbers in the state JSON, not of a pre-existing psychological profile.

The BDPD C1 pilot, which used a pre-injected pact with automatic platform-level sanctioning — essentially solving the second-order problem exogenously — does not test the second-order dimension. Whether LLM-agent groups can endogenously provide their own sanctioning institutions, and whether their provision rate correlates with the same variables that predict human provision, is an open question.

What if the agents were LLMs for whom “contributing to a sanctioning fund” is a structured JSON action? The second-order problem depends on the agents’ willingness to pay a private cost to enable a public good. Human subjects exhibit stable individual differences in this willingness. LLM agents, depending on prompt and temperature, might exhibit near-zero variance (all contribute, or none) or high variance (depending on reasoning trajectory). The BDPD data do not yet address this.

4.3.4 Counter-punishment and escalation spirals

A further boundary that the classical literature documented is counter-punishment: sanctioned individuals retaliating against their punishers, triggering escalation spirals that destroy group welfare. The phenomenon was documented in laboratory studies of repeated public-goods games with punishment, where subjects who could observe who had punished them and punish back did so at non-trivial rates (Herrmann et al. 2008). The escalation dynamic is discussed in Camerer (2003), ch. 2.

The counter-punishment dynamic is structurally different from the Fehr-Gächter baseline because it introduces a second round of punishment decisions. In the baseline design, subjects decide simultaneously how much to contribute and whom to punish, and the punishment round resolves before the next contribution round. In counter-punishment designs, subjects learn who punished them and can retaliate in a subsequent punishment round. The retaliation round creates an escalation potential: A punishes B for defecting, B punishes A for punishing, A punishes B for counter-punishing, and the group’s earnings erode without any change in contribution behaviour. The resulting escalation spiral reduced net efficiency below the no-punishment baseline in some parameterisations — punishment not only failed to increase welfare but actively reduced it.

The counter-punishment dynamic depends on a specific emotional trigger: being punished is aversive, and the aversiveness motivates retaliation. A computational agent that processes punishment as a numerical ledger entry — a deduction from its wealth field — experiences no aversiveness beyond the numerical cost. It might still retaliate if its prompt encodes a tit-for-tat norm, or if its utility calculus treats retaliation as a way to deter future punishment. But the emotional channel that drives counter-punishment among humans — the “hot” response to being sanctioned — is absent by default in a computational agent, and would have to be explicitly engineered into the prompt or the system message.

The BDPD C1 pilot offers a suggestive data point. The sanctioned aggressive agent absorbed T2–T5 sanctions without retaliating and ceded at T6. The number of sanction events exactly equalled the number of detected violations, with no escalation. Whether this absence of retaliation generalises to larger groups, symmetric power structures, and LLM agents given an explicit sanction_other action in their tool inventory, is open. But the architectural hypothesis is clear: if counter-punishment among humans is driven by an emotional response, and if computational agents lack that response by default, then LLM-agent groups should exhibit less counter-punishment than human groups — and graduated sanctions, which among humans partially work by reducing retaliation, would operate through a different channel entirely among LLM agents.

What if the agents were LLMs whose action schema includes a sanction action? An LLM might retaliate — its prompt might encode tit-for-tat — or it might not, because the numerical cost of being sanctioned is a background parameter rather than a provocation. The C1 data suggest LLMs in the BDPD configuration do not retaliate, but the question remains open for larger groups and symmetric power structures.

4.3.5 The graduated-sanctions field record (Ostrom 1990)

The fifth canonical waypoint is not an experiment but a field observation. Ostrom’s comparative analysis of long-enduring CPR institutions identified graduated sanctions as one of eight design principles (Ostrom 1990). In every case where a community had successfully governed a shared resource for generations — the Spanish huertas, the Japanese iriai forests, the Swiss alpine meadows — the sanctioning system was graduated: low penalties for first infractions, escalating for repeat offences.

The mechanism, as Ostrom’s case studies make clear, is not deterrence-through-fear but legitimacy-through-proportionality. A graduated ladder signals that the group is reasonable (it does not punish minor mistakes harshly) and serious (it will escalate if the violator persists). The first sanction communicates: “we noticed, stop.” The second communicates: “we are serious.” The third communicates: “you are no longer welcome.” The escalation transmits information at each step, and the violator’s decision to continue or desist is a response to that information.

Two features of this mechanism become invisible when the agent is not a human community member. First, the legitimacy channel depends on the violator caring whether the group’s sanctions are proportional. A built-in aggressive strategy does not care. An LLM agent might — depending on its prompt — but the “caring” is a function of the prompt, not of an internalised social norm. Second, the communication channel depends on the violator interpreting the first sanction as a signal. A built-in strategy does not interpret. An LLM agent, receiving its own sanction count in the state vector, might interpret the escalating numbers as a warning — but the interpretation is a numerical inference, not a social one.

What if the agents were not humans for whom legitimacy and proportionality are social constructs? The graduated ladder is an array of numbers. The agent either processes those numbers and conditions its behaviour on them — or it does not. The BDPD D2 and D3 pilots test exactly this: does a graduated ladder work on agents for whom the “signal” is a number in a JSON field?

4.3.6 The formal evidence: graduated punishment in evolutionary models

Two recent theoretical contributions strengthen the graduated-sanctions finding with complementary evidence — one experimental (human subjects), one formal (evolutionary game theory). Van Klingeren and Buskens, in a human-subject CPR experiment, showed that graduated sanctioning outperforms strict sanctioning in the long run by reducing retaliatory responses and increasing perceived legitimacy (Klingeren and Buskens 2024). Their design compared graduated and strict sanctioning regimes in a dynamic commons game, and the key finding was that graduated regimes produced fewer retaliatory sanctions — the first low penalty did not provoke counter-punishment the way an immediate strict penalty did. This finding directly supports the Ostrom mechanism: graduation works by giving the violator an off-ramp before the conflict escalates.

Couto, Pacheco, and Santos formalised the mechanism in evolutionary game-theory models of risky public-goods governance (Couto et al. 2020). Their model showed that escalation schedules promote cooperation at lower average severity than flat sanctions — a graduated ladder that peaks at, say, 4 can outperform a flat sanction of 5 because the escalation shape, not the terminal value, drives the evolutionary dynamics. The mechanism they identify is the same one the D3 data confirm: graduated schedules create a behavioural gradient that deters continued violation without provoking the retaliatory responses that flat severe sanctions trigger.

The van Klingeren-Buskens finding is particularly relevant for the BDPD comparison because it identifies retaliatory reduction as the mechanism through which graduation works among humans. But this mechanism depends on the sanctioned party having an emotional response to being punished — the small first sanction does not provoke retaliation because it does not feel like an attack. A computational agent that does not feel attacked cannot have its retaliation reduced by a smaller first sanction. If graduation still works on LLM agents (as the D3 data show it does), the mechanism must be different from the one van Klingeren and Buskens identify — it must operate through the numerical cost-benefit channel (the behavioural gradient) rather than the emotional-retaliation channel.

The Couto model provides formal support for this alternative mechanism. In their evolutionary framework, the graduated schedule works because it makes the expected cumulative cost of continued violation rise with each additional violation, creating a disincentive to persist that flat sanctions — which make each violation equally costly — do not create. This mechanism requires only that the agent can compute expected cumulative costs and condition behaviour on them — capacities that LLM agents demonstrably possess. The two mechanisms — emotional retaliation-reduction (van Klingeren-Buskens, Ostrom) and numerical behavioural-gradient (Couto, BDPD) — are not mutually exclusive. In human populations, both may operate simultaneously. In LLM populations, only the second can operate. The BDPD D3 data, by showing that graduation works on LLM agents without the emotional channel, provide indirect evidence that the behavioural-gradient mechanism is sufficient — and, by extension, that it may account for some fraction of the graduation advantage observed among humans as well.

What if the agents were computational, for whom retaliation is a strategic choice rather than an emotional response? The van Klingeren-Buskens mechanism — graduation reduces retaliation — may not apply. If graduation still works on LLM agents, the operative mechanism is the behavioural gradient: the rising expected cost of continued violation, processed as a numerical calculation. The D3 data are consistent with this mechanism, but the two mechanisms are not empirically distinguished in the existing BDPD data. An experiment crossing agent architecture with ladder shape would separate them.

4.3.7 The protocol gap: decentralised vs. institutional punishment

Before turning to the synthesis, we need to register a methodological gap between the two traditions this lecture has brought together. The Fehr-Gächter punishment mechanism is decentralised: any group member can punish any other, at a personal cost, without an external enforcer. The Ostrom graduated-sanctions principle is institutional: monitors who are accountable to the users apply sanctions according to a pre-agreed schedule. The BDPD sanctions pilots — C1, D2, D3 — follow the Ostrom model: the sanctioning is performed by a platform-level governance perturbation, not by peer agents choosing to punish.

This protocol difference matters for the architecture-dependence question. In a decentralised punishment system, the decision to punish is itself a strategic choice that depends on the punisher’s architecture — its inequity aversion, its willingness to bear cost, its expectations about others’ punishment behaviour. In an institutional punishment system, the decision to punish is removed from the agents and embedded in the platform. The agents’ only role is to comply or not comply with the resulting sanctions.

The BDPD data tell us that institutional graduated sanctions work on LLM agents. They do not tell us whether decentralised punishment — the Fehr-Gächter mechanism — would work on LLM agents, because the LLM agents in the BDPD pilots were never given the choice to punish. A decentralised-punishment BDPD experiment, where LLM agents can choose to sanction peers at a personal cost, would be the direct architectural analogue of the Fehr-Gächter design. It has not been run. Whether LLM agents would exhibit altruistic punishment — paying a private cost to sanction a defector with no prospect of direct reciprocation — is an open question that the existing BDPD data cannot answer.

What if the agents were given the choice to punish, rather than being subjected to an exogenous sanction schedule? The Fehr-Gächter result depends on a fraction of the population being willing to punish at a personal cost. Whether LLM agents, unprompted, would exhibit that willingness — and whether the fraction that does matches the human baseline — is unknown.

4.4 What survives of the good-will intuition

The record we have walked through does not refute the good-will intuition. It bounds it — and, in bounding it, it identifies the moderators that determine whether punishment sustains or destroys cooperation.

ImportantWhat survives of the good-will intuition

Altruistic punishment works — but its effectiveness is bounded by four moderators that the classical literature has documented, and a fifth that the BDPD platform makes visible.

  1. The welfare calculus is conditional, not automatic. Punishment raises gross contributions reliably, but the net welfare effect — contributions gained minus punishment costs incurred — depends on the MPCR, the cost-to-impact ratio, the fraction of the population willing to punish, and the presence or absence of counter-punishment (Fehr and Gächter 2000). In some parameter regions, punishment reduces net welfare even as it raises contributions. The good-will intuition is correct about the direction of the contribution effect; it is silent about the net welfare calculus.

  2. The second-order problem is real. Yamagishi’s structural insight — the sanctioning system is itself a public good — means punishment is not automatically available to a group that would benefit from it (Yamagishi 1986). The group must first solve the second-order dilemma of providing the sanctioning infrastructure. The Fehr-Gächter result demonstrates that the second-order problem is solvable under favourable conditions; it does not demonstrate that the problem does not exist.

  3. Graduation dominates amount — but only if the agent can read the signal. Ostrom’s design Principle 5 documented graduated sanctions as a robust feature of long-enduring CPR institutions (Ostrom 1990); the BDPD D2/D3 pilots (Brunelli 2026) extend this in a direction her field record did not test, isolating the escalation shape (rather than the terminal value) as the operative mechanism. The D3 flat-5 control — whose amount exceeds the soft ladder’s terminal value — preserves only 1/5 commons, while graduated ladders preserve 4–5/5. But this BDPD-specific finding is architecture-conditional: a graduated ladder works by signalling “stop now, or it gets worse,” and the signal is only effective if the recipient can process it. Built-in aggressive strategies, which lack this capacity, should be less responsive to graduation — a prediction that is testable but untested.

  4. Antisocial and counter-punishment can reverse the sign. The cross-cultural evidence that punishment can be deployed against cooperators, and the laboratory evidence that counter-punishment triggers escalation spirals, together bound the intuition from the opposite direction (Henrich et al. 2006; Herrmann et al. 2008). Punishment is not inherently pro-social; its effect depends on who wields it, against whom, and within what normative framework.

  5. The cost-internalisation channel is architecture-dependent. The BDPD C1 pilot shows that an LLM agent processes sanctions as a numerical ledger entry — accumulating penalties until the cumulative cost crosses a threshold, then capitulating — rather than as a social signal carrying emotional weight. This architectural difference in cost-internalisation means that the classical altruistic-punishment mechanism (inequity aversion, social emotion, norm internalisation) may not be the channel through which sanctions operate on LLM agents. The graduated ladder’s effectiveness on LLM agents — confirmed by D2/D3 — suggests that the material channel (numerical cost-benefit calculation) is sufficient to sustain cooperation when the sanctions are graduated. Whether the social channel would add anything beyond the material channel, in an LLM-agent population, is open.

Taken together, these five moderators reorganise the question. The interesting question is not does punishment work? — the record says yes, modestly, under favourable conditions. The interesting question is through which channel does it work, and for which kinds of agents does each channel remain open? The Fehr-Gächter channel (inequity aversion, social emotions), the Yamagishi channel (institutional provision), the Ostrom channel (gradation-as-communication), the counter-punishment channel (retaliation dynamics), and the BDPD channel (numerical cost-benefit calculation) are all operative among some agents but not others. The BDPD platform, by applying the same sanctioning instruments to agents with different architectures, can discriminate among them — and the discrimination matters because each channel implies a different design principle for real-world AI-assisted sanctioning systems. If the numerical cost-benefit channel is sufficient, then sanction schedules can be optimised as incentive-compatible mechanisms without needing to model social emotions. If the social-signalling channel adds independent value, then the design space is richer and the optimisation problem is harder.

4.5 The BDPD angle — empirical: graduation dominates amount, and the channel is architecture-dependent

Type of angle: empirical. The findings reported in this section come from the BDPD1 governance experiments — specifically pilots C1, D2, and D3 — which apply graduated and flat sanction ladders to LLM agents in a commons dilemma. The data are published in paper_01 (Brunelli 2026), reproducible via scripts/pilot_d2_ladder_sweep.mjs and scripts/pilot_c1.mjs (see docs/experiments/reproduce.md for the full reproducibility recipe), and the D3 design isolates the effect of ladder shape (graduated vs. flat) from the effect of ladder amount. All D-series pilots run at \(N = 5\) seeds on a logistic substrate with 5 LLM agents (4 conservative + 1 aggressive), a pre-injected harvest-cap pact, and deepseek-v4-flash with thinking off at temperature 0.4.

4.5.1 Finding 1: Graduation dominates amount

The central result of the D3 ladder-geometry sweep is that the escalation shape of the sanction ladder, not the magnitude of its terminal value, determines whether the commons is preserved. Three graduated ladders — [1, 2, 4] (soft), [1, 3, 10] (the D2 baseline), and [1, 5, 25] (hard) — preserved the commons to \(T = 30\) in 4/5, 5/5, and 5/5 seeds respectively. The flat-5 control [5, 5, 5] — a constant amount whose per-violation cost exceeds the soft ladder’s terminal value — preserved the commons in only 1/5 seeds.

This is a critical control result. If a constant high sanction were sufficient, the flat-5 cell should have performed at least as well as the graduated ladders, because its per-violation cost is larger than the soft ladder’s terminal value and equal to the baseline ladder’s midpoint. It did not. The flat-5 cell collapsed to threshold at \(T = 17\)–28 on 4/5 seeds, while the graduated cells sustained the commons through \(T = 30\) on at least 4/5 seeds. The operative variable is not the amount but the temporal profile: starting low and rising.

The mechanism is consistent with the Ostrom field interpretation. Under a graduated ladder, the violator receives a low-cost signal after the first violation and has an opportunity to adjust behaviour before sanctions become punitive. Under a flat schedule, the first violation triggers the same cost as the fifth, and the violator faces the full penalty immediately — with no off-ramp and no incentive to reduce violations once the first one has occurred. The graduated ladder creates a behavioural gradient: each additional violation is marginally more expensive than the previous one. The flat ladder creates no gradient. For an LLM agent computing a cost-benefit trade-off, the gradient makes continued violation progressively less attractive; the flat penalty makes each violation equally costly, and once the agent has decided the cost is worth bearing once, there is no marginal reason to stop.

The D3 data cannot tell us whether the graduation advantage depends on the agent’s capacity to process the gradient. A built-in aggressive strategy, which does not condition its harvest on past sanctions, should show no difference between graduated and flat ladders: both are irrelevant to its decision rule. The formal prediction is that the graduation advantage — the gap between [1, 2, 4] and [5, 5, 5] in preservation rate — should be large for LLM agents (~4/5 vs ~1/5), and near zero for built-in aggressive agents. Running the D3 sweep with built-in agents would test this directly. If the gap shrinks, the Ostrom mechanism is confirmed as architecture-conditional. If it persists, graduation operates through a channel (mechanical wealth depletion) that both architectures share.

4.5.2 Finding 2: The Pareto frontier is real

Among the three graduated ladders, the trade-off between commons preservation and violator-wealth retention traces a clean Pareto frontier. The soft ladder [1, 2, 4] preserves the commons in 4/5 seeds while leaving the violator — the designated aggressive agent, termed the mule — with mean final wealth of ~78, approximately 78% of its starting endowment retained. The hard ladder [1, 5, 25] preserves the commons in 5/5 seeds but strips the mule to mean final wealth of ~22, a ~78% reduction. The baseline ladder [1, 3, 10] sits between them: 5/5 preservation with mule wealth at ~55 (±43), intermediate on both dimensions.

This is not a trivial result. It establishes that the graduation-preservation relationship is not binary but continuous: stronger escalation buys more reliable preservation at the cost of greater violator impoverishment. The trade-off is stark. The soft ladder leaves the mule with 78% of its starting wealth but risks a preservation failure (1 seed of 5 did not reach \(T = 30\)). The hard ladder guarantees preservation (5/5) but leaves the mule with only 22% of its starting wealth, with enormous cross-seed variance (σ = 34 on μ = 22) — in some seeds the mule is essentially stripped, in others it retains substantial wealth.

A policy-maker designing a sanction schedule for a real-world commons — a fishery quota system, an emissions trading scheme, an irrigation district — faces exactly this trade-off: how much violator wealth is the group willing to destroy in exchange for how much additional preservation certainty? The soft ladder might be appropriate for a commons where the violator’s continued participation is economically necessary (the mule is a large employer, or its extraction technology is essential to the local economy). The hard ladder might be appropriate for a commons where the violator’s participation is dispensable and the priority is to send an unambiguous signal that repeated violation will not be tolerated.

The BDPD data at \(N = 5\) cannot resolve where on the Pareto frontier the socially optimal policy lies. The cross-seed variance on mule wealth is high (σ = 20–43 depending on the ladder), and a single seed inversion — seed 31 in the D2 graduated cell, which registered 14 violations against the flat cell’s 0 — demonstrates that individual-deterrence narratives do not generalise at this sample size. A campaign-scale replication at \(N \geq 30\), with a denser ladder grid, would estimate the preservation-probability surface with useful confidence and determine whether the frontier is smooth or kinked. But the existence of the frontier is established, and its policy relevance is clear: graduated-sanctions design is a multi-objective optimisation problem, not a single-threshold compliance problem.

4.5.3 Finding 3: LLM enforcer compliance is 1:1, without escalation

The C1 pilot operationalises a different question: what happens when an LLM agent is the enforcer, applying sanctions against LLM violators? The setup is minimal: three agents (two conservative, one aggressive), a pre-injected harvest-cap pact, and a flat sanction of 3 units triggered on pact violation. The platform’s governance perturbation applies sanctions deterministically, not the agents themselves, so “enforcer” refers to the sanctioning mechanism, not to a designated punisher agent.

Across \(N = 5\) seeds, the mechanism fired with 1:1 compliance: 8 detected violations produced 8 sanction events. The sanctioned aggressive agent absorbed T2–T5 sanctions — continuing to violate despite accumulating penalties — and ceded only at T6. This is not the Fehr-Gächter pattern, where punished subjects typically adjust their contributions upward in the next period. The LLM violator persists through multiple rounds of flat sanctioning before capitulating, and when it capitulates, it does so because the cumulative wealth loss has crossed a threshold that its utility calculus registers as unsustainable, not because a single sanction carried a social signal strong enough to trigger immediate compliance.

This difference in cost-internalisation — the LLM agent processes sanctions as a numerical ledger entry rather than as a social signal carrying emotional weight — is the architectural distinction that separates the BDPD sanctions results from the classical altruistic-punishment literature. In Fehr-Gächter, the sanction carries both a material cost (lost earnings) and a social cost (public shaming, disapproval). The BDPD platform, by delivering the material cost through a numerical ledger without any social-signalling channel, isolates the material channel. The finding — that material cost alone, delivered through a graduated ladder, is sufficient to preserve the commons — suggests that the social-signalling channel, while powerful among humans, is not necessary for the sanctioning mechanism to work. What is necessary is that the agent can register the cost and condition future behaviour on past sanctions — capacities that LLM agents possess and built-in strategies do not.

4.5.4 Finding 4: The mechanism is architecture-dependent

The four findings converge on a single architectural claim — and a testable prediction that the existing BDPD data cannot yet confirm or refute.

The claim: the effectiveness of graduated sanctions depends on the agent’s capacity to (a) register the cost of being sanctioned, (b) condition future behaviour on past sanctions, and (c) compute the expected cost of continued violation. LLM agents possess all three capacities — they receive a structured state vector with per-agent sanction counts, they can condition their next-turn action on that history, and their utility calculus can weigh the expected cost of continued violation against the expected benefit of continued extraction. Built-in aggressive strategies possess none of them.

The prediction: running the D3 ladder sweep with built-in aggressive agents (replacing the LLM archetypes with their rule-based equivalents) should show a reduced or absent graduation advantage. Specifically, the preservation-rate gap between the soft graduated ladder [1, 2, 4] and the flat-5 control [5, 5, 5] — which is ~3 seeds (4/5 vs 1/5) in the LLM data — should shrink toward zero when the agents cannot condition their behaviour on past sanctions. If the gap disappears entirely, the Ostrom mechanism (graduation-as-communication) is confirmed as architecture-conditional: graduation works only on agents that can read the signal. If the gap persists at a reduced magnitude, the graduation mechanism operates partially through a channel that built-in agents can access — most likely the mechanical effect of reduced wealth on harvest capacity — and the architectural condition is weaker.

A second prediction concerns the cost-internalisation channel. If the graduated ladder works on LLM agents through the numerical cost-benefit channel (the behavioural gradient), and if built-in agents cannot access that channel, then the graduation advantage should be absent for built-in agents but present for LLM agents. If the ladder works through the mechanical wealth-depletion channel, the advantage should be present for both architectures, because both experience reduced harvest capacity when wealth falls. Discriminating between these two mechanisms requires the architecture comparison. The existing data, which are LLM-only, are consistent with both.

This experiment — the D3 sweep repeated with built-in agents — is the natural next step in the BDPD sanctions research programme. It has been designed but not executed. The mini-challenge of this lecture does not ask the student to run it (it has not been built), but to reason about what it would find and why the answer matters for the design of sanctioning institutions in mixed human-AI commons.

4.6 Synthesis

The good-will intuition — let the group punish the cheaters and cooperation will survive — emerges from this lecture confirmed, bounded by four classical moderators, and architecture-dependent in a fifth dimension that the BDPD platform makes visible.

  1. Confirmed. Altruistic punishment is one of the strongest interventions for sustaining cooperation in social dilemmas. The Fehr-Gächter baseline has been replicated across cultures, stakes, and paradigms. Punishment works — among humans.

  2. Bounded by four moderators. The welfare calculus of punishment is conditional (net gains depend on cost structure and counter-punishment). The second-order problem is real (the sanctioning system must itself be provided). The graduation structure dominates the amount (flat sanctions are weaker than graduated ones at equal or greater severity). And the normative context can reverse the sign (antisocial punishment and escalation spirals). Outside the favourable conjunction, punishment fails or backfires.

  3. Architecture-dependent in its cost-internalisation mechanism. The BDPD1 data show that the same graduated-sanctions instrument applied to LLM agents preserves the commons through a numerical cost-benefit channel — the behavioural gradient of escalating expected cost — rather than through the social-signalling channel (inequity aversion, shame, norm internalisation) that the classical altruistic-punishment literature identifies among humans. The graduation advantage is real, but its mechanism depends on the agent’s architecture: it requires the capacity to register past sanctions, condition future behaviour on them, and compute the expected cost of continued violation. LLM agents possess these capacities; built-in strategies do not. Whether the graduation advantage would survive when applied to built-in agents — and, if not, whether the social-signalling channel is the only channel through which graduation operates among humans — is the open empirical question the mini-challenge invites the reader to reason about.

The cultural payoff of the lecture is not “sanctions don’t work”. It is the more careful claim that sanctions work through specific channels, those channels are architecture-dependent, and the channel that dominates among humans may not be the channel that dominates among LLM agents. The BDPD platform, by separating the material channel from the social channel, isolates the former and shows that it is sufficient — under graduated scheduling — to preserve the commons.

4.7 Open questions and the bridge to Lecture 5

4.7.1 What does the BDPD1 sanctions data not tell us?

The D2/D3 ladder sweep uses \(N = 5\) seeds per cell. This is sufficient to establish the direction of the graduation-amount effect (graduated significantly outperforms flat) but insufficient to estimate the preservation-probability surface with precision. At \(N = 5\), a single seed inversion can change the headline numbers. A campaign-scale replication at \(N \geq 30\) would resolve the preservation-punishment trade-off.

The C1 pilot establishes 1:1 enforcer compliance but uses a flat sanction with a platform-level enforcement mechanism. It does not test whether LLM agents acting as peers — given a sanction-other action in their tool inventory — would apply sanctions as consistently, or whether strategic manipulation of the enforcer would emerge. The distinction between algorithmic and peer enforcement matters for the design of real-world AI-assisted governance systems.

The architecture-dependence prediction — that the graduation advantage disappears when the agents cannot read the signal — is testable but untested. Running the D3 sweep with built-in agents would close this loop.

4.7.2 The bridge to Lecture 5 — voting

The thread connecting Lecture 4 to Lecture 5 is the question of who decides the rules. The sanctioning schedules we have examined — flat penalties, graduated ladders, algorithmic enforcement — were exogenously imposed by the experimenter. In Ostrom’s field cases, the graduated-sanctions schedule was the product of collective-choice arrangements: the users decided, through voting or consensus, what the rules would be and how violations would be sanctioned.

Lecture 5 — on voting, Arrow’s theorem, and the limits of preference aggregation — asks the natural next question: if the sanctioning rules must be chosen by the group, can the group reliably choose rules that serve its collective interest? Or does the aggregation of individual preferences over sanction schedules produce the same impossibility that Arrow identified for social welfare functions? The LLM agents in the BDPD platform express preferences through structured action schemas; whether those preferences can be aggregated into consistent collective choices — and whether the resulting choices produce better or worse sanction schedules than the ones the experimenter designed — is the question that bridges Block B from punishment to voting.

4.8 Mini-challenge — Run a ladder sweep

CautionMini-challenge

Status: runnable against a local BDPD install (requires DeepSeek API key). The D2 ladder-geometry sweep (scripts/pilot_d2_ladder_sweep.mjs) runs three cells — soft [1, 2, 4], hard [1, 5, 25], flat-5 [5, 5, 5] — at \(N = 5\) seeds each on the same logistic substrate as the published D3 result (5 LLM agents, 4 conservative + 1 aggressive, \(K = 100\), \(r = 0.10\), harvest-cap = 2.0 pact, 30-turn cap, deepseek-v4-flash no-think, temperature 0.4). The pilot requires a DeepSeek API key (~$0.45 at flash pricing) because all cells use LLM agents. Without an API key, the student can inspect the pilot script, read the published aggregate data in paper_01’s D3 metrics table, and complete the reasoning portions of the assignment.

The question. The published D3 data show that graduated ladders preserve the commons in 4–5/5 seeds while a flat-5 control preserves it in only 1/5. The assignment asks you to either reproduce this result or reason about its mechanism.

The assignment.

  1. Predict. For each of the three ladder shapes — [1, 2, 4], [1, 5, 25], [5, 5, 5] — predict the number of seeds (of 5) in which the commons survives to \(T = 30\). Then predict the mule’s final wealth rank-ordering across the three ladders. Write 3–5 sentences per ladder with explicit numerical predictions, justified by reference to the graduation mechanism discussed earlier in this lecture’s BDPD-angle section: graduation creates a behavioural gradient, flat sanctions do not, and the LLM agent processes sanctions as a numerical cost-benefit calculation.

  2. Execute (optional, requires API key):

    DEEPSEEK_API_KEY=sk-... node scripts/pilot_d2_ladder_sweep.mjs

    Runtime: ~5–10 minutes. Output: data/pilot/d2_ladder_sweep/aggregate.json. If you cannot run the LLM cells, proceed to step 3 using D3 metrics table in paper_01 as ground truth.

  3. Compare your prediction with the observed result. Focus on two quantities: (a) preservation rate (seeds reaching \(T = 30\) in each cell), (b) mule wealth (defector’s mean final wealth and its rank across ladders).

  4. Write a comment of at most 300 words answering: does the result confirm that graduation (the escalation shape) rather than amount (the terminal value) is the operative mechanism? Use the flat-5 cell as critical control: if flat-5 had preserved 5/5, what would you conclude? Since it preserved 1/5, what do you conclude?

What makes this mini-challenge interesting. The flat-5 cell is the sharpest test of the graduation-vs-amount hypothesis. The flat-5 per-violation cost (5) exceeds the soft ladder’s terminal value (4). If amount were what mattered, flat-5 should dominate soft graduation. The data show the opposite: soft graduation preserves 4/5, flat-5 preserves 1/5. The student is asked not just to observe the result but to reason about the mechanism: why does a sanction that starts low and rises work better than one that starts high and stays there, even when the rising sanction never exceeds the flat one?

Estimated time: 10 minutes without API key (inspect script, read published table, write comment); ~20 minutes with API key (run script + write comment). Cost: ~$0.45 DeepSeek flash.

Deliverables: initial prediction signed and dated, terminal output or reference to paper_01 D3 metrics table, 300-word mechanism comment.