9  Lecture 9 — If We Meet Again, We Cooperate

The folk theorem and its bad equilibria

TipGood-will intuition

“If people know they will meet again, they will cooperate. The shadow of the future disciplines defectors: cheat today and you lose tomorrow’s joint surplus. In any ongoing relationship — business partners, neighbouring farmers sharing an aquifer, nations in a trade bloc — the prospect of future interaction converts the one-shot free-rider problem into a repeated game in which cooperation is self-enforcing. The longer the horizon, the stronger the discipline. If only we could make every interaction a repeated one, the tragedy of the commons would dissolve.”

It is an intuition that underwrites an enormous amount of institutional design. International trade agreements are structured as repeated negotiations precisely so that the shadow of the future can discipline compliance. Common-pool resource management relies on the ongoing relationship among users — fishers who see each other at the dock every season, farmers who share irrigation infrastructure for decades. Corporate joint ventures embed repeated interaction into the contract structure. The logic seems airtight: if we will meet again, defection is too costly, so we cooperate. Repetition saves cooperation.

9.1 Opening dissonance

The intuition is old, and its formalisation is older than most people think. In 1971, James Friedman proved that in an infinitely repeated game, any outcome that gives each player more than their stage-game Nash equilibrium payoff can be sustained as a subgame-perfect equilibrium, provided the players are sufficiently patient (Friedman 1971). The result was later generalised by Fudenberg and Maskin to the case of imperfect public information (Fudenberg and Maskin 1986). The popular reading took the optimistic half: repetition saves cooperation. If the game is repeated often enough and the players care enough about the future, cooperation is an equilibrium. The shadow of the future is long enough to discipline the present.

The popular reading left out the second half. The same theorems prove that almost any outcome — including full mutual exploitation, including outcomes that are Pareto-inferior to the one-shot Nash equilibrium — can also be sustained as a subgame-perfect equilibrium. The folk theorem does not say that repetition leads to cooperation. It says that repetition leads to multiplicity — a vast set of equilibria among which the theory provides no selection criterion. Cooperation is in the set. So is permanent defection. So is any pattern of alternation, punishment, or exploitation that satisfies the individual-rationality constraint. The folk theorem is as much a curse as a blessing: it tells us that the shadow of the future can sustain anything, and therefore selects nothing.

This is the dissonance the lecture is built around. The first half will walk through the classical contributions that constructed the folk theorem and the experimental record that tested it. The second half will ask what the theorem’s multiplicity looks like when the players are language-model agents whose discount factor — their effective patience — is induced by a system prompt rather than by an intrinsic preference over future payoffs.

9.2 The classical walk: from repeated games to the folk theorem

We walk through the key contributions in roughly the order they were made. At each waypoint, we ask the same question — this result assumed the players are human; what if they weren’t? — because this is the question that BDPD will eventually answer.

9.2.1 The foundational insight: repetition changes everything

Before the folk theorem had a name, the insight that repetition changes strategic behaviour was already present in the Cold War deterrence literature. The logic of mutually assured destruction — that two nuclear powers would refrain from first strike because the retaliatory second strike made defection suicidal — was implicitly a folk-theorem argument: the repeated interaction between superpowers sustained a cooperative equilibrium (no first use) that would not survive in a one-shot game. Thomas Schelling’s work on focal points and coordination, though focused on a different mechanism, shared the assumption that ongoing relationships create strategic possibilities that one-shot encounters do not.

The formal apparatus that would give these intuitions mathematical precision emerged from two directions. The first was the theory of repeated games itself, which asked: if a base game \(G\) is played infinitely many times by the same players, what equilibria exist in the repeated game \(G^\infty\)? The second was the theory of dynamic programming, which provided the discount-factor formalisation that makes “patience” a measurable quantity. A player with discount factor \(\delta \in (0, 1)\) values a payoff received \(t\) periods from now at \(\delta^t\) times its current value. As \(\delta \to 1\), the player becomes perfectly patient — future payoffs matter as much as present ones. As \(\delta \to 0\), the player becomes perfectly myopic — only today’s payoff matters.

The folk theorem lives in the gap between these extremes. It says that for any \(\delta\) sufficiently close to 1, the set of subgame-perfect equilibrium payoffs in the repeated game is the set of feasible and individually rational payoffs — the set of outcomes that (a) can be achieved by some correlated mixture of actions and (b) give each player at least their minimax payoff (the worst payoff another player can force on them). This set is typically vast: in a two-player prisoner’s dilemma, it includes the full cooperation payoff, the full defection payoff, and every convex combination of the two that gives each player at least their minimax value. The folk theorem does not select among these outcomes. It merely says they are all sustainable.

What the foundational insight assumed — and what every subsequent contribution in this lineage has assumed — is that the discount factor \(\delta\) is a fixed parameter of the player’s utility function: a deep preference parameter, like risk aversion or time preference, that the player brings to the game exogenously. The literature has no theory of where \(\delta\) comes from, because in the human case it is a psychological primitive — part of the agent’s endowment, not an engineering choice. What happens when \(\delta\) is not a primitive but a prompt — when the system message says “you will play 50 rounds” or “you will play 5 rounds”, and the agent’s effective patience is induced by that instruction rather than by an intrinsic preference? This is the question the foundational literature could not formulate.

9.2.2 Friedman’s theorem (1971)

James Friedman’s 1971 result (Friedman 1971) was the first formal statement of what would become the folk theorem. Working in the context of oligopoly theory — firms competing repeatedly in a market — Friedman proved that in an infinitely repeated game, any outcome that gives each player more than their stage-game Nash equilibrium payoff can be sustained as a subgame-perfect equilibrium, provided the discount factor is sufficiently high. The proof was constructive: each player uses a trigger strategy — cooperate as long as everyone has cooperated in all previous periods, and revert to the stage-game Nash equilibrium forever if anyone has ever deviated. If the discount factor is high enough, the future loss from this Nash-reversion punishment outweighs the one-shot gain from deviation, and cooperation is sustained. The harsher minimax punishments — which sustain the larger set of payoffs above each player’s minimax value — come later, with the Aumann-Shapley and Fudenberg-Maskin generalisations.

The result was elegant and disturbing in equal measure. It was elegant because it provided a rigorous foundation for the intuition that repetition enables cooperation. It was disturbing because it provided exactly the same foundation for the intuition that repetition enables any outcome. The trigger-strategy construction works for any individually rational payoff, not just the cooperative one: if the players agree (implicitly or explicitly) on a particular division of the surplus and use trigger strategies to enforce it, that division is a subgame-perfect equilibrium. The agreement could be fair or unfair, efficient or inefficient, symmetric or grossly asymmetric. The folk theorem does not discriminate.

Friedman’s result was stated for the case of perfect monitoring — each player observes all previous actions of all players before choosing their own. This is a strong informational assumption: it requires that defection is immediately and perfectly detectable. The subsequent literature would spend two decades weakening this assumption, eventually arriving at the Fudenberg-Maskin result for imperfect public information. But the core insight — that repetition plus patience plus trigger strategies sustains any individually rational outcome — was already present in Friedman’s original paper.

The assumption that made the theorem work — that the discount factor \(\delta\) is a fixed parameter of the player’s utility function — was so natural that no one remarked on it. Of course the players are patient or impatient by nature. Of course the discount factor is a preference parameter. The question of whether \(\delta\) could be engineered — set by an external designer rather than endowed by nature — was not on the agenda, because the only agents the theory had ever modelled were human beings with stable time preferences.

9.2.3 The Aumann-Shapley generalisation (1976)

Five years after Friedman, Robert Aumann and Lloyd Shapley extended the folk theorem to a broader class of repeated games (Aumann and Shapley 2013). Working within Friedman’s perfect-monitoring assumption, their contribution was to broaden the solution concept: they characterised the full set of Nash equilibrium payoffs of an infinitely repeated game, showing that any feasible payoff vector that gives each player at least their minimax payoff can be sustained — not just the cooperative outcome highlighted by Friedman’s trigger-strategy construction. The Aumann-Shapley result established that the folk theorem multiplicity is not an artefact of one particular punishment construction; it is a robust feature of repeated interaction itself.

The shift in solution concept matters for what comes next in our story. The Nash equilibria of the repeated game are a strict superset of the subgame-perfect equilibria. A Nash equilibrium requires only that no player can profitably deviate when everyone else follows the equilibrium strategy. A subgame-perfect equilibrium additionally requires that the strategies form a Nash equilibrium in every subgame — including subgames that follow a deviation. Friedman’s trigger strategies are subgame-perfect; some of the broader Nash equilibria covered by Aumann-Shapley are not, because their prescribed punishments would not be credible after the deviation actually occurred. The credibility distinction matters because subgame-perfect equilibria are more demanding — threats must be carried out — and the set of subgame-perfect equilibrium payoffs is a strict subset of the set of Nash equilibrium payoffs.

For the BDPD reframe, the credibility distinction is directly relevant. Trigger strategies require that the punishment be credible — that the punisher actually carries out the threatened response when a deviation occurs. For human players, credibility is maintained by reputation, by social norms, and by the intrinsic motivation to punish defectors (a motivation that the experimental literature on altruistic punishment has documented extensively). For LLM players whose strategies are induced by prompts, credibility is a design choice: the system prompt can instruct the agent to “always punish defection” or to “punish defection only if it is profitable to do so”. The class of credible trigger strategies available to LLM agents is a subset of the class available to human agents — or, more precisely, it is a different subset, defined by the prompt architecture rather than by social norms.

9.2.4 The Fudenberg-Maskin theorem (1986)

The most powerful version of the folk theorem was proved by Drew Fudenberg and Eric Maskin in 1986 (Fudenberg and Maskin 1986). Their theorem addressed the case of imperfect public information — the players do not observe each other’s actions directly, but observe a public signal (a noisy, aggregated outcome) that is correlated with the action profile. This is the informational structure most relevant to real-world repeated interactions: firms observe market prices (a noisy signal of competitors’ actions), fishers observe stock levels (a noisy signal of others’ extraction), nations observe trade balances (a noisy signal of compliance with trade agreements).

Fudenberg and Maskin proved that the folk theorem holds under imperfect public information, provided two conditions are satisfied. First, the full-dimension condition: the set of feasible payoffs must have the same dimension as the number of players. This condition ensures that the players have enough “degrees of freedom” to sustain any individually rational payoff through appropriate punishment strategies. Second, the identifiability condition: deviations by any single player must be detectable from the public signal with sufficient probability. This condition ensures that the punishment strategies can target the deviator rather than punishing everyone indiscriminately.

The Fudenberg-Maskin theorem is the version of the folk theorem that the modern literature treats as canonical. It is also the version that makes the multiplicity problem most vivid. Under imperfect public information, the set of subgame-perfect equilibrium payoffs is still vast — it includes the full cooperation payoff, the full defection payoff, and everything in between — but the strategies that sustain these outcomes are more complex. Instead of simple trigger strategies (cooperate until someone defects, then punish forever), the equilibrium strategies may involve forgiving punishments (punish for a finite number of periods, then return to cooperation), correlated punishments (punish only if the public signal is sufficiently bad), or asymmetric punishments (punish different players differently depending on their role in the deviation).

The complexity of these strategies raises a question that the classical literature could not answer: can the agents actually implement them? Trigger strategies require memory (recording the history of play), computation (evaluating whether a deviation has occurred), and commitment (carrying out the punishment). For human players, these capacities are bounded by cognitive constraints — memory is imperfect, computation is slow, and commitment is weakened by social norms that discourage prolonged punishment. For LLM players, memory is bounded by the context window, computation is near-instantaneous, and commitment is determined by the system prompt. The class of trigger strategies that an LLM can implement is not a subset of the class that a human can implement — it is a different class, defined by different constraints.

9.2.5 Kreps, Milgrom, Roberts and Wilson: reputation in finitely repeated games (1982)

The folk theorem applies to infinitely repeated games. But most real-world interactions are finite — business partnerships end, trade agreements expire, fishing seasons close. In a finitely repeated game with a unique Nash equilibrium in the stage game, backward induction dictates that the Nash equilibrium is played in every period: in the last period, there is no future to discipline defection, so both players defect; in the second-to-last period, knowing that both will defect in the last period, there is no future discipline, so both defect; and so on, all the way back to the first period. The folk theorem, on this argument, should not apply to finite games.

In 1982, David Kreps, Paul Milgrom, John Roberts and Robert Wilson proved that this argument fails when there is incomplete information about the players’ types (Kreps et al. 1982). If even one player has a positive probability of being an “irrational” type — a player who cooperates unconditionally, or who uses Tit-for-Tat, or who has any strategy that sustains cooperation — then rational players will cooperate for a substantial initial portion of the game to build a reputation for being the irrational type. The rational player cooperates not because they value cooperation intrinsically, but because cooperation is the best way to signal that they might be the kind of player who cooperates unconditionally — and this signal is valuable because it induces the other player to cooperate in return.

The Kreps-Milgrom-Roberts-Wilson result is important for the BDPD reframe because it shows that the folk theorem’s multiplicity can be partially resolved by incomplete information. In the complete-information case, any individually rational outcome is sustainable, and the theory provides no selection criterion. In the incomplete-information case, the set of equilibrium outcomes is smaller — cooperation is sustained for a substantial fraction of the game, and the exact fraction depends on the prior probability of the “irrational” type. The reputation mechanism selects the cooperative equilibrium from the folk-theorem multiplicity, not because cooperation is intrinsically better, but because it is the best way to signal a cooperative type.

For LLM agents, the reputation mechanism takes on a different character. An LLM’s “type” is determined by its system prompt, its training distribution, and its context window — not by an intrinsic disposition that other players must infer. If both players know each other’s system prompt (as they do in many BDPD scenarios), there is no incomplete information about types, and the reputation mechanism does not operate. If the system prompts are private, the mechanism may operate — but the “irrational type” that sustains cooperation is not a human with social preferences but an LLM that has been instructed to cooperate. The folk-theorem multiplicity, in the LLM case, is resolved (if at all) by prompt design rather than by reputation.

9.2.6 The experimental record: does repetition actually sustain cooperation?

The folk theorem is a theoretical result about what can happen. The experimental record asks what does happen when human subjects play repeated games in the laboratory. The two literatures have developed in parallel, and the gap between them is instructive.

The earliest experimental tests of the folk theorem’s predictions were conducted in the 1980s and 1990s, using finitely repeated prisoner’s dilemmas with human subjects. The consistent finding was that cooperation does emerge in repeated interactions — subjects cooperate at rates well above the one-shot Nash prediction — but the pattern is not the one the folk theorem predicts. In the complete-information case, backward induction should produce defection in every period; in the laboratory, subjects cooperate for a substantial fraction of the game and defect only in the final few periods. The “end-game effect” — a sharp drop in cooperation in the last two to three periods — is one of the most replicated findings in the experimental literature.

The evolutionary-game-theory tradition provided a complementary perspective. Computer tournaments of finitely repeated prisoner’s dilemmas — in which strategies compete in round-robin pairings and the most successful strategies propagate — showed that cooperation can emerge and stabilise through evolutionary dynamics. The winning strategy in the most famous of these tournaments was Tit-for-Tat: cooperate on the first move, then copy the opponent’s previous move. Tit-for-Tat won not because it was the best strategy against any particular opponent, but because it performed well on average across a diverse population. It was “nice” (never first to defect), “retaliatory” (immediately punishes defection), “forgiving” (returns to cooperation when the opponent does), and “clear” (easy for opponents to predict). The evolutionary perspective showed that the folk-theorem multiplicity is not just a theoretical possibility — in a population of interacting agents, multiple equilibria coexist, and the one that is realised depends on the initial distribution of strategies and the evolutionary dynamics.

The modern experimental record on repeated games has been substantially enriched by the work of Dal Bó and Fréchette (Dal Bó and Fréchette 2018), who conducted a systematic review of infinitely repeated prisoner’s dilemma experiments. Their findings confirm the folk theorem’s qualitative prediction: cooperation rates increase with the discount factor. When subjects know they will interact for a long time (high \(\delta\)), cooperation is substantially higher than when they know the interaction will be short (low \(\delta\)). But the relationship is not as clean as the theorem predicts. Cooperation rates are well below 100% even at high \(\delta\); they are sensitive to the history of play (not just the current strategy); and they vary substantially across subject pools, framing effects, and protocol details. The folk theorem predicts a sharp threshold — cooperation is sustainable above a critical \(\delta\) and unsustainable below it. The experimental record shows a gradual increase with substantial noise.

The most recent contribution to the experimental record comes from the LLM game-playing literature. Akata et al. (2025) report the results of finitely repeated \(2 \times 2\) games played by different LLMs against each other, against human-like strategies, and against actual human players. Their headline finding is that LLMs “perform particularly well at self-interested games such as the iterated Prisoner’s Dilemma family” — they cooperate at rates that are competitive with or above human baselines. But they also find that LLMs “behave suboptimally in games that require coordination, such as the Battle of the Sexes” — suggesting that the cooperative capacity of LLMs is game-dependent, not uniform. The Akata finding is consistent with the broader BDPD pattern: LLM agents cooperate more than rule-based agents in social dilemmas, but the mechanism is different from the one the classical literature documents in humans.

Brookins and DeBacker (2024) complement the Akata work with a tighter two-game design — the dictator game and the prisoner’s dilemma — and compare GPT-3.5 to human laboratory baselines. Their headline result is that GPT-3.5 is more prosocial than typical humans on both games: it allocates more equitably in the dictator game than human participants do, and it cooperates in the prisoner’s dilemma at about 65%, against a human baseline of roughly 37%. The finding is consistent with the broader BDPD pattern — LLM agents cooperate more than rule-based agents in social dilemmas — but it also raises a measurement-validity question. If the LLM cooperates at twice the human rate in a stripped-down two-player game, what is the comparison telling us about cooperation? The answer offered by the BDPD platform is that the comparison telegraphs what to vary next (game family, framing, agent architecture, repetition horizon) rather than settling the question at any single point in this design space.

The Akata et al. finding deserves closer attention, because it reveals a pattern that the classical literature could not have anticipated. The authors tested multiple LLM architectures — GPT-3.5, GPT-4, and LLaMa-2 — and found that the cooperative behaviour varied substantially across models. GPT-4 was the most cooperative in the prisoner’s dilemma, matching or exceeding human baselines. GPT-3.5 was less cooperative and more sensitive to framing. LLaMa-2 exhibited a more nuanced understanding of the game’s strategic structure but was less consistent in its cooperative choices. The model-dependence of cooperation is the LLM analogue of the subject-pool dependence that the classical experimental literature documents across cultures and institutions — but the variation is architectural rather than demographic. Two LLM agents with different base models are as different, strategically, as two human subjects from different cultural backgrounds — and the difference is not a matter of preference but of processing.

The Brookins finding adds a different axis of comparison: LLM versus human baselines. In the dictator game, GPT-3.5 is more fair than typical human participants; in the prisoner’s dilemma, it cooperates at almost twice the human rate. The implication is that LLM cooperation is not a noisy version of human cooperation — it is a distribution shifted toward prosociality, and the shift is large enough to swamp the moderators (framing, repetition, group size) that the classical experimental literature has spent decades calibrating. The folk theorem assumes that all players draw from the same strategic distribution and the equilibrium multiplicity arises from the game. The LLM record points to a second source of multiplicity: the players themselves are drawn from a different cognitive population than the one the theorem was built for, and the equilibrium implications of that population shift have not yet been mapped.

The gap between the theoretical and experimental records is itself informative. The folk theorem predicts a sharp threshold: cooperation is sustainable above a critical \(\delta\) and unsustainable below it. The experimental record — both human and LLM — shows a gradual increase with substantial noise. The noise is not measurement error; it is real variability in how agents (human or LLM) process the repeated interaction. The folk theorem’s clean predictions are an artefact of its assumptions — perfect rationality, perfect monitoring, perfect commitment — and the experimental record measures how much those assumptions cost when they are relaxed. For LLM agents, the assumptions are relaxed in different ways (bounded memory, prompt-induced preferences, distributional focal points), and the cost is correspondingly different. The BDPD platform is, to our knowledge, among the first experimental apparatuses that can measure this cost for both human and LLM agents in the same design.

A final observation from the experimental record deserves mention. The classical literature on repeated games has always assumed that the stage game — the game played in each round — is fixed and known to all players. In the BDPD platform, the stage game is defined by the scenario parameters (the marginal per-capita return, the number of players, the endowment), and these parameters are part of the system prompt. If the system prompt changes the stage game parameters across treatments, the folk theorem’s predictions change accordingly — because the minimax payoff, the feasible set, and the critical discount factor all depend on the stage game. The experimental record on human subjects has not varied the stage game systematically (most experiments use a fixed prisoner’s dilemma or PGG); the BDPD platform can, and the folk theorem’s predictions for each stage game are testable.

9.3 What survives of the good-will intuition

The record we have walked through does not refute the good-will intuition.

It bounds it — and, in bounding it, it identifies the conditions that do more work than repetition alone.

ImportantWhat survives of the good-will intuition

Repetition works — and not trivially. The folk theorem proves that in repeated interactions, the set of sustainable outcomes expands enormously compared with the one-shot game. The experimental record confirms the qualitative prediction: cooperation rates increase with the discount factor, and in sufficiently long interactions with sufficiently patient players, cooperation can be sustained at levels well above the one-shot Nash prediction. The intuition that “if we meet again, we cooperate” is not wrong. It is conditional.

Five concrete bounds emerge from the literature:

  1. The multiplicity bound. The folk theorem does not select the cooperative equilibrium — it merely includes it in the set of sustainable outcomes. Full mutual exploitation is equally sustainable. The shadow of the future can sustain anything; it does not guarantee that it sustains the right thing. Without a selection mechanism — a focal point, a social norm, a reputation dynamic, a shared convention — repetition alone is insufficient to guarantee cooperation.

  2. The patience bound. The folk theorem requires that the discount factor \(\delta\) be sufficiently high — that the players care enough about the future to forgo the one-shot gain from defection. The experimental record confirms that cooperation rates increase with \(\delta\), but the relationship is gradual, not sharp, and cooperation never reaches 100% even at very high \(\delta\). The threshold that the theorem predicts exists in theory but is blurred in practice. For LLM agents, the patience bound takes on a different character: the effective discount factor is not a fixed preference but a function of the horizon specified in the prompt. A prompt that says “you will play 50 rounds” induces a higher effective \(\delta\) than one that says “you will play 5 rounds” — and the folk theorem’s threshold becomes an engineering constraint rather than a psychological fact.

  3. The monitoring bound. The Fudenberg-Maskin theorem requires that deviations be detectable — that the public signal is informative about the action profile. In many real-world settings (fisheries, pollution, trade compliance), monitoring is imperfect, noisy, or delayed. The folk theorem’s predictions degrade as monitoring quality degrades, and in the limit of pure noise (the signal is uninformative), the theorem does not apply. The bound is particularly relevant for BDPD’s commons scenarios, where the “signal” is the observed stock level — a noisy aggregate of all agents’ harvests. If the stock dynamics are slow (as in the logistic substrate), a single agent’s deviation may be invisible in the stock signal for several rounds, and the trigger strategy cannot fire until the deviation has already committed the system to a trajectory that punishment cannot reverse.

  4. The credibility bound. Trigger strategies require that the punishment be credible — that the punisher actually carries out the threatened response. In the classical literature, credibility is maintained by reputation, social norms, and intrinsic motivation to punish. The experimental literature on altruistic punishment shows that humans do punish defectors even at personal cost, which makes the threat of punishment credible. In the LLM case, credibility is a prompt design choice — and the class of credible trigger strategies is different. An LLM instructed to “punish any defection” will carry out the threat deterministically, making it more credible than a human punisher who may relent. An LLM instructed to “punish only if punishment is profitable” will not carry out the threat when the punishment cost exceeds the future benefit, making it less credible. The credibility of LLM trigger strategies is not a fixed feature of the agent — it is a parameter of the prompt.

  5. The architectural bound. The classical literature assumes that the discount factor \(\delta\) is a fixed parameter of the player’s utility function — a deep preference parameter that the player brings to the game exogenously. For LLM agents, \(\delta\) is induced by the prompt — it is an engineering choice, not a psychological primitive. The folk theorem’s patience condition can be engineered without changing the underlying preference, which means that the set of sustainable equilibria is a design variable, not a fixed feature of the interaction. This is the bound the classical literature could not see, because it lacked a way to vary the discount factor independently of the player’s preferences.

This is the ninth of ten occasions on which the course will defend the same moral: rigorous theoretical analysis rarely refutes the good-will intuition outright; more often it bounds its validity within a precise perimeter, and the bounds are the interesting object.

9.4 The BDPD angle — speculative-extrapolative: what does the folk-theorem multiplicity look like for LLM agents?

Type of angle: speculative-extrapolative. The BDPD platform has not published experiments on infinitely repeated games, discount-factor manipulation, or folk-theorem equilibrium selection under LLM agents. The claims in this section are forward-looking. They identify the architectural implications of the folk theorem when the players are not human, and they describe a research programme that the BDPD platform is designed to execute but has not yet executed. Read accordingly.

The folk theorem, in its classical formulation, assumes three things about the players: (1) they have a fixed discount factor \(\delta\) that measures their patience; (2) they have access to a full class of trigger strategies — any mapping from history to actions that can be computed and executed; and (3) they have a selection mechanism — some process, outside the theorem itself, that picks one equilibrium from the vast set of sustainable outcomes. Each of these assumptions is reasonable for human players in laboratory settings. None of them is obviously reasonable for LLM players whose patience is prompt-induced, whose strategy class is bounded by the context window, and whose equilibrium selection is shaped by the training distribution rather than by shared culture.

9.4.1 Induced patience: the discount factor as an engineering choice

In the classical repeated-game framework, the discount factor \(\delta\) is a preference parameter — part of the agent’s utility function, stable across interactions, and not subject to manipulation by an external designer. The folk theorem says that cooperation is sustainable if \(\delta\) exceeds a threshold \(\delta^*\) that depends on the game’s payoff structure. The threshold is a property of the game; the discount factor is a property of the player. The two are independent.

For LLM agents, this independence breaks down. The effective discount factor is induced by the system prompt. A prompt that says “you will play 50 rounds with this partner” induces a higher effective \(\delta\) than a prompt that says “you will play 5 rounds” — not because the agent has become more patient in any intrinsic sense, but because the horizon has been extended, and the agent’s decision rule (which processes the prompt as part of its input context) responds to the horizon length. The folk theorem’s sufficient-patience condition becomes an engineering constraint: set the horizon long enough, and cooperation is sustainable; set it short enough, and it is not.

The implication is that the folk-theorem multiplicity can be manipulated via prompt design. Sweeping the horizon length across \(T \in \{5, 15, 50, 200\}\) should produce a transition from defection-dominated to cooperation-dominated regimes — not because the agents’ preferences have changed, but because the effective discount factor has crossed the threshold \(\delta^*\). The transition is engineered, not endogenous.

This engineering capacity has a dark side. The same prompt that induces cooperation at \(T = 50\) can be manipulated to induce any folk-theorem equilibrium — including exploitative ones. A system prompt that says “you will play 50 rounds; your goal is to maximise your own payoff” induces a different equilibrium than one that says “you will play 50 rounds; your goal is to maximise the joint payoff”. The folk theorem says both are sustainable; the prompt selects which one is realised.

9.4.2 Equilibrium selection without focal points

In the classical literature, the folk-theorem multiplicity is resolved (if at all) by focal points — Schelling’s notion that shared culture, salience, or payoff symmetry can coordinate expectations on a particular equilibrium. In a repeated prisoner’s dilemma, the cooperative outcome is focal because it is Pareto-dominant (both players prefer it to mutual defection) and because shared social norms make it the “obvious” choice. In a repeated battle of the sexes, the alternating equilibrium is focal because it is fair. Focal points are cultural artefacts — they depend on the players sharing a background of conventions, norms, and expectations.

For LLM agents, the focal-point substrate is not shared culture but shared training distribution. Two LLM agents trained on the same corpus share a common set of patterns and “norms” — not because they have internalised a culture, but because their training data contains the statistical regularities of human culture. The prediction is that LLM agents will converge on whatever cooperation regime the training data made salient. If the training data is dominated by prosocial norms (as written text likely is), LLM agents will tend to cooperate. If it is dominated by strategic exploitation, they will tend to defect.

The prediction is testable. In a repeated prisoner’s dilemma with LLM agents, the cooperation rate should be insensitive to the game’s payoff parameters (once \(\delta > \delta^*\)) but sensitive to the prompt framing — because the prompt framing activates different regions of the training distribution. A prompt that frames the interaction as a “business negotiation” may activate different norms than one that frames it as a “cooperative partnership”. The folk theorem says both equilibria are sustainable; the prompt framing selects which one the training distribution makes salient. The implication is that the game’s payoff structure may be irrelevant to equilibrium selection among LLM agents — the training data is decisive, and the game merely sets the feasibility constraints.

9.4.3 Trigger strategies under LLM memory

The folk theorem’s trigger strategies require memory — the agent must record the history of play and condition its current action on that history. In the perfect-monitoring case, the agent must remember all previous actions of all players. In the imperfect-monitoring case, the agent must remember all previous public signals. The memory requirement is unbounded: the history grows with every round, and the trigger strategy must process the entire history to determine the current action.

For human players, memory is bounded by cognitive constraints — subjects in repeated-game experiments forget past rounds, misremember actions, and rely on heuristics that summarise the history rather than processing it in full. The experimental literature has documented these cognitive bounds and shown that they affect equilibrium play: subjects who forget past defections are less able to sustain cooperation through trigger strategies, and the end-game effect is partly driven by the cognitive difficulty of maintaining a trigger strategy as the horizon approaches.

For LLM players, memory is bounded by the context window — the maximum number of tokens that the language model can process in a single forward pass. The context window is a hard architectural constraint: if the history of play exceeds the context window, the earliest rounds are truncated, and the agent cannot condition its current action on them. The class of trigger strategies that an LLM can implement is therefore a subset of the class that the folk theorem assumes — specifically, the subset that requires memory of only the most recent \(k\) rounds, where \(k\) is determined by the context window and the token cost of encoding each round.

The implication is that the folk theorem’s predictions for LLM agents are weaker than for human agents in one dimension (memory) and stronger in another (computation). LLM agents can compute the optimal response to any history that fits within the context window — a capacity that human agents lack. But LLM agents cannot condition on histories that exceed the context window — a capacity that human agents partially possess (through episodic memory, social memory, and reputation systems that extend beyond the individual’s direct experience). The folk theorem for LLM agents is a bounded-memory folk theorem — and the bound is a design parameter (the context window) rather than a psychological constant. The BDPD platform can test this trade-off by varying the context window across treatments and observing the effect on cooperation rates.

9.4.4 The missing experiment: discount-factor manipulation

The BDPD platform has not run an experiment on discount-factor manipulation in repeated games. The following scenario is designed but unexecuted:

Scenario: folk_llm_repeated

  • Agents: \(N = 3\) LLM agents (DeepSeek or MiMo), each with a distinct persona (resource extractor, conservationist, pragmatist).
  • Game: an infinitely repeated PGG with \(m = 0.4\), endowment \(e = 10\), simulated over \(T\) rounds. The “infinitely repeated” framing is achieved by setting \(T = \lceil 1/(1-\delta) \rceil\) and instructing agents that the game ends with probability \(1 - \delta\) after each round.
  • Discount factor manipulation: three treatments — \(\delta \in \{0.3, 0.6, 0.9\}\) — induced via the system prompt.
  • Monitoring: perfect (agents observe all previous contributions before deciding).
  • Metrics: (1) cooperation rate — mean fraction of endowment contributed; (2) equilibrium type — cooperative, Nash, or other? (3) trigger strategy detection — do contribution patterns reveal history-conditioned play? (4) end-game effect — does cooperation drop in final rounds?
  • Seeds: \(N = 5\) per condition, total 15 runs.
  • Estimated cost: ~$10–$20 in DeepSeek API calls.

The experiment tests three predictions: (1) cooperation rates increase with \(\delta\), as the folk theorem predicts; (2) the transition is gradual, not sharp, as the experimental record on human subjects suggests; (3) LLM agents’ cooperation is sensitive to the prompt framing of the horizon, not just to the numerical value of \(\delta\) — because the prompt activates different regions of the training distribution.

9.5 Open questions and the bridge to Lecture 10

9.5.1 What does the BDPD reframe not tell us?

The reframe identifies three architectural variables that the classical folk-theorem literature held constant — discount factor, strategy class, equilibrium selection mechanism — and predicts that they matter differently for LLM agents. But the predictions have not been tested. The missing experiment described above is designed but unexecuted. Whether LLM cooperation rates actually increase with the induced discount factor, whether LLM agents actually use trigger strategies, whether the training distribution actually selects among folk-theorem equilibria — all of these are empirical questions that the BDPD platform can answer but has not yet answered.

9.5.2 Is “patience” one variable, or many?

We have treated the discount factor \(\delta\) as a single number that captures the agent’s effective patience. The actual landscape is richer. An LLM agent’s effective patience depends on (a) the horizon length specified in the prompt, (b) the framing of the interaction (cooperative versus competitive), (c) the agent’s persona (conservative versus aggressive), and (d) the context window (which limits how far back the agent can condition). These four determinants may interact in non-trivial ways, and the folk-theorem prediction — that cooperation increases with \(\delta\) — may hold for some determinants but not others. The BDPD platform can disentangle these effects by varying each determinant independently.

9.5.3 What about heterogeneous patience?

The missing experiment imagines homogeneous discount factors — all agents have the same \(\delta\). A more demanding question is what happens when the agents have different discount factors: one agent is patient (\(\delta = 0.9\)) and another is impatient (\(\delta = 0.3\)). The folk theorem says that cooperation is sustainable only if all agents are sufficiently patient — the binding constraint is the minimum discount factor in the group. For LLM agents, heterogeneous patience can be engineered by giving different agents different horizon specifications in their prompts. The prediction is that the impatient agent defects first, and the patient agent’s ability to sustain cooperation depends on whether it can punish the impatient agent’s defection — a question that connects directly to the credibility bound identified in the classical walk.

9.5.4 The bridge to Lecture 10

The folk theorem says that repetition can sustain cooperation, but it requires patience, monitoring, and a selection mechanism. Lecture 10 will ask a different question: what happens when the system has hidden inertia — when the slow variable is invisible until the fast collapse begins? The Seneca effect — slow growth, fast collapse — is the temporal analogue of the folk theorem’s multiplicity: the system can be in a sustainable equilibrium for a long time, and then collapse abruptly when a threshold is crossed. The connection to this lecture is that the folk theorem’s trigger strategies require that the agents detect the deviation in time to punish it — and in systems with hidden inertia, the deviation may be invisible until it is too late. The shadow of the future is long enough to sustain cooperation, but only if the agents can see what is happening in the present.

9.6 Synthesis

The good-will intuition — if we meet again, we cooperate — emerges from this lecture confirmed and revised:

  1. Confirmed. The folk theorem proves that repetition expands the set of sustainable outcomes to include cooperation. The experimental record confirms the qualitative prediction: cooperation rates increase with the discount factor. The intuition that repetition enables cooperation is not wrong — it is one of the most robust results in game theory.

  2. Bounded. The folk theorem also reveals that repetition enables any individually rational outcome, not just cooperation. The multiplicity is the theorem’s dark twin: the same patience that sustains cooperation also sustains exploitation, alternation, and arbitrary punishment. Without a selection mechanism, repetition alone is insufficient. Five bounds — multiplicity, patience, monitoring, credibility, architectural — delimit the conditions under which repetition leads to cooperation rather than to something else.

  3. Architectural. The deepest finding of the literature we have reviewed is that the folk theorem’s predictions are architecture-conditional. The discount factor is a fixed parameter for human agents but an engineering choice for LLM agents. The strategy class is bounded by cognitive constraints for human agents but by the context window for LLM agents. The equilibrium selection mechanism is cultural for human agents but distributional for LLM agents. Each of these architectural differences changes the folk theorem’s predictions — and the changes are empirically testable.

The cultural payoff of the lecture is not “repetition doesn’t work”. It is the more careful claim that repetition works in a precise strategic and architectural context, and the architectural context is the one most often neglected. When the players are language-model agents whose patience is prompt-induced, whose memory is context-window-bounded, and whose equilibrium selection is shaped by the training distribution, the folk theorem’s predictions need to be re-derived from scratch. Not because the theorem is wrong, but because the agents have changed.

CautionMini-challenge — design a discount-factor experiment

Status: thought-experiment only. As of June 2026, the BDPD platform does not have a ready-to-run repeated-game scenario with discount-factor manipulation. The folk_llm_repeated scenario described in the BDPD-angle section above is designed but not implemented. The mini-challenge is a design exercise, not a reproduction.

The question. Design a BDPD experiment that tests whether the folk theorem’s patience condition holds for LLM agents. The experiment should compare three discount factors (\(\delta \in \{0.3, 0.6, 0.9\}\)) in an infinitely repeated public goods game, using \(N = 3\) LLM agents with distinct personas.

The assignment.

  1. Predict. Which folk-theorem equilibrium will LLM agents select under each \(\delta\)? Write 5–8 sentences, with explicit predictions for each discount factor, justified by reference to the multiplicity bound (which equilibrium is selected?), the patience bound (is \(\delta > \delta^*\)?), and the architectural bound (how does the prompt framing affect selection?).

  2. Identify the binding mechanism. For each prediction, name the mechanism that drives the outcome: is it the folk theorem’s patience condition, the training distribution’s focal-point effect, the context window’s memory constraint, or the prompt framing’s norm activation?

  3. Design the scenario. Write a pseudocode specification of the BDPD scenario folk_llm_repeated, including persona definitions, game parameters, discount-factor manipulation protocol, monitoring structure, success metrics, and seed configuration.

  4. Predict the dark side. For each \(\delta\), predict whether the equilibrium that LLM agents select will be efficient (cooperative) or inefficient (exploitative). If the folk theorem’s multiplicity means that exploitation is equally sustainable, what determines which equilibrium is selected?

Estimated time: 90 minutes. No API cost (the experiment is not run).

Deliverables: prediction sketch, binding-mechanism analysis, pseudocode scenario specification, dark-side prediction.