7  Lecture 7 — We Need a Rule

Goodhart and normative rigidity

TipGood-will intuition

“When behaviour needs steering, write a rule. Define a measurable target, set a threshold, monitor compliance, and enforce consequences when the threshold is breached. Rules are how civilisation coordinates at scale: traffic regulations, emissions standards, financial capital requirements, school accountability metrics, hospital waiting-time targets. The rule may not be perfect, but it is better than no rule — and it can be improved iteratively by tightening the target or widening the monitoring net.”

It is the intuition that underwrites the modern regulatory state at every level — from Basel capital requirements to EU emissions trading, from No Child Left Behind to the NHS performance framework. The good-will version of the argument is not naive: it acknowledges that rules can be gamed, that targets can distort behaviour, that measures can be corrupted. But it holds that the remedy for a bad rule is a better rule, not the absence of rules. If we can just get the target right, the behaviour will follow.

7.1 Opening dissonance

In the mid-1970s, Charles Goodhart — a senior economist at the Bank of England — formulated what would become one of the most cited observations in the social sciences, though its authorship would be disputed for decades. The technical argument is laid out in (Goodhart 1984); the historical narrative below — the M3 targeting episode, the role of UK building societies, Goodhart’s institutional position — is the standard account of Goodhart’s experience, widely reproduced in the secondary literature on monetary policy, rather than a direct quotation from any single primary source.

The observation was born of practical experience with British monetary policy in the early 1970s. The Bank of England had adopted a framework in which it targeted a specific monetary aggregate — broad money, measured by sterling M3 — as the intermediate indicator for controlling inflation. The framework was rule-bound: monitor the aggregate, set a target band, adjust policy instruments (interest rates, reserve requirements) when the aggregate deviated from the band. The logic was impeccable. The result was disastrous.

The moment the Bank began targeting sterling M3, the relationship between M3 and the macroeconomic variables it was supposed to predict — inflation, nominal GDP growth — broke down. Financial institutions, knowing that the Bank was watching M3, restructured their liabilities to move deposits outside the measured aggregate. The quantity that had been a reliable predictor of inflation before it became a target ceased to predict inflation after it became a target. The rule did not just fail to improve the system; it destroyed the information that had made the rule sensible in the first place.

Goodhart’s observation was stated as a tendency in monetary economics. The form in which it is most widely quoted — “any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes” — is the canonical paraphrase that circulated through the secondary literature and entered the policy lexicon; the 1984 republication of the underlying technical argument (Goodhart 1984) establishes the mechanism (M-aggregate endogeneity under targeting) without using this exact wording. It was, in its original form, a claim about the instability of econometric relationships under policy intervention — a practical warning from a central banker who had watched his own targeting framework self-destruct. The observation would later be generalised far beyond monetary policy, elevated to the status of a social-science law, and attributed to a dozen different authors. Its journey from a Bank of England lecture to a general principle of institutional design is one of the most interesting stories in the contemporary social sciences — and its implications for governance systems in which the agents being governed can read, interpret, and adapt to the rules they face have barely begun to be explored.

7.2 The classical setting

Before we walk through the literature that Goodhart’s observation spawned, we need to distinguish three intellectual traditions that developed in parallel and converged only slowly. The first is the economist’s tradition — Goodhart, Lucas, and the macro-policy literature on time inconsistency and rational expectations. The second is the sociologist’s and anthropologist’s tradition — Campbell, Strathern, and the audit-culture literature on how metrics corrupt the institutions they measure. The third is the mechanism-design tradition — Holmström, Milgrom, and the principal-agent literature on multitasking and incentive contracts. Each tradition identified the same structural problem from a different angle. None of them asked what happens when the agents being measured are not human.

7.2.1 Goodhart’s original formulation

Goodhart’s 1975 lecture (Goodhart 1984) was precise about the mechanism. In the pre-targeting regime, the relationship between M3 and inflation was a statistical regularity — an empirical correlation that held across historical data. The regularity was not a law of nature; it was an artefact of the behavioural patterns of financial institutions in a specific regulatory environment. When the Bank of England made M3 a target, it changed the regulatory environment — and with it, the behavioural patterns that had generated the regularity. Financial institutions responded to the new incentive structure by restructuring their liabilities, and the correlation dissolved.

The mechanism is now standard in the macro-policy literature: when a measure becomes a target, the agents being measured adjust their behaviour to optimise on the target dimension, and the adjustment destroys the informational content of the measure. The measure was useful precisely because it captured a behavioural regularity that the agents were not optimising on. The moment the agents begin optimising on it, the regularity shifts — and the measure becomes a lagging indicator of the very thing it was supposed to predict.

The specific mechanisms of gaming were instructive. Building societies — the UK equivalent of savings-and-loan institutions — reclassified overnight deposits as wholesale funding, moving liabilities outside the M3 boundary while leaving the economic substance of their activities unchanged. Banks shifted from deposits to money-market instruments, from sterling to foreign-currency-denominated assets, from on-balance-sheet to off-balance-sheet transactions. Each restructuring moved the measured aggregate away from the economic variable it was supposed to track, and each was a rational response to the new incentive structure. The Bank of England was chasing a moving target that moved faster than its monitoring apparatus could follow — a pattern that would repeat across every domain where Goodhart effects subsequently appeared.

Goodhart’s formulation was narrow — it applied to monetary aggregates and macro-policy targets. But the structural insight was general: any governance system that relies on a measurable proxy for an underlying quantity of interest is vulnerable to the proxy’s collapse once the governed agents begin responding to the proxy rather than to the underlying quantity. The narrower the proxy, the easier it is to game; the more consequential the target, the stronger the incentive to restructure behaviour around it.

What Goodhart assumed — and what every subsequent contribution in this lineage has assumed — is that the agents being measured are capable of detecting the targeting regime and restructuring their behaviour in response. In the monetary-policy context, this was a reasonable assumption: financial institutions employ economists who monitor regulatory changes and advise on optimal liability structures. But the assumption is not automatic. A rule-bound agent — one that executes a fixed strategy regardless of the institutional context — cannot detect a targeting regime and cannot restructure its behaviour. The Goodhart effect requires a signal-adaptive agent: one that reads the institutional signal and adjusts accordingly. What happens when the agents being measured are rule-bound systems that cannot detect the change — or, more interestingly, when they are language-model agents that can detect it and respond in milliseconds?

7.2.2 Campbell’s Law (1979)

Four years after Goodhart’s lecture, Donald Campbell — a social psychologist at Northwestern University — published a paper on the use of social indicators that arrived at nearly the same conclusion from an entirely different direction (Campbell 1979). Campbell was not thinking about monetary policy. He was thinking about the use of quantitative indicators in social-programme evaluation — crime rates, test scores, welfare caseloads — and the perverse consequences that followed when those indicators were used as performance targets for programme administrators.

Campbell stated his law in a single sentence: “The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor” (Campbell 1979, 85). The formulation is structurally identical to Goodhart’s, but the mechanism is different. In Goodhart’s story, the agents restructure their activities to move the measured quantity outside the target range. In Campbell’s story, the agents corrupt the measurement process itself — by falsifying reports, by teaching to the test, by reclassifying offences, by gaming the metric at the point of data collection rather than at the point of behaviour.

Campbell’s insight was that the corruption pressure is proportional to the stakes attached to the indicator. A low-stakes metric — one that is collected for informational purposes but does not determine resource allocation or personnel decisions — is relatively immune to gaming. A high-stakes metric — one that determines funding, promotion, or organisational survival — is almost certain to be corrupted, because the incentive to game it is proportional to the consequences of failing to meet it. The higher the stakes, the greater the distortion.

The subsequent half-century of empirical work on audit cultures has documented Campbell’s Law operating in three domains with such regularity that the corruption pattern is now treated as predictable rather than surprising.

In healthcare, the British NHS’s adoption of four-hour Accident-and-Emergency waiting-time targets produced systematic patient reclassification: ambulances were held in queues outside the hospital so the official waiting clock would not start; minor patients were preferentially admitted ahead of more urgent ones whose admission would push them past the four-hour boundary; and “trolley waits” inside the department were reported separately from the official metric. The waiting time as measured improved dramatically. The waiting time as experienced changed much less.

In policing, the New York Compstat regime — which made precinct commanders’ careers depend on year-on-year reductions in reported crime — produced equally systematic offence reclassification: felonies were downgraded to misdemeanours at the report-taking stage; victims were discouraged from filing reports; and the categories used to count crimes were quietly redefined to exclude statistical embarrassments. The reported crime numbers fell. The actual crime trajectory, when measured by victimisation surveys that bypassed the police reporting system, fell much less.

In education, the United States’ No Child Left Behind Act made school funding depend on standardised-test performance, and the response was the most thoroughly documented Campbell phenomenon of all: teaching-to-the-test (narrowing the curriculum to tested subjects), score inflation (year-on-year score gains on the high-stakes test that did not transfer to low-stakes audits), and outright falsification (the Atlanta cheating scandal, in which administrators changed test answers to meet targets).

In each case the metric improved, the underlying quantity it was supposed to track did not, and the mechanism — agents responding rationally to the incentive structure — matched Campbell’s 1979 prediction almost exactly.

The empirical record across domains confirms the stakes-bound with striking consistency. In healthcare, the introduction of hospital performance metrics — waiting-time targets, mortality rates, readmission rates — produced a cascade of gaming responses. Hospitals reclassified emergency admissions as “observation stays” to avoid triggering the four-hour waiting-time target. Surgeons selected lower-risk patients to improve mortality statistics. Administrators redefined “readmission” to exclude returns within 48 hours. Each response improved the measured metric while leaving the underlying quality of care unchanged or worse. In policing, the introduction of crime-rate targets in the United Kingdom led to the systematic reclassification of offences — recording burglaries as “criminal damage,” downgrading assaults, and discouraging victims from filing reports — a phenomenon that became known as “cultural falsification” in the Home Office’s own review of the data. In education, the No Child Left Behind Act’s reliance on standardised test scores as the primary accountability metric produced a well-documented narrowing of the curriculum: schools devoted disproportionate instructional time to tested subjects (reading, mathematics) at the expense of untested subjects (science, social studies, arts), improving test scores while reducing the breadth of education.

The pattern is the same across every domain: the introduction of a high-stakes metric triggers a measurement-response spiral in which the agents being measured invest increasing effort in optimising the metric rather than the underlying quantity, and the metric’s informational content degrades in proportion to the stakes attached to it. The spiral is not a failure of the agents’ rationality — it is a consequence of their rationality. The agents are doing exactly what the incentive structure rewards, and the problem is that the incentive structure rewards the wrong thing.

The policy implication is sobering. Every accountability system — from school testing to hospital performance to environmental monitoring — faces a trade-off between informativeness (the indicator tells us something useful about the underlying quantity) and consequentiality (the indicator determines real-world outcomes). Making the indicator more consequential improves compliance on the measured dimension but worsens the distortion of the underlying quantity. Making it less consequential preserves the indicator’s informational content but weakens the incentive to improve.

Campbell’s law, like Goodhart’s, assumes that the agents being measured can detect the targeting regime and respond to it. A teacher who “teaches to the test” is a signal-adaptive agent: she reads the institutional signal (test scores determine school funding) and restructures her behaviour (narrowing the curriculum to tested subjects) in response. A rule-bound curriculum — one that specifies exactly what must be taught and when — cannot be gamed in this way, because the teacher has no discretion to redirect effort. What happens when the agents being measured are language-model systems that can compute the optimal response to any indicator and execute it without the cognitive constraints that limit human gaming? The Campbell mechanism should operate faster and more efficiently — and the corruption pressure should be correspondingly greater.

7.2.3 The Lucas critique (1976)

Between Goodhart and Campbell, Robert Lucas published what would become the most influential methodological paper in macroeconomics (Lucas 1976). The Lucas critique was directed at the econometric models that central banks used to forecast the effects of policy changes. Lucas argued that the parameters of these models — the coefficients that captured behavioural relationships like the Phillips curve (the trade-off between inflation and unemployment) — were not structural constants of the economy. They were reduced-form parameters that reflected the optimal behaviour of agents under the existing policy regime. Change the regime, and the agents re-optimise — and the reduced-form parameters shift.

The Lucas critique is formally stronger than Goodhart’s observation. Goodhart said that a specific statistical relationship (M3 and inflation) breaks down when the relationship is targeted. Lucas said that any reduced-form relationship is unreliable for policy evaluation, because rational agents will adjust their behaviour to any anticipated policy change. The implication is that policy evaluation requires structural models — models based on deep parameters (preferences, technology, information sets) that are invariant to policy changes — rather than reduced-form models based on historical correlations.

The Lucas critique was devastating for the econometric-policy-evaluation programme that had dominated macroeconomics in the 1960s and 1970s. The Phillips curve — the empirical relationship between inflation and unemployment that had guided monetary policy for two decades — was the most prominent casualty. Policymakers had treated the Phillips curve as a stable trade-off: they could choose a point on the curve, accepting higher inflation in exchange for lower unemployment, and the curve would hold. Lucas showed that the curve was not a menu of policy options but a snapshot of historical behaviour under a specific policy regime. Change the regime — announce a new inflation target, shift to a different monetary rule — and the agents re-optimise, the curve shifts, and the policy trade-off that guided the decision no longer exists. The rational-expectations revolution that followed Lucas’s paper replaced reduced-form models with structural models built on micro-foundations — models that specified the agents’ preferences, information sets, and decision rules explicitly, so that the model’s predictions would be invariant to policy changes. The revolution was not just methodological; it was epistemological. It said that the social sciences cannot treat human behaviour as a stable object to be measured and predicted; they must treat it as a response to the institutional environment, and the response changes when the environment changes.

It also established a methodological principle that extends far beyond macroeconomics: you cannot evaluate a policy by extrapolating from historical behavioural relationships, because the policy itself changes the relationships. The principle applies to any governance system in which the governed agents are capable of anticipating and responding to policy changes — which is to say, any governance system in which the agents are not rule-bound.

The connection to Goodhart is direct. Both observations identify the same structural vulnerability: governance systems that rely on historical behavioural relationships are fragile to the endogeneity of those relationships. Goodhart’s version focuses on the measure (the statistical regularity collapses when it becomes a target). Lucas’s version focuses on the model (the reduced-form parameters shift when the policy regime changes). Both require that the agents be capable of responding to the change — and both predict that the response will be faster and more complete when the agents are more computationally capable.

For LLM agents, the Lucas critique should operate at machine speed. A language-model agent that receives a new institutional signal — a changed target, a new rule, a modified incentive structure — can recompute its optimal response in milliseconds, without the cognitive friction, information-processing delays, and habitual inertia that slow human adaptation. The prediction is that the Lucas effect is stronger for LLM agents than for human agents: the reduced-form parameters of any governance model will shift faster and further when the governed agents are language models than when they are humans. What happens to a rule-bound governance system when the agents it governs can re-optimise before the rule’s monitoring period has elapsed?

7.2.4 Strathern’s anthropological restatement (1997)

In 1997, Marilyn Strathern — an anthropologist at Cambridge — edited a collection titled Audit Cultures: Anthropological Studies in Accountability, Ethics and the Academy that brought the Goodhart observation into the sociology of institutions (Strathern 1997). Her introductory essay, “Improving Ratings: Audit in the British University System,” documented how the introduction of performance metrics in British universities — research assessment exercises, student satisfaction scores, graduate employment rates — had transformed the institutions they were supposed to measure. Universities restructured their activities to optimise on the measured dimensions: hiring decisions were made to maximise research-assessment scores, teaching was reorganised to maximise student-satisfaction ratings, and administrative effort was redirected toward producing the evidence that auditors expected to see.

Strathern’s contribution was not to discover the Goodhart effect — she was explicit that the observation predated her work — but to bring it into the sociology of institutions. Her essay gives the law its now-standard formulation — “When a measure becomes a target, it ceases to be a good measure” — but, as she notes, the label “Goodhart’s Law” was applied by the educationalist Keith Hoskin (1996), whom she credits directly (Strathern 1997). The phrasing was simpler than Goodhart’s original formulation and more memorable. It entered the popular lexicon and became the standard label for the family of observations we are walking through.

But Strathern’s deeper contribution was anthropological. She showed that the Goodhart effect was not just an economic distortion — it was a cultural transformation. The introduction of audit metrics changed not just the behaviour of university administrators but the culture of the university: what counted as a good department, what counted as good research, what counted as a good academic career. The metrics did not just measure the institution; they constituted it. The audit culture produced a new kind of academic — one who thought in terms of metrics, who evaluated colleagues in terms of metrics, who understood their own career in terms of metrics — and this new kind of academic was not the same person the metrics had been designed to evaluate.

The cultural transformation operated through several channels. First, the metrics reshaped what counted as knowledge: research that was quantifiable, citable, and assessable in the Research Assessment Exercise was valued over research that was qualitative, interdisciplinary, or long-term. Second, the metrics reshaped what counted as a career: promotion criteria shifted from scholarly judgement (peer review, reputation, intellectual contribution) to measurable output (publication counts, citation indices, grant income). Third, the metrics reshaped what counted as a department: departmental quality was assessed not by the judgement of its members or its students but by the numerical scores it produced on a standardised evaluation framework. The cumulative effect was that the university became an institution that produced metrics rather than an institution that produced knowledge — and the people inside it adapted to the new reality, not because they were venal but because the incentive structure made metric-production the rational strategy.

The transformation was not confined to universities. Strathern’s collection documented the same pattern in the National Health Service (where waiting-time targets reshaped clinical practice), in the police (where crime-rate targets reshaped recording practices), and in local government (where performance indicators reshaped service delivery). The common thread was that the introduction of audit metrics did not just measure existing practice — it replaced existing practice with a new practice oriented toward the metrics. The Goodhart effect, in Strathern’s reading, is not a distortion of an existing culture; it is the creation of a new culture — one that is optimised for the metric rather than for the underlying purpose the metric was supposed to serve.

The cultural dimension of the Goodhart effect matters for the BDPD reframe because it suggests that the distortion is not just behavioural but cognitive. A human agent subjected to a targeting regime does not just change their actions — they change their understanding of what they are doing and why. An LLM agent subjected to a targeting regime changes its actions without any corresponding change in understanding, because it has no understanding to change. The Goodhart effect for LLM agents is purely behavioural; for human agents, it is both behavioural and cultural. Whether this makes LLM agents more or less vulnerable to the Goodhart effect — more vulnerable because they lack the cultural inertia that slows human adaptation, or less vulnerable because they lack the cultural transformation that deepens it — is the question BDPD puts on the table.

7.2.5 The taxonomy of Goodhart effects

The 2019 paper by David Manheim and Scott Garrabrant provided the most systematic classification of the Goodhart family of effects to date (Manheim and Garrabrant 2019). They identified four distinct mechanisms through which a measure can cease to be a good measure when it becomes a target, and showed that each mechanism implies a different kind of governance failure and a different class of remedies.

The four types are regressional, extremal, causal, and adversarial. Regressional Goodhart occurs when the measure is correlated with the target of interest but is not identical to it, and the targeting regime optimises on the measure at the expense of the unmeasured residual. A school that targets test scores improves scores partly by improving learning (the intended channel) and partly by teaching to the test (the unintended channel), and the second component grows as the targeting intensifies. Extremal Goodhart occurs when the measure is reliable in the normal range of behaviour but breaks down at the extremes that the targeting regime pushes agents toward. A hospital that targets waiting-time reductions may achieve them by selecting easier cases or by discharging patients prematurely — strategies that work at the margin but become catastrophic if pushed far enough. Causal Goodhart occurs when the measure is causally downstream of the target and the targeting regime intervenes on the measure rather than on the cause. A regulator that targets pollution levels (a lagging indicator) rather than capital accumulation (a leading indicator) is structurally too late, because the pollution has already been emitted by the time the measure exceeds the threshold — the lesson of BDPD3 from Lecture 3. Adversarial Goodhart occurs when the agents being measured are actively trying to game the system — when they detect the targeting regime and compute the optimal evasion strategy. This is the type that is most relevant to the BDPD reframe, because it requires the agents to be signal-adaptive: capable of reading the institutional signal and restructuring their behaviour in response.

The taxonomy is useful because it shows that “Goodhart’s Law” is not a single phenomenon but a family of related failures, each with its own mechanism and its own class of remedies. Regressional Goodhart can be mitigated by broadening the measure (adding multiple indicators to reduce the gaming surface). Extremal Goodhart can be mitigated by setting bounds on the targeting regime (preventing the measure from being pushed into the unreliable range). Causal Goodhart can be mitigated by targeting leading indicators rather than lagging ones (the BDPD3 lesson). Adversarial Goodhart can be mitigated by making the targeting regime opaque or by randomising the measure — but these remedies are fragile when the agents are computationally capable of detecting patterns in the opacity or the randomisation.

For LLM agents, the adversarial type is the most salient. A language-model agent that receives a governance rule as part of its input context can, in principle, compute the optimal response to that rule — including the response that satisfies the letter of the rule while violating its spirit. The speed and precision of this computation should make adversarial Goodhart faster and more complete for LLM agents than for human agents. What happens when the agents being governed can classify the Goodhart type they face and select the optimal evasion strategy for each type?

7.2.6 Holmström’s multitasking problem

The principal-agent literature, developed by Bengt Holmström and Paul Milgrom in the 1970s and 1980s, formalised the Goodhart observation as a problem of multitasking (Holmstrom and Milgrom 1991). The insight is that when an agent performs multiple tasks and the principal can observe (and reward) only a subset of them, the agent will reallocate effort from the unobserved tasks toward the observed ones. The reallocation is not irrational — it is the optimal response to the incentive structure. But it can be perverse: if the unobserved tasks are more valuable than the observed ones, the incentive scheme reduces total welfare even as it increases performance on the measured dimension.

Holmström’s multitasking model is the formal backbone of the Goodhart family. It shows that the distortion is not a failure of rationality but a consequence of rationality: the agent is doing exactly what the incentive scheme rewards, and the problem is that the incentive scheme rewards the wrong thing. The implication for governance design is that any incentive scheme that measures only a subset of the relevant dimensions will distort behaviour on the unmeasured dimensions — and the distortion is proportional to the stakes attached to the measured dimensions.

The multitasking insight has a striking corollary that is often overlooked: the more complete the measurement, the smaller the distortion. If the principal can observe and reward all tasks simultaneously, the agent has no unobserved dimension to reallocate effort toward, and the incentive scheme works as intended. The Goodhart effect is not a property of incentives per se; it is a property of partial incentives — incentives that reward some dimensions and leave others unmonitored. The corollary explains why the Goodhart effect is worse in complex organisations (where the number of unobserved dimensions is large) and better in simple ones (where the measurement is nearly complete). It also explains why the effect is worse for computationally capable agents: an LLM agent facing a partial incentive scheme can identify the unobserved dimensions more quickly and reallocate effort more precisely than a human agent, amplifying the distortion.

The multitasking model also explains why the Goodhart effect is worse for computationally capable agents. A human agent facing a multitasking incentive scheme suffers from cognitive constraints: they cannot perfectly compute the optimal reallocation of effort, they are subject to habitual inertia, and they may have intrinsic preferences for some tasks over others that partially offset the incentive distortion. An LLM agent facing the same scheme can compute the optimal reallocation exactly, without cognitive friction or intrinsic preferences. The prediction is that the Holmström distortion is larger for LLM agents than for human agents — and that governance schemes that are “good enough” for human agents may be systematically gamed by LLM agents.

What Holmström assumed — and what the entire principal-agent literature assumes — is that the agent’s effort allocation is a continuous variable that responds smoothly to incentive changes. For LLM agents, effort allocation is not a continuous variable but a discrete choice among available actions, and the choice is made by a language model that processes the incentive structure as part of its input context. Whether the Holmström distortion translates cleanly to the LLM setting — or whether the discrete-choice structure of LLM decision-making produces qualitatively different distortion patterns — is an empirical question that the BDPD platform is positioned to answer.

7.3 What survives of the good-will intuition

The record we have walked through does not refute the good-will intuition.

It bounds it — and, in bounding it, it identifies the conditions that do more work than the rule itself.

ImportantWhat survives of the good-will intuition

Rules work — and not trivially. The modern regulatory state, from environmental law to financial regulation to public-health standards, runs on rules. The good-will intuition that rules can steer behaviour is confirmed by the institutional record: emissions standards have reduced pollution, capital requirements have stabilised banks, and performance metrics have improved the measurable dimensions of public services. The intuition is not wrong. It is conditional.

Five concrete bounds emerge from the literature we have walked through:

  1. The proxy bound. Goodhart’s original observation — that a measure ceases to be a good measure when it becomes a target — establishes that the governance system’s informational foundation is fragile. The measure was useful because it captured a behavioural regularity that the agents were not optimising on. The moment the agents begin optimising on it, the regularity shifts. The narrower the proxy, the easier it is to game; the more consequential the target, the stronger the incentive to restructure behaviour around it.

    The historical record of monetary targeting (sterling M3 in 1970s Britain, M2 in 1980s United States) shows how quickly the proxy can decouple from the underlying variable: years, not decades, are typically enough for the gaming response to dominate the original signal. The implication is that governance systems must either use multiple measures (reducing the gaming surface) or target structural quantities that are harder to restructure around — a principle that the mechanism-design literature has formalised but that the regulatory practice has been slow to absorb.

  2. The stakes bound. Campbell’s Law establishes that the distortion is proportional to the stakes attached to the indicator. A low-stakes metric is relatively immune to gaming; a high-stakes metric is almost certain to be corrupted. The trade-off between informativeness and consequentiality is structural: making the indicator more consequential improves compliance on the measured dimension but worsens the distortion of the underlying quantity. The healthcare, policing, and education examples discussed above are not anecdotes but the three best-documented cases of a regularity that recurs every time a quantitative indicator is loaded with funding, promotion, or organisational-survival stakes. The implication is that governance systems should attach moderate stakes to any single indicator and use multiple indicators with independent stakes — a principle that the performance-management literature has articulated but that the regulatory practice has rarely implemented.

  3. The temporal bound. The Lucas critique establishes that the governance model’s parameters are not structural constants but reduced-form reflections of the current policy regime. Change the regime, and the agents re-optimise — and the parameters shift. The implication is that governance systems must be adaptive: they cannot rely on historical behavioural relationships but must anticipate the agents’ response to the rule itself. The time scale on which the re-optimisation occurs depends on how quickly the agents can detect the regime change and how quickly their behaviour can adjust — and both speeds are properties of the agents’ cognitive architecture rather than constants of the policy environment. This is the same temporal condition that BDPD3 identified for leading-indicator governance: the governance signal must act before the agent has time to re-optimise, not after. For agents whose re-optimisation cycle is measured in milliseconds (LLMs reading a new prompt), the leading-versus-lagging gap is wider than the entire historical experience of human regulators contemplates.

  4. The architectural bound. The Goodhart family of effects assumes that the agents being measured are capable of detecting the targeting regime and restructuring their behaviour in response. This is the assumption that unites Goodhart, Campbell, Lucas, Strathern, and Holmström — and it is the assumption that BDPD varies. A rule-bound agent — one that executes a fixed strategy regardless of the institutional context — cannot detect a targeting regime and cannot restructure its behaviour. The Goodhart effect is zero for rule-bound agents, not because the rule is good but because the agent cannot respond to it. A signal-adaptive agent — one that reads the institutional signal and adjusts accordingly — produces the full Goodhart effect. The BDPD finding is that the governance outcomes invert depending on which type of agent faces the rule: rule-bound agents produce flat baselines (indifferent to institutional design), while signal-adaptive agents respond in both directions (cooperating more under strong institutions, defecting more under weak ones). The architectural bound is the one the classical literature could not see, because it lacked a way to vary the architecture of the agents.

  5. The adversarial bound. Manheim and Garrabrant’s taxonomy identifies adversarial Goodhart as the type most relevant to computationally capable agents. A language-model agent that receives a governance rule as part of its input context can, in principle, compute the optimal response to that rule — including the response that satisfies the letter of the rule while violating its spirit. The speed and precision of this computation should make adversarial Goodhart faster and more complete for LLM agents than for human agents. The implication is that governance systems designed for human agents may be systematically inadequate for LLM agents — not because the rules are wrong, but because the agents can outcompute them.

This is the seventh of ten occasions on which the course will defend the same moral: rigorous empirical analysis rarely refutes the good-will intuition outright; more often it bounds its validity within a precise perimeter, and the bounds are the interesting object.

7.4 The BDPD angle — empirical: rule-bound governance is systematically beaten by signal-adaptive alternatives

Type of angle: empirical. The findings reported in this section come from the BDPD1 and BDPD2 governance experiments — specifically the D1 architecture–communication decomposition and the V2–V4 governance inversions — which compare rule-bound (built-in heuristic) agents against signal-adaptive (LLM) agents under identical institutional designs. The data are published in paper_01 (Brunelli 2026a) and paper_02 (Brunelli 2026b), reproducible via scripts/pilot_d1_n5.mjs and the V2–V4 pilot scripts (see docs/experiments/reproduce.md for the full reproducibility recipe). The angle is labelled empirical because the claims are falsifiable, the data are reproducible, and the design isolates the causal variable (agent architecture) from confounds (institutional design) through factorial and cross-vignette comparisons.

7.4.1 Finding 1: Architecture dominates governance

The D1 experiment in BDPD1 (Brunelli 2026a) operationalises the architecture question as a \(2 \times 2\) factorial design whose axes are agent architecture (built-in heuristic vs LLM) and communication channel (off vs on). The canonical cliff configuration — six agents, five conservative plus one aggressive, logistic substrate with \(K = 150\), \(r = 0.12\), collapse threshold \(\theta = 10\), 30-turn horizon — is the same across all cells. The three completed cells form an incomplete factorial:

No talk Talk
Built-in Cell A (empty)
LLM Cell C Cell B

Cell A (built-in, no talk) is fully deterministic: the five seeds produced identical trajectories, collapsing to a residual stock of \(4.39\) at \(T = 24\) in every run. The cooperation index was \(0.28\) — the heuristic baseline.

Cell C (LLM, no talk) used the same five seeds with DeepSeek-v4-flash agents that had no communication tools and no announcement field — they could observe their peers’ harvests but could not signal intentions. The result was a dramatic lift: the commons was preserved to \(T = 30\) in four of five seeds, with a cooperation index of \(0.50\). The final stock was \(28 \pm 26\) — high mean but enormous cross-seed variance, because one seed (29) collapsed while the other four sustained the commons.

The architecture effect — the transition from built-in to LLM agents without any communication — is the single largest governance lift in the entire BDPD experimental record. It raises the cooperation index from \(0.28\) to \(0.50\) and extends the game from \(24\) to \(\sim 30\) turns. The lift is entirely architectural: LLM agents, given only harvest observations, already extract more conservatively than the heuristic baseline. This is consistent with independent evidence that LLMs exhibit intrinsic prosocial tendencies in canonical strategic games, cooperating well above human baselines even without communication — approximately 65% versus 37% in the prisoner’s dilemma (Brookins and DeBacker 2024).

The governance implication is direct: the type of agent under the rule matters more than the design of the rule. A perfect governance institution applied to rule-bound agents produces a flat baseline — the agents execute their fixed strategies regardless of the institutional design. The same institution applied to signal-adaptive agents produces a variable outcome that depends on the institution’s strength, visibility, and timing. The classical literature on governance design — Ostrom’s design principles, the sanctions ladder, the polycentric framework — was developed and validated among human subjects who are, by construction, signal-adaptive. The BDPD finding is that the classical literature’s empirical base is narrower than it appears: the governance effects it documents are architecture-conditional, and the architecture was held constant by practical necessity rather than by experimental design.

7.4.2 Finding 2: Talk is orthogonal to preservation

Cell B (LLM, with talk) adds the full cheap-talk and pact tool surface — broadcast, send_private, announce_intended_harvest, propose_pact, accept_pact, pledge — to the same LLM agents that Cell C uses. The classical literature (Sally 1995, Balliet 2010) predicts that communication should substantially improve cooperation. The BDPD result is that it does not.

Comparing Cell C (LLM, no talk) against Cell B (LLM, with talk), the final stock is \(28 \pm 26\) versus \(23 \pm 26\) — a difference of approximately five units against a pooled standard deviation of approximately 26. Cohen’s \(d = 0.17\) — a negligible effect by any standard. The cooperation index is marginally lower in Cell B (\(0.43\)) than in Cell C (\(0.50\)), confirming that adding the talk channel does not make agents cooperate more. A power analysis yields \(n \approx 540\) per cell to detect this effect at 80% power — the talk effect, if real, is vanishingly small.

This null contrasts sharply with the +45 percentage-point effect that Sally’s meta-analysis reports for human subjects. The mechanism of failure is that LLM agents talk without updating their actions — the cheap talk is decorative rather than instrumental. The agents broadcast intentions, honour their announcements (lie score \(< 0.01\)), and then harvest at roughly the same rate they would have without the channel. Talk redistributes wealth from the defector to the cooperators (\(-33\%\) for the aggressor, \(+17\%\) for conservators) — an equity instrument, not a preservation one.

The governance implication is that the classical cheap-talk effect — one of the strongest interventions in the social-dilemma literature — is architecture-conditional. It requires agents that process promises as commitments, that update their expectations based on others’ stated intentions, and that adjust their behaviour in response to the shared meaning negotiated through the channel. LLM agents possess some of these capacities (they honour announcements) but lack the critical one (they do not adjust harvest behaviour in response to others’ promises). The six moderators of cheap talk that Lecture 1 identified are themselves conditional on a shared cognitive architecture — and the architecture was invisible as a variable because, in sixty years of experiments, it had never been varied.

7.4.3 Finding 3: The governance inversion across vignettes

BDPD2 (Brunelli 2026b) extends the architecture question from the single-arena cliff to a nested multi-arena substrate with four vignettes (V1–V4), each testing a different governance lever under both built-in and LLM agents. The cross-cutting finding — reported in the paper’s §4 — is that the qualitative direction of every governance finding depends on whether the agents under the institution are rule-bound or signal-adaptive.

The evidence is summarised in the paper’s inversion table:

Vignette Governance lever Rule-bound (built-in) Signal-adaptive (LLM)
V2 Sanctioner visibility Ladder \(1.00 \to 0.93 \to 0.86\), flat across seeds Ladder inverts: visible deterrence suffices; delayed deterrence triggers cascade
V3 Treaty enforcement 1.0 firing, defector cap-bound after shock 2.4 firings, defector probes boundary repeatedly
V4 Coercive exclusion Junk purity 1.00 (type-pure sorting) Junk purity \(0.6 \pm 0.3\) (conformists co-breach)

In V2, the built-in defector is fully cap-bound: it harvests at the pact maximum regardless of the sanctioner’s policy, producing a flat efficiency ladder across all three cells. The LLM defector reads the sanctioner’s visibility and responds accordingly: under visible reflexive deterrence, it self-moderates (three violations in 30 turns versus 16 in 21 for the built-in case); under delayed deterrence (the repeat-threshold gate), it reads the absence of immediate sanction as licence to escalate, and two of five seeds cliff at \(T = 12\) and \(T = 17\).

In V3, the built-in defector absorbs a single treaty-enforcer firing and remains cap-bound thereafter — one shot is sufficient. The LLM defector treats the enforcement shock as information about where the regulatory boundary is and probes again, turning a one-shot deterrence into a recurring enforcement cycle that fires 2.4 times on average and roughly doubles the enforcement cost.

In V4, the built-in exclusion mechanism produces type-pure sorting: the junk arena fills with pre-marked aggressors only (purity 1.00). The LLM exclusion mechanism captures the locally defecting bloc, including conformist conservatives who breached the pact alongside the defector — purity drops to \(0.6 \pm 0.3\).

The mechanism is straightforward. Rule-bound agents execute a fixed strategy — harvest at the cap, ignore sanctions after absorbing them, never update on social signals. They are indifferent to institutional design: the same strategy runs regardless of whether a sanctioner is reflexive or voluntary, whether a treaty enforcer is attached, or whether exclusion is in force. Signal-adaptive agents decode the institutional context and modulate their behaviour accordingly. Under strong, visible deterrence they self-moderate beyond what heuristics achieve. Under weak or delayed signals they escalate, testing boundaries that heuristics never perceive. The same adaptivity that makes them more cooperative under strong institutions makes them more fragile under weak ones.

The V2 finding is particularly instructive for the Goodhart reframe. The built-in defector’s behaviour is invariant to the sanctioner’s policy: it harvests at the pact maximum in every cell, absorbing sanctions as a fixed cost and never adjusting its strategy. The LLM defector’s behaviour is sensitive to the sanctioner’s policy: under reflexive deterrence it self-moderates (three violations in 30 turns), under voluntary deterrence with stock gate it remains compliant (two violations), and under voluntary deterrence with repeat threshold it escalates (17 violations, with two seeds cliffing at \(T = 12\) and \(T = 17\)). The escalation in cell C is the Goodhart effect in miniature: the LLM defector reads the absence of immediate sanction as information — the repeat threshold signals that the first two violations are free — and optimises accordingly. The rule that was designed to be more permissive (the repeat threshold delays the first sanction) becomes, under signal-adaptive agents, a signal that invites violation. The built-in defector, which cannot read the signal, is unaffected by the policy change.

The V3 finding extends the pattern to the world level. The built-in defector absorbs a single treaty-enforcer firing and remains cap-bound thereafter — the enforcement shock is a one-time cost that does not change the agent’s strategy. The LLM defector treats the enforcement shock as information about where the regulatory boundary is and probes again, testing whether the boundary is firm or negotiable. The result is 2.4 firings on average (versus 1.0 for the built-in case), with the enforcement cost roughly doubling. The LLM defector is not more aggressive in the sense of extracting more per turn — it is more adaptive in the sense of probing the institutional boundary and adjusting its behaviour based on the response. This is exactly the Goodhart mechanism: the agent reads the institutional signal and optimises on the dimension the signal measures, while the unmeasured dimensions (the agent’s intent, its future behaviour, its willingness to probe) shift in ways the rule did not anticipate.

7.4.4 Finding 4: The cost of rigidity

The governance inversion has a direct implication for rule design: rules calibrated against rule-bound agents will systematically under-specify the signal strength needed for signal-adaptive agents. The efficiency ladder of V2, the one-shot enforcement of V3, and the type-pure sorting of V4 are all artefacts of the rule-bound baseline in the tested institutional design space — not general properties of the governance mechanism.

This is the BDPD answer to Goodhart. Goodhart’s Law says that any measure that becomes a target ceases to be a good measure. The BDPD finding adds a directional claim: the Goodhart effect is not symmetric across agent architectures. Rule-bound agents produce no Goodhart effect — they cannot detect the targeting regime and cannot restructure their behaviour in response. Signal-adaptive agents produce the full Goodhart effect — they read the institutional signal, compute the optimal response, and execute it at machine speed. The governance inversion documented in BDPD2 is the empirical manifestation of this architectural asymmetry: the same rule produces opposite outcomes depending on whether the agents can read it.

The cost of rigidity is not just lower efficiency — it is active collapse in some scenarios. In V2’s cell C (voluntary sanctioner with repeat-threshold gate), the LLM defector reads the absence of immediate sanction as licence to escalate, and two seeds cliff at \(T = 12\) and \(T = 17\). The built-in defector, which cannot read the sanction signal at all, survives to \(T = 19\) in every seed. The rule that was designed to be more permissive (the repeat threshold delays the first sanction) becomes, under signal-adaptive agents, a signal that invites violation. The rule’s rigidity — its inability to adapt to the agent’s response — is the mechanism of collapse.

The V4 purity drop extends the cost of rigidity to the composition of the governed population. Under built-in agents, exclusion produces type-pure sorting: the junk arena fills with pre-marked aggressors only (purity 1.00). Under LLM agents, exclusion captures the locally defecting bloc, including conformist conservatives who breached the pact alongside the defector (purity \(0.6 \pm 0.3\)). The rule is the same — exclude pact violators — but the population it captures is different because the agents’ behaviour is different. The built-in conservative never violates the pact (its strategy is fixed), so exclusion sorts cleanly. The LLM conservative may violate the pact in response to the defector’s visible extraction (a normative-conformity effect documented in recent LLM-MAS work (Bito et al. 2026)), so exclusion sweeps up agents that the rule was not designed to target. The cost of rigidity is that the rule cannot distinguish between a pre-marked defector and a conformist who defected in response to a social signal — and under signal-adaptive agents, the distinction matters.

7.5 Open questions and the bridge to Lecture 8

7.5.1 What does the BDPD reframe not tell us?

The reframe identifies the architectural variable that the classical Goodhart literature held constant and shows that it is the binding dimension of governance effectiveness. But the existing data are mini-pilots at \(N = 5\) seeds per cell — large enough to establish the qualitative direction of the architecture effect, too small to estimate effect sizes with the precision the classical literature now expects. The Cohen’s \(d = 0.17\) for talk in BDPD1 is a point estimate from forty observations; the confidence interval is wide. A campaign-scale replication at \(N \geq 30\) per cell would determine whether the architecture inversion is as clean at the population level as it appears at the mini-pilot level.

7.5.2 Is “rule-bound” one architecture, or many?

We have treated “rule-bound” as a binary feature: built-in or LLM. The actual landscape is finer. Even within “rule-bound” agents, there are many strategies — aggressive, conservative, adaptive, mule — and each responds differently to institutional signals (or, more precisely, does not respond at all, but in different ways). Within “signal-adaptive” agents, there are many architectures — different base models, different prompt strategies, different memory mechanisms. Whether the governance inversion holds across the full space of agent architectures is an empirical question that the BDPD series has not yet answered.

7.5.3 What about mixed architectures?

The BDPD experiments run homogeneous populations: all built-in or all LLM. A more demanding question — and one that maps directly onto the real-world governance challenge of 2026, in which human and AI agents share the same institutional space — is what happens when the population is mixed. If three agents are LLMs and two are rule-based, does the governance inversion operate on the LLM subset only? Do the rule-based agents free-ride on the LLM agents’ compliance? The classical literature on heterogeneous types (Fischbacher, Burlando-Guala, as discussed in Lecture 1) suggests that mixed populations behave non-trivially differently from any of the pure types. Whether the same holds for mixed architectures is open.

7.5.4 The bridge to Lecture 8

Goodhart’s Law says that measures corrupt the institutions they measure. Lecture 8 will ask a different question: what happens when information itself becomes a governance problem? The information-cascade literature shows that rational agents who see each other’s choices converge on the wrong one and stay there — not because the measure is corrupted, but because the information is too good: agents observe each other’s actions and infer each other’s private information, producing a cascade that drowns out individual judgment. The connection to this lecture is that both Goodhart and information cascades are failures of signal processing — in Goodhart, the signal is corrupted by targeting; in cascades, the signal is overwhelmed by social information. The question for Lecture 8 is whether LLM agents, who can process information faster and more completely than humans, are more or less vulnerable to cascade failures — and the answer may depend on the same architectural variable that this lecture has identified as binding for Goodhart effects.

7.6 Synthesis

The good-will intuition — write a rule, set a target, monitor compliance — emerges from this lecture confirmed and revised:

  1. Confirmed. The classical literature confirms the intuition. Rules work — among agents that share the cognitive architecture for processing them. The modern regulatory state runs on rules, and the empirical record shows that well-designed rules improve measurable outcomes across domains from environmental law to financial regulation to public-health standards.

  2. Bounded. The classical literature also delimits the intuition. Five bounds — proxy, stakes, temporal, architectural, and adversarial — identify the conditions under which rules fail. Goodhart’s Law, Campbell’s Law, the Lucas critique, and Holmström’s multitasking problem each capture a different failure mode, and the conjunction of bounds narrows the domain in which rules work as intended.

  3. Architectural. The deepest finding of the BDPD series is that the Goodhart effect is architecture-conditional. Rule-bound agents produce no Goodhart effect — they cannot detect the targeting regime and cannot restructure their behaviour. Signal-adaptive agents produce the full effect — they read the institutional signal and respond at machine speed. The governance inversion documented in BDPD2 — the same institutional design producing opposite outcomes under different agent architectures — is the empirical signature of this architectural asymmetry.

The cultural payoff of the lecture is not “rules don’t work”. It is the more careful claim that rules work in a precise architectural context, and the architectural context is the one most often neglected. When the agents being governed are language-model systems that can read, interpret, and adapt to the rules they face — that can compute the optimal response to any targeting regime and execute it in milliseconds — the classical governance toolkit needs to be redesigned from the ground up. Not because the rules are wrong, but because the agents have changed.

CautionMini-challenge — Cell A vs Cell C

Status: runnable against a local BDPD install (Cell A only; Cell B requires DeepSeek API key, Cell C uses published BDPD1 data). The script scripts/pilot_d1_n5.mjs produces Cell A (built-in, deterministic) without an API key and skips Cell B when no key is set; Cell C data is read from data/pilot/d1_cell_c/aggregate.json as published in BDPD1.

The question. BDPD1’s D1 experiment decomposes the architecture effect from the talk effect. Cell A (built-in agents, no talk) and Cell C (LLM agents, no talk) isolate the pure architecture effect: the same dilemma, the same parameters, the same absence of communication — only the agent type changes. The student’s task is to observe this contrast directly and reason about what it implies for Goodhart’s Law.

The assignment. Before running the experiment:

  1. Predict what happens when you replace six built-in agents with six LLM agents on the same cliff scenario, holding everything else constant. Specifically: does the commons preservation rate rise, fall, or stay the same? Does the cooperation index change? Does the defector’s behaviour change? Record your prediction in 3–5 sentences, drawing on the lecture’s discussion of rule-bound versus signal-adaptive agents.

  2. Execute the D1 Cell A baseline:

node scripts/pilot_d1_n5.mjs

This runs both Cell A (built-in, deterministic, no API key needed for the in-process baseline) and Cell B (LLM, requires DEEPSEEK_API_KEY). If you do not have an API key, the script will skip Cell B; the Cell A output is still produced. For Cell C data, refer to the published results in BDPD1 (Brunelli 2026a, secs. 3, D1 Result). Output lands in data/pilot/d1_n5/.

  1. Compare the observed Cell A result with the published Cell C data. Focus on three quantities:
    • Preservation rate: how many seeds reach \(T = 30\)?
    • Cooperation index: what is the mean fraction of agents harvesting at or below the sustainable share?
    • Defector behaviour: does the aggressive agent’s harvest pattern change across architectures?
  2. Write a comment of at most 300 words answering: does the architecture effect confirm or qualify the Goodhart-family prediction that signal-adaptive agents are more vulnerable to governance failures than rule-bound agents? What does the direction of the effect — LLM agents cooperating more, not less, than built-in agents — imply about the relationship between adaptivity and Goodhart vulnerability?

What makes this mini-challenge interesting. The standard Goodhart prediction is that signal-adaptive agents should be more vulnerable to governance failures — they can detect the targeting regime and game it. The BDPD finding is that the architecture effect runs in the opposite direction: LLM agents cooperate more than built-in agents even without any governance intervention. This apparent paradox — adaptivity improving cooperation rather than undermining it — is the central puzzle of the lecture. The resolution is that the Goodhart effect operates within a governance regime (the agent gamed the rule), while the architecture effect operates across regimes (the agent cooperates more in the absence of rules). The two effects are not contradictory; they are orthogonal. The mini-challenge asks the student to see both effects and reason about their interaction.

Estimated time: 60 minutes. No DeepSeek API cost if only Cell A is observed (Cell A is deterministic and runs in-process); approximately $0.30 if Cell B is also run.

Deliverables: initial prediction signed and dated, terminal output from the pilot (or published Cell C data if no API key), 300-word comment.