Beyond Nudge: The Case for Decision
Decision Accounting
Beyond Nudge: The Case for Decision Accounting
core-claim
Core claim
Nudges can change behavior while making welfare costs harder to audit
The paper argues that standard nudge evaluation tracks the choice architect and the target agent, but leaves system welfare outside the payoff space. That frame can count a behavioral change as successful even when costs move into household finance, public budgets, credit markets, privacy, or the environment.
- The problem is not that nudges never work; the problem is that many are evaluated inside a bilateral frame.
- The paper estimates that current nudge evaluation practices undercount welfare costs by 40-70%.
- Decision Accounting is proposed as the rule change that makes system welfare a required part of intervention records.
retirement-default
Opening case
A 3% to 6% retirement default raised participation but hurt debt-constrained workers
The paper opens with a 2018 US retirement plan default change for new employees. Participation rose from 72% to 89% within the first year, which made the intervention look successful under standard nudge metrics.
- 14% of newly enrolled employees had high-interest credit card debt averaging $8,400.
- The higher contribution reduced liquid cash flow by $1,440 per year for affected workers.
- Credit card delinquency in that cohort rose 22% over 18 months.
- Two percent of the cohort defaulted on secured debt obligations.
flawed-game
Flawed game
The nudge game has two players, but three affected parties
The paper models nudge interventions as a bilateral game G between the choice architect A and the target agent B. The broader system C is affected by the intervention but is not represented in the payoff functions.
- A is scored on behavioral change metrics such as participation, opt-out, click-through, and compliance rates.
- B is scored on individual outcomes such as savings, energy use, or medication adherence.
- C includes the economic, social, environmental, or institutional system that absorbs spillover costs.
hollow-win
Outcome taxonomy
The key failure mode is the Hollow Win: C=0, A=1, B=1
Using the paper's 8-outcome taxonomy, the Hollow Win is the outcome where the architect and target both register gains while system welfare deteriorates. The retirement default example fits this pattern when higher savings are offset by debt costs and credit damage.
- Win-Win-Win: default enrollment with adequate liquidity protections.
- Hollow Win: default enrollment that increases credit card debt.
- Corrosive Win-Lose: an aggressive default that harms the target and the system.
- Misery: intervention failure with system damage.
mst
Missing System Theorem
MST says system welfare is outside the bilateral payoff space
The Missing System Theorem states that in any bilateral economic game G between A and B, the system-welfare dimension W sits outside the payoff space. The paper applies this to behavioral policy to explain why apparently Pareto-improving nudges can still degrade C.
- Only 12% of nudge studies measure any welfare consequence beyond targeted behavior, according to Mertens et al. (2022).
- Large-scale government nudge effects are 60-80% smaller than academic-study effects, according to DellaVigna and Linos (2022).
- Postnieks (2024) reports the same exclusion pattern in financial regulation, with βW estimates of 4.1-7.3.
nit
1
The Nudge Intractability Theorem follows from payoff exclusion
The theorem states that any nudge designed inside the standard behavioral economics framework excludes C from the bilateral payoff space. A and B can reach a Nash equilibrium even when system welfare is degraded.
- A1: π depends on behavioral change metrics, and π depends on individual choice outcomes.
- Neither π nor π contains a term for C.
- A3: disclosure about system consequences does not change the equilibrium set unless payoffs change.
- A6: when designs yield equivalent bilateral metrics, architects select without penalty for C=0.
prevalence
2
Hollow Wins are predicted to outnumber Win-Win-Win outcomes
The Hollow Win Prevalence Theorem states that implemented nudge equilibria are more likely to be C=0, A=1, B=1 than C=1, A=1, B=1. The reason is selection bias: C is outside the payoff function, so C=0 designs are not screened out by the bilateral game.
- The theorem does not require every nudge to be harmful.
- It predicts a systematic bias in implemented interventions, not a universal outcome.
- The paper makes the claim falsifiable through a pre-registered meta-analysis threshold.
evidence
base
Four intervention classes generate at least $6.7 billion in documented welfare destruction
The βW Lower Bound Theorem aggregates documented welfare losses across four classes of behavioral intervention. The paper uses these figures to estimate the lower bound for welfare destruction per dollar of industry revenue.
- Retirement default effects: at least $4.2 billion.
- Energy report effects: at least $0.8 billion.
- Administrative burden effects: at least $1.2 billion.
- Digital friction effects: at least $0.5 billion.
beta-w
βW lower bound
The paper estimates βW at 3.2-7.8 for the behavioral intervention industry
βW is defined as −dW/dΠ, where W is system welfare and Π is industry revenue. Using 6.7 billion in documented welfare destruction and industry revenue of no more than 2.1 billion, the paper proves βW ≥ 3.2.
- The abstract reports a βW range of 3.2-7.8.
- The lower-bound proof uses 6.7B divided by 2.1B.
- The paper states this is a lower bound because documented cases capture only a subset of system welfare effects.
decision-accounting
Decision Accounting
DA transforms G into G1 by adding C to the required record
Decision Accounting is the paper's proposed rule change R. It is a 17-field structured decision protocol that turns system welfare from an optional afterthought into a mandatory dimension of each choice-architecture decision.
- Field 17, SYSTEMWELFARE , requires analysis of hidden costs, burden-shifting, and affected stakeholders.
- Field 16, PREDICTION, ties material claims to scoreable outcomes.
- Field 9, REVIEW, creates reversal triggers when C=0 outcomes appear.
- The transformed game is G1 = A, B, C .
mandate
Mandate theorem
Voluntary DA adoption fails because costs fall on A while gains accrue to C
The Mandate Necessity Theorem states that a DA mandate is necessary and sufficient for transforming G into G' with a structurally better equilibrium. Voluntary adoption is intractable because architects bear compliance costs, while system welfare gains accrue outside the bilateral game.
- A4: upfront compliance cost is borne by A.
- Welfare gains accrue primarily to C, which has no bargaining power in G.
- Non-adopting competitors gain a cost advantage.
- A mandate changes π by adding penalties for DA noncompliance and documented C=0 outcomes.
reform-dividend
Reform dividend
The paper estimates at least $12 billion per year from a DA mandate in OECD countries
The Reform Dividend Theorem defines the gain as Wmandate minus Wcurrent. The lower bound combines avoided welfare losses from the four intervention classes plus a conservative estimate for unmeasured system welfare improvements.
- Avoided retirement default welfare destruction: $4.2-8.8 billion.
- Avoided energy report moral licensing: $0.8-1.6 billion.
- Reduced administrative burden improper payments: $2.8-5.6 billion.
- Improved digital consent quality: $0.5-1.2 billion.
- Unmeasured system welfare improvements: $3.7-10.8 billion.
falsification
Falsifiability
The framework names empirical tests that would overturn its claims
The paper closes by specifying falsification conditions rather than asking readers to accept the framework as unfalsifiable. These tests target Hollow Win prevalence, voluntary adoption, disclosure-only interventions, βW, DA mandates, and the formal status of C in bilateral games.
- 2 is refuted if a pre-registered meta-analysis of at least 50 interventions finds (1,1,1) exceeds (0,1,1) by a statistically significant margin at p < 0.05.
- Corollary 1.2 is refuted if a survey of at least 200 organizations finds at least 20% voluntary DA-like adoption maintained for 3 years.
- 3 is refuted if an independent audit of at least 50 major organizations estimates βW below 1.0 with a 95% confidence interval excluding values ≥ 1.0.