A Disciplined Cross-Section
Decision Accounting
A Disciplined Cross-Section: Bayesian Model Averaging over the Factor Zoo Under the Joint-Hypothesis Constraint
intro
Core claim
Fama's joint-hypothesis problem is formalized as a Bayesian posterior decomposition, delivering a stopping rule for factor inclusion
The paper constructs a Bayesian model-averaging framework over the full factor zoo (K ≥ 200) that treats the joint-hypothesis constraint as an explicit posterior object, not a rhetorical device.
- Likelihood decomposed into pricing-model adequacy and market-efficiency components, both estimable
- Posterior-odds stopping rule derived from the joint-hypothesis structure, asymptotically equivalent to FDR control
- Nests FF3, FF5, HLZ t=3, and KMZ complexity as special-case prior configurations
problem
The factor zoo problem
Over 400 candidate factors exist, with 65% replication failure, creating a triple crisis
Harvey, Liu, and Zhu (2016) catalogued 316 factors; Hou, Xue, and Zhang (2020) found 65% of 452 anomalies fail replication. The field faces replication, multiple-testing, and model-selection crises simultaneously.
- Chen-Zimmermann (2022) open-source library provides 207 replicable factors
- Fama's own progression from 3 to 5 factors implies a need for a stopping rule
- Existing answers (HLZ t=3, FDR, dense models) are not cast in Fama's language
formalization
Joint-hypothesis formalized
The joint-hypothesis constraint is written as a Bayesian likelihood factorization
Fama (1970) stated the constraint verbally; this paper writes it as p(DMj,E) = Lpricing · Lefficient, making both components independently estimable.
- Posterior on efficiency is model-averaged: p(ED) = Σ p(Mj, E=1D)
- Previous Bayesian work (Barillas-Shanken 2018) limited to 6 models, not 2K
- Spike-and-slab priors enable stochastic search over 1060 model space
priors
Prior specification
Two Fama-consistent hyperparameters encode parsimony and theory motivation
π encodes preference for compact models; π boosts prior inclusion probability for theory-motivated factors (MKT, SMB, HML, RMW, CMA, momentum).
- Parsimony prior: p(fk ∈ Mj) = π · g(Mj; π ), decreasing in model size
- Theory prior: p(fk ∈ Mj) ∝ exp(π ) for theory factors, 1 otherwise
- No assumption hard-coded; every assumption is a tunable prior
decomposition
Likelihood decomposition
Residual alphas are decomposed into omitted-risk vs. genuine-mispricing components
Lefficient compares probability that observed alphas arise from omitted systematic risk versus true mispricing, using persistence, cross-sectional structure, and international covariation.
- Persistence: omitted-risk alphas persistent; mispricing alphas transient (McLean-Pontiff 2016)
- Cross-sectional structure: omitted-risk alphas load on firm characteristics
- International covariation: systematic risk appears across markets (Fama-French 2012, 2017)
empirical
Posterior inference
Gibbs sampler with spike-and-slab priors identifies 7–9 factors in U.S. data
500,000 MCMC draws (100,000 burn-in) over 207 factors from Chen-Zimmermann library, using CRSP/Compustat 1963–2025. Baseline hyperparameters π =1.0, π =1.0.
- 8 factors exceed posterior inclusion threshold τ*=0.72 at 5% Bayesian FDR
- FF5 factors all survive (MKT 0.998, SMB 0.94, HML 0.91, RMW 0.89, CMA 0.85)
- Momentum (0.82), short-term reversal (0.78), accruals (0.74) also included
theorem
Stopping rule theorem
Posterior-odds stopping rule controls Bayesian FDR and is asymptotically equivalent to Benjamini-Hochberg
1: Include factor fk if p(fkD) > τ*, where τ* is derived from priors and data via the joint-hypothesis decomposition — not chosen by convention.
- Bayesian FDR control follows Newton et al. (2004) direct-posterior-probability approach
- Posterior consistency: p(fkD) → 1 for true factors, 0 for false (Narisetty-He 2014)
- τ* = f(π , π , Lefficient) — emerges from framework, not arbitrary
nesting
Nesting results
Framework nests FF3, FF5, HLZ t=3, and KMZ complexity as special prior cases
Under strong parsimony and 1963–1991 data, posterior mode recovers FF3. Weakened parsimony with 2015 data recovers FF5. Flat priors reproduce HLZ t=3. Zero parsimony reproduces KMZ dense-model result.
- FF3: π →∞, π >>0, data 1963–1991 → MKT, SMB, HML
- FF5: π moderate, data to 2015 → MKT, SMB, HML, RMW, CMA
- HLZ: flat priors, τ → t≈3 threshold; KMZ: π →0 → ridge regression
replication
International replication
7–9 factor posterior holds across European, Japanese, and Asia-Pacific markets
Using Fama-French international datasets, the posterior identifies similar factor sets, confirming robustness. Out-of-sample cross-sectional R² competitive with dense ML approaches.
- International replication follows Fama-French (2012, 2017) methodology
- Posterior mode at 8 factors, 90% credible interval [6, 11]
- Models with >15 factors have <0.02 posterior probability
welfare
Welfare implications
Factor model misspecification propagates welfare costs through pension, insurance, and sovereign debt
The Private Pareto Theorem (Postnieks 2026a) shows bilateral optimality does not guarantee system-level welfare. Getting factor pricing right is necessary for efficient capital allocation.
- Factor models determine cost of capital, pension expected returns, insurance risk assessments
- Mispricing propagates through credit markets, pension funding, insurance, sovereign debt
- BMA framework is a methodological contribution with welfare implications across 61 domain papers
conclusion
Conclusion
The framework resolves a sixty-year gap: Fama's implied stopping rule is now explicit
The posterior-odds rule answers 'when do you stop adding factors?' within Fama's own methodological frame. The answer is a procedure, not a number.
- Rational-equilibrium null maintained throughout; no behavioral content
- Five falsification bounties (F1–F5) specify exact conditions for each core claim to fail
- Open-source replication: seed=20260419, 500,000 draws, full repository documented