Validity of the Hypergeometric
Decision Accounting

Validity of the Hypergeometric Convergence Test

intro
Core finding

Sixteen regimes jointly require 14 documentation fields with P < 0.001

The paper defends the Decision Accounting convergence result against three methods objections: the wrong null model, coder dependence, and selection bias. The reported statistic is the intersection size of regime-level requirement vectors, not a reliability score.

objections
Reviewer objections

The paper answers three threats to the convergence test

The reviewer argues that hand-coded prose is not an urn draw, that the same two coders could manufacture overlap, and that the sixteen regimes may have been chosen because they already looked similar.

vectors
Data structure

Each regime is reduced to a fixed binary vector over the same 47-field taxonomy

The test operates after coding. Each regime i has a binary vector vi of length M, where vij = 1 if the regime requires field j and 0 otherwise. The test uses those fixed vectors as inputs.

null
Null model

The null is independent random subsets with each regime’s observed size held fixed

The hypergeometric null asks how much overlap would occur if each regime independently selected exactly ki fields from the shared pool of M fields. It does not claim regulators actually choose fields randomly.

hypergeometric
Why hypergeometric

The test matches the overlap-of-random-subsets problem

The paper’s target question is set overlap: how many fields would all N regimes share by chance, given each regime’s fixed number of required fields. That is the combinatorial setting for the hypergeometric distribution.

conservative
Conservative assumption

Equal field probability makes rejection harder under the paper’s argument

The null treats all 47 candidate fields as equally likely before selection. The paper argues this inflates the chance of spurious overlap compared with a world where some fields are naturally more likely than others.

reliability
Separate quantities

C = 14 is cross-regime overlap, while alpha = 0.89 is coder reliability

The paper says the methods dispute comes partly from conflating two forms of agreement. Krippendorff’s alpha measures whether coders agree on field attributions. C measures whether regimes share required fields.

coder
Coder dependence

The hypergeometric P-value depends on C, M, and ki, not on who coded the vectors

1 states that once the regime vectors are fixed, the convergence statistic and its P-value are functions only of the vectors and their sizes. Coder dependence can bias inputs, but it does not enter the conditional null calculation.

selection
Selection

The sixteen regimes were selected as major documentation regimes, not on overlap

The paper’s selection defense is that the sample was purposive but not selected on the dependent variable. The criteria were binding regulatory framework, documentation requirements, and public English source text.

isomorphism
Isomorphism

The paper rejects copying because shared structure appears with different vocabulary

The institutional-isomorphism alternative says regulators may have copied one another. The paper argues the corpus has the opposite signature: the same field appears through different institutional language.

lineage
Lineage evidence

Disjoint agencies weaken the diffusion explanation

The paper adds an institutional-lineage argument against copying. The relevant agencies and regimes come from different eras, continents, and professional settings, with no formal coordination mechanism for decision documentation requirements.

falsify
Falsification

The claim fails if blind coders cannot reproduce the regime vectors

The paper makes reproducibility the empirical test. Independent coders must use the published source texts and Appendix F protocol to reproduce the requirement vectors, then the convergence result must survive the same test.