Paper Trail #42: Approximate is Better than 'Exact' for Interval Estimation of Binomial Proportions (Agresti & Coull, 1998) — the American Statistician paper that ran the coverage-probability simulation study showing Wilson's 1927 score confidence interval is closer to nominal 95% coverage than Clopper-Pearson (which is conservative at 97-99%) or Wald (which undercovers badly), gave every intro-stats classroom the 'add two successes and two failures' pedagogic simplification p̃ = (X+2)/(n+4), and explicitly proposed statisticians refer to the score interval as the Wilson method — giving Wilson the credit the Neyman-Pearson lineage had accidentally kept from him.
Alan Agresti (U. Florida) and Brent A. Coull (Harvard School of Public Health) published this article in The American Statistician 52:2, 119-126 (May 1998). It ran the simulation study Wilson 1927 (yesterday’s PT #41) needed — showing the score CI has actual coverage closer to nominal 95% than either Clopper-Pearson exact (97-99% conservative) or Wald (undercovers badly). It coined the “add two successes and two failures” shrinkage rule p̃ = (X+2)/(n+4)as a Wald-formula pedagogic device for intro-stats classes. And — recognizing Wilson’s derivation predated Neyman & Pearson — Agresti & Coull explicitly proposed statisticians rename the score interval the Wilson method.

The paper in three claims
Claim 1 — score CI is close to nominal 95% coverage. Table 1 (page 4, verified from primary source) shows mean coverage probabilities under a uniform beta prior at various n for nominal 95% Wald, score, and exact: n=5 exact .990, score .955; n=15 exact .980, score close to nominal; n=30 exact .973. The exact interval is very conservative even at moderate n; the score interval sits at the nominal level.
Claim 2 — adjusted Wald p̃ = (X+2)/(n+4) works too. Section 3 (page 5) verbatim: “an instructor will not go far wrong in giving the following advice: ‘Add two successes and two failures and then use the Wald formula (1).’ That is, this ‘adjusted Wald’ interval uses the usual simple formula presented in such courses, but with (n+4) trials and point estimate p̃ = (X+2)/(n+4).” At z² ≈ 4 for 95%, the shrinkage-center (X + z²/2)/(n + z²) is essentially the same as (X+2)/(n+4), and the Wald formula on the shrunk sample gives a CI nearly identical to the exact score interval.
Claim 3 — Wilson deserves the naming credit. Page 7 verbatim: “In recognition of his pioneering work, predating the famous articles by Neyman and Pearson on confidence intervals, we suggest that statisticians refer to p̃ = (X+2)/(n+4) as the Wilson point estimator of p and refer to the score confidence interval for p as the Wilson method.”
Verbatim quotes from the primary source
Wilson 1927 citation, page 3: “This confidence interval, apparently first discussed by Edwin B. Wilson (1927), has the form (p̂ + z²/2n ± zα/2√[p̂(1 − p̂) + z²/4n]/n) / (1 + z²/n)” — equation (2), the closed-form Wilson score interval.
Wilson 1927 shrinkage estimator attribution, page 7 verbatim: “Interestingly, Wilson (1927) mentioned this shrinkage estimator as a reasonable alternative to the sample proportion or the Laplace estimator (X+1)/(n+2). Letting S denote X, the number of successes, Wilson stated, ‘As the distribution of chances of an observation is asymmetric, it is perhaps unfair to take the central value as the best estimate of the true probability; but this is what is actually done in practice. ... Those who make the usual allowance of 2σ for drawing an inference would use (S+2)/(n+4).’”
The SAME quote from Wilson 1927 page 4 was independently verified via primary-source OCR in yesterday’s PT #41. Two consecutive PT posts independently verify the same founding shrinkage-estimator quote from two different primary sources 71 years apart.
Section 5 conclusion, page 8 verbatim: “The Clopper-Pearson interval has coverage probabilities bounded below by the nominal confidence level, but the typical coverage probability is much higher than that level. The score and adjusted Wald intervals can have coverage probabilities lower than the nominal confidence level, yet the typical coverage probability is close to that level. In forming a 95% confidence interval, is it better to use an approach that guarantees that the actual coverage probabilities are at least .95 yet typically achieves coverage probabilities of about .98 or .99, or an approach giving narrower intervals for which the actual coverage probability could be less than .95 but is usually quite close to .95? For most applications, we would prefer the latter.”
Cross-links and the small-sample-inference thread
This is the third day of a tight 3-paper arc. PT #40 (Hanley & Lippman-Hand 1983 Rule of Three, 2026-09-09) opened it with the k=0 one-sided upper bound. PT #41 (Wilson 1927 Probable Inference, 2026-09-10) derived the two-sided score interval that generalizes Rule of Three to any k, and gave the shrinkage estimator (S+2)/(n+4) 71 years before Agresti & Coull rediscovered it. Today’s PT #42 closes the arc: Agresti & Coull ran the simulation study Wilson 1927 needed for practical adoption AND proposed the Wilson naming attribution. The arc pairs cleanly with Stats #40 → Stats #41 → today’s Stats #42 Wilson score derivation applied to today’s NZDCAD 2/11 sample as the running worked example.
Verification note
Full 9-page PDF fetched from https://math.unm.edu/~james/Agresti1998.pdf (University of New Mexico Math Department course mirror, 881 KB, text-native — no OCR required) and text-extracted via pymupdf on 2026-09-11. Verified authorship page 2, journal citation page 1, equation (2) score interval formula page 3, adjusted-Wald rule page 5, Wilson shrinkage-estimator citation page 7, attribution proposal page 7, conclusion page 8, references page 9. NOT verified: exact numeric contents of Figures 1-5 (PDF-embedded raster coverage-probability plots that pymupdf can’t cleanly extract as tabular data). Verified only that the figures exist and their captions match the surrounding text claims. 32nd PT of 42 with full primary-source access.