Statistics for Traders #46: The seven methods Newcombe (1998) compared for a single-proportion 95% CI — the methodological survey that unifies Stats #40-#44. Applied to today's slot 1 GBPJPY 5m big_miss 0-of-4 UP: methods 3 Wilson [0%, 48.99%] and 7 Likelihood [0%, 38.13%] reject fair-coin; methods 4 Wilson+CC, 5 Clopper-Pearson, and 6 Mid-p all fail. Newcombe's mean coverage on 96,000 parameter-space points (Table II): method 3 = 0.952, method 6 = 0.957 — both closest to nominal 0.95. Newcombe's Section 8 recommendation: for the CP=1-α criterion, methods 3 (Wilson) or 6 (mid-p); method 3 has the simpler closed form.
Newcombe (1998) evaluated seven CI methods for a single proportion on 96,000parameter-space points and recommended methods 3 (Wilson score) or 6 (mid-p) for the CP=1-α criterion, method 5 (Clopper-Pearson) for the strict- conservative min-CP ≥ 1-α criterion, and argued against methods 1 and 2 (Wald with and without continuity correction) for the scientific literature. Applied to today’s slot 1 GBPJPY 5m big_miss 0-of-4 UP sample, only Newcombe’s method 3 (Wilson, [0%, 48.99%]) and method 7 (likelihood, [0%, 38.13%]) reject fair-coin at 95%; methods 4, 5, and 6 all fail because their upper bounds sit above 50% at this boundary-of-support sample.
![Horizontal comparison of 7 Newcombe (1998) methods' 95% confidence intervals on the GBPJPY 5m big_miss 0-of-4 UP sample. Method 1 Wald: [0%, 0%] — degenerate zero-width interval. Method 2 Wald+CC: [0%, 24.50%] — lower bound would violate 0 before truncation. Method 3 Wilson score: [0%, 48.99%] — rejects fair-coin. Method 4 Wilson+CC: [0%, 60.42%] — fails. Method 5 Clopper-Pearson exact: [0%, 60.24%] — fails. Method 6 Mid-p: [0%, 52.71%] — fails. Method 7 Likelihood: [0%, 38.13%] — rejects with TIGHTEST upper bound. A vertical dashed line at 50% marks the fair-coin decision threshold. Methods 3 and 7 sit entirely to the left of the 50% line.](/insights/stats-for-traders-newcombe-1998/methods.png)
The seven Newcombe (1998) methods, numbered as in Section 2
| # | Method | Stats series link | Primary source |
|---|---|---|---|
| 1 | Wald (simple asymptotic, no CC) | — | textbook standard |
| 2 | Wald with continuity correction | — | Fleiss 1981; Blyth-Still 1983 |
| 3 | Wilson score (no CC) | Stats #42 | Wilson 1927 (PT #41) |
| 4 | Wilson score with continuity correction | — | Fleiss 1981; Blyth-Still 1983 |
| 5 | Clopper-Pearson exact | Stats #41 + Stats #40 one-sided form | Clopper-Pearson 1934 |
| 6 | Mid-p binomial-based | — | Berry-Armitage 1995; Miettinen 1985 |
| 7 | Likelihood-based | — | Miettinen-Nurminen 1985 |
Methods 3 and 6 (coral) are Newcombe’s Section 8 recommendation for the CP=1-α criterion. Stats #43 Agresti-Coull and Stats #44 Jeffreys are NOT in Newcombe’s seven — AC is Wald-with-adjusted-count (Newcombe cites Wilson 1927 as the primary source of the +λ²/2 adjustment on p. 861 and does not count it separately); Jeffreys is Bayesian (Newcombe’s Section 2 excludes Bayesian intervals on frequentist grounds).
Applied to today’s slot 1 GBPJPY 5m big_miss 0-of-4 UP
| Method | 95% CI on up-rate | Rejects fair-coin? | Newcombe aberration? |
|---|---|---|---|
| 1 Wald | [0%, 0%] | REJECTS (spurious) | ZWI (§2, p. 858) |
| 2 Wald+CC | [0%, 24.50%] | REJECTS (spurious) | overshoot (§2, p. 858) |
| 3 Wilson | [0%, 48.99%] | REJECTS | none |
| 4 Wilson+CC | [0%, 60.42%] | no | none |
| 5 Clopper-Pearson | [0%, 60.24%] | no | none |
| 6 Mid-p | [0%, 52.71%] | no (just fails) | none |
| 7 Likelihood | [0%, 38.13%] | REJECTS (TIGHTEST) | none but min CP 0.802 |
Methods 1 and 2 “reject” because their CI collapses below the fair-coin threshold — but only through the ZWI and overshoot aberrations that Newcombe explicitly rules out. The legitimate rejections are methods 3 (Wilson) and 7 (likelihood). Method 7 gives the TIGHTEST upper bound (38.13%), but Newcombe warns method 7 “is in fact slightly anti-conservative, with average coverage probability 0)948” and can dip to min CP = 0.802 (Newcombe Table II) — the aggressive rejection carries genuine risk. Method 3 is the safer rejection here.
Newcombe Table II — mean coverage across 96,000 PSPs
| Method | Mean CP | Min CP | Reading |
|---|---|---|---|
| 1 Wald | 0.8814 | 0.0002 | heavily anti-conservative |
| 2 Wald+CC | 0.9257 | 0.3948 | still anti-conservative |
| 3 Wilson | 0.9521 | 0.8322 | CLOSEST to 0.95 (recommended) |
| 4 Wilson+CC | 0.9707 | 0.9491 | conservative |
| 5 Clopper-Pearson | 0.9710 | 0.9501 | strict-conservative gold standard |
| 6 Mid-p | 0.9572 | 0.9121 | CP=1-α recommended |
| 7 Likelihood | 0.9477 | 0.8019 | slightly anti-conservative |
Newcombe’s decision rule for choosing between methods 3, 5, and 6 comes from the explicit decisionwhether “nominal 1-α should represent a minimum, [in which case] methods that are strictly conservative… should be chosen” (p. 869 — method 5 or 4) or “1-α is construed as an average, [in which case] CP should approximate to 1-α” (methods 3 or 6). The Wilson score interval’s simple closed form (Stats #42 formula reproduced verbatim from Wilson 1927 Section 2) is what makes it the practical CP=1-α default over method 6.
Cross-checks on three past-week samples
| Sample | 3 Wilson | 5 CP | 6 Mid-p | 7 Likelihood |
|---|---|---|---|---|
| NZDCAD 2/11 UP (2026-09-11) | [5.14, 47.70]✓ | [2.28, 51.78]✗ | [3.17, 48.27]✓ | [3.28, 46.30]✓ |
| JPY CPI 4h 4/16 UP (2026-09-11) | [10.18, 49.50]✓ | [7.27, 52.38]✗ | [8.49, 49.89]✓ | [8.52, 48.90]✓ |
| GBPJPY 5m 0/4 UP (today) | [0, 48.99]✓ | [0, 60.24]✗ | [0, 52.71]✗ | [0, 38.13]✓ |
✓ = rejects fair-coin at 95%; ✗ = fails. Method 3 (Wilson) rejects on ALL three samples; method 5 (CP) rejects on NONE; methods 6 and 7 line up with method 3 except at the small-n boundary (0/4). Same pattern the Stats #42 (Wilson) and Stats #43 (Agresti-Coull) posts documented, now placed inside Newcombe’s full 7-method framework.
Verification note
All seven Newcombe methods implemented from Newcombe (1998) Section 2 formulae and cross-verified: method 3 vs statsmodels.stats.proportion.proportion_confint(method="wilson"); method 5 vs scipy.stats.beta.ppf; methods 6 and 7 vs custom binomial-PMF root-finding via scipy.optimize.brentq. Table II mean coverage numbers verified verbatim against today’s PT #46 (Newcombe 1998 primary source, Statistics in Medicine 17:857-872, pp. 866 Table II). Chart built via a one-off script reusing scripts/insights-charts/svg.ts + theme.ts + sharp, not committed under scripts/.