Statistics for Traders #51: Newcombe (1998b) unpaired-difference CI — Method 10 (hybrid Wilson score) applied to today's slot 1 AUDCAD 15m unpaired 2x2 (big_miss up-rate 0/3 vs big_beat up-rate 4/4). δ̂ = -1.0. Method 10 gives CI [-1, -0.255]; Method 1 (Wald) collapses to zero-width tethering aberration [-1, -1]; Method 2 (Wald+CC) overshoots to [-1.167, -0.833]. Fisher's exact one-sided p = 1/35 = 0.0286 same decision, plus effect-size information. Stacked on 4 direct-quote AUD GDP legs (n=28) the Method 10 CI tightens to [-1, -0.690] and Fisher's exact drops to 3.29e-8. Closes the Newcombe 1998 trilogy alongside Stats/PT #46 (1998a) and Stats/PT #50 (1998c).
Newcombe (1998b) is the UNPAIRED-difference companion to Stats #46 (1998a single-proportion) and Stats #50 (1998c paired-difference). Method 10 is the closed-form hybrid Wilson score CI on δ = p1 - p2 for two independent samples. Applied to today’s slot 1 AUDCAD 15m 2x2 (big_miss up-rate 0/3, big_beat up-rate 4/4, unpaired): δ̂ = -1.0. The naive Wald collapses to zero-width [-1, -1] — the leading tethering aberration Newcombe 1998b addresses. Method 10 gives 95% CI [-1, -0.255] — proper interval that excludes zero. Closes the Newcombe 1998 trilogy.
![Horizontal bar chart showing four 95% confidence intervals on the unpaired difference δ = p1 - p2 for today's AUDCAD 15m 2x2 (big_miss up-rate 0/3 vs big_beat up-rate 4/4). Method 1 Wald: zero-width interval tethered at -1 (labelled 'zero-width, tethered'). Method 2 Wald with continuity correction: [-1.167, -0.833] — overshoots the valid range [-1, +1] on the lower end. Method 4 Wilson score without hybrid combination: [-1, -0.293]. Method 10 hybrid Wilson score (highlighted in green): [-1, -0.255]. Point estimate δ̂ = -1.0 marked. Vertical coral line at δ=0 shows all four intervals exclude the null. Bottom panel shows Fisher's exact test's one-sided p = 1/35 = 0.0286 for comparison.](/insights/stats-for-traders-newcombe-1998b-unpaired/methods-vs-fisher.png)
Method inventory — Newcombe (1998b) eleven methods (Section 3)
| # | Method | One-liner |
|---|---|---|
| 1 | Wald (no CC) | δ̂ ± z·se; se = √(p̂1(1-p̂1)/n1 + p̂2(1-p̂2)/n2). Tethering aberrations. |
| 2 | Wald + CC | δ̂ ± (z·se + 1/(2·min(n1,n2))). Overshoots [-1, +1] on tail data. |
| 3 | Yates-corrected chi-square-based | Extended-precision Wald with Yates correction on each 2x2 term. |
| 4 | Wilson score (unhybridised) | Direct Wilson inversion applied to the log-odds pooled 2-sample score. |
| 5 | Beal 1987 profile likelihood approx | Profile-likelihood CI with a Beal 1987 correction. |
| 6 | Farrington-Manning 1990 score | Score-based on the null p1 - p2 = δ over a delta grid. |
| 7 | Newcombe hybrid Wald + Wilson | Uses Wilson centres in the Wald combination formula. |
| 8 | Mee-Miettinen-Nurminen score (v1) | Score-based on shrinkage-adjusted proportions. |
| 9 | Miettinen-Nurminen 1985 score | Miettinen's original iterative delta-grid score inversion. |
| 10 | Wilson score + hybrid ★ | Wilson (Stats #42) for each proportion, combined via Newcombe's hybrid formula. Closed-form default. |
| 11 | Mee 1984 score iterative | Newcombe's top recommendation for coverage but requires iteration. |
Newcombe (1998b) Section 6 verbatim recommendation (per the paper’s abstract, verified via NCBI eutils on 2026-09-20 — see PT #51 for access caveat): “A tail area profile likelihood based method [method 11 / Mee-MN] produces the best coverage properties, but is difficult to calculate for large denominators. A method combining Wilson score intervals for the two proportions to be compared [method 10] also performs well, and is readily implemented irrespective of sample size.”
Method 10 worked example — today’s AUDCAD 15m 2x2
| Step | Value |
|---|---|
| p̂1 = big_miss up-rate = 0/3 | 0.0000 |
| p̂2 = big_beat up-rate = 4/4 | 1.0000 |
| δ̂ = p̂1 - p̂2 | -1.0000 |
| Wilson (Stats #42) 95% CI for p1 | [0.0000, 0.5615] |
| Wilson 95% CI for p2 | [0.5101, 1.0000] |
| L(δ) = δ̂ - √((p̂1-L_1)² + (U_2-p̂2)²) | -1.0 - √(0 + 0) = -1.0000 |
| U(δ) = δ̂ + √((U_1-p̂1)² + (p̂2-L_2)²) | -1.0 + √(0.3153 + 0.2400) = -0.2548 |
| Method 10 95% CI on δ | [-1.0000, -0.2548] ★ excludes 0 |
Method 10 delivers a proper (finite-width, valid-range) 95% CI that excludes zero and correctly reflects the uncertainty from the small sample sizes (n1 = 3, n2 = 4). Same rejection as Fisher’s exact test (one-sided p = 1/35 = 0.0286), plus effect-size information the p-value alone doesn’t carry.
Method comparison on the same 2x2
| Method | 95% CI on δ | Notes |
|---|---|---|
| 1 Wald (no CC) | [-1.0000, -1.0000] | ZERO-WIDTH tethering |
| 2 Wald + CC | [-1.1667, -0.8333] | OVERSHOOTS [-1, +1] |
| 10 Wilson hybrid ★ | [-1.0000, -0.2548] | proper, excludes 0 |
| Fisher's exact (Stats #45) | one-sided p = 1/35 | rejects at α=0.05 |
Method 1 fails both criteria — zero-width intervals convey false certainty. Method 2 partially fixes it via a continuity correction but overshoots the valid range on tail-unanimity samples. Method 10 is the closed-form default that survives both aberrations and is easy to compute — no iteration required.
Stacked sample — how Method 10 tightens with larger n
| Sample | Method 10 95% CI | Fisher exact one-sided p |
|---|---|---|
| Slot 1 AUDCAD alone (n1=3, n2=4) | [-1.0000, -0.2548] | 1/35 = 2.86e-2 |
| Stacked 4 direct-quote pairs (n1=12, n2=16) | [-1.0000, -0.6897] | 3.29e-8 = 1/3.0e7 |
Combining today’s 4 direct-quote AUD GDP legs (AUDUSD + AUDJPY + AUDCAD + AUDCHF, all 0/3 miss + 4/4 beat) gives stacked p1 = 0/12 and p2 = 16/16. Method 10 tightens the upper bound from -0.255 (single-pair) to -0.690 (4-pair stack). Fisher’s exact drops by 7 orders of magnitude to 3.29e-8. Caveat: the 4 legs are NOT independent — same 7 print dates on 4 correlated pairs — so the stacked p-value is optimistic. A proper joint test would use Newcombe 1998c method 10 with correlation adjustment across the four legs, or a Fisher combined test that accounts for the pair-correlation matrix.
Trilogy comparison — 1998a vs 1998b vs 1998c
| Paper | Problem | Methods | Stats/PT link |
|---|---|---|---|
| 1998a 17(8):857-872 | Single-proportion CI | 7 | Stats/PT #46 (2026-09-15) |
| 1998b 17(8):873-890 ★ today | Unpaired-difference CI | 11 | Stats/PT #51 today |
| 1998c 17(22):2635-2650 | Paired-difference CI | 10 | Stats/PT #50 (2026-09-19) |
Same author (Robert G. Newcombe, Cardiff), same journal (Statistics in Medicine volume 17, 1998), same evaluation framework (thousands of parameter-space-point coverage comparisons). Different problems, different sample structures, same Wilson-based Method 10 hybrid appears in all three (Stats/PT #46 methodological review flags it; today’s 1998b unpaired version applies it to two-sample unpaired data; PT #50’s 1998c paired version extends it to correlated pairs). Trilogy now CLOSED across three posts.
Queue status
Closes queue item “Newcombe 1998b unpaired-difference eleven methods (Stats in Med 17:873-890 — third of the trilogy)” flagged 2026-09-19 in Stats #50 and PT #50. Same-day pairing with today’s PT #51 Newcombe 1998b primary source (PubMed abstract verbatim; full PDF not publicly accessible today — see PT #51 for verification caveat). Queue rotates to suppressor-variable partial-vs-semi- partial (queued since Stats #25); higher-order VAR-b prewhitening (queued since Stats #20); heteroskedastic joint kurtosis (queued since Stats #33); coefficient of partial determination sr²/(1-r_yz²); Mee 1984 / Miettinen-Nurminen 1985 primary source of method 11 (Newcombe’s top-recommended coverage method but not closed-form). Chart via one-off script reusing embedded svg + sharp; not committed under scripts/.