Paper Trail #44: An Invariant Form for the Prior Probability in Estimation Problems (Jeffreys, 1946) — Proceedings of the Royal Society A 186:453-461. FULL PRIMARY SOURCE VERIFIED via WebFetch + pdfminer.six on 2026-09-13 from a Wayback Machine capture of the Royal Society Publishing PDF (9 pages, 1.59 MB, 23,168 characters). 34th PT of 44 with full primary-source access. Derived the SQUARE-ROOT-OF-FISHER-INFORMATION prior that today's Stats #44 uses on the binomial: Section 2 eqs (19)-(21) verbatim show that reparametrizing p = sin²α with uniform prior on α in [0, π/2] recovers exactly the Beta(0.5, 0.5) density. Directly closes the four-paper small-sample-inference arc PT #40/#41/#42/#43 with the Bayesian sibling BCD 2001 recommends.
Harold Jeffreys (1946), “An Invariant Form for the Prior Probability in Estimation Problems.” Proc. Royal Society A 186:453-461. Full primary source verified via WebFetch + pdfminer.six on 2026-09-13 from a Wayback Machine capture of the Royal Society Publishing PDF (9 pages, 1.59 MB, 23,168 characters extracted). 34th PT of 44 with full primary-source access. Derived the SQUARE-ROOT-OF-FISHER-INFORMATION prior that today’s Stats #44 uses on the binomial: Section 2 eqs (19)-(21) verbatim show that reparametrizing p = sin²(alpha) with a uniform prior on alpha in [0, pi/2] recovers exactly the Beta(0.5, 0.5) density. Directly closes the four-paper small-sample- inference arc PT #40/#41/#42/#43 with the Bayesian sibling BCD 2001 recommends alongside Wilson for n ≤ 40.
![Three prior densities on the unit interval for a binomial proportion p: Jeffreys Beta(0.5, 0.5) shown in orange with a U-shape with poles at 0 and 1 (density 3.20 at p=0.01, 0.64 at p=0.5), Bayes-Laplace uniform Beta(1, 1) shown as a coral horizontal line at density 1.0, and Haldane improper dp/[p(1-p)] shown as a light-grey dashed line that goes to infinity at both endpoints (density 100 at p=0.01, 4 at p=0.5, 100 at p=0.99). Jeffreys is intermediate between uniform and Haldane.](/insights/paper-trail-jeffreys-1946/prior-comparison.png)
The invariance derivation — Section 2 eq. (12)
Jeffreys 1946 abstract (p. 453) verbatim: “It is shown that a certain differential form depending on the values of the parameters in a law of chance is invariant for all transformations of the parameters when the law is differentiable with regard to all parameters ... This form has the properties required to give a general rule for stating the prior probability in a large class of estimation problems.” Section 2 equation (12) verbatim: “if the prior probability density of the parameters is taken as proportional to ||g_ik||^(1/2), the prior probability over any region will be invariant for all ways of choosing the parameters.” g_ik is defined in eq. (7) as the Fisher information matrix: g_ik = lim(δx_r → 0) Σ (1/p_r)(∂p_r/∂α_i)(∂p_r/∂α_k). The prior is sqrt(det(I(theta)))— later canonized as “the Jeffreys prior.”
The binomial special case — eqs (19)-(21)
Section 2, p. 457, eqs (19)-(21) verbatim: “Next consider simple sampling. Denote the chance of a success by sin²(alpha), that of a failure by cos²(alpha). Then for variations of alpha I_1 = (sin alpha − sin alpha’)² + (cos alpha − cos alpha’)² = 4 sin²(½(alpha’−alpha)), I_2 = (sin²alpha’ − sin²alpha) log[sin²alpha’ / sin²alpha] = 4(alpha’−alpha)². Then the rule gives, since 0 ≤ alpha ≤ ½pi, P(dalpha | H) = dalpha/(½pi). According to the Bayes-Laplace rule the prior probability of p should be taken uniform. Haldane has suggested dp/[p(1−p)]... The present rule is intermediate.” Reparametrizing p = sin²(alpha) with uniform prior on alpha in [0, pi/2] recovers a prior on p whose density is proportional to |d(alpha)/dp| = 1/(2 * sqrt(p(1-p))) — the Beta(0.5, 0.5) density up to a factor of 2.
Direct pairing with today’s Stats #44
Prior Beta(0.5, 0.5) × Binomial(n, p) likelihood → posterior Beta(k + 0.5, n - k + 0.5) by conjugacy. Today’s Stats #44 uses this to compute equal-tailed 95% credible intervals for the past-week samples: on today slot 1 GBPNZD 9/11 up, the Jeffreys interval is [53.28%, 96.02%] — the tightest lower bound of the four methods. BCD 2001 (yesterday’s PT #43) Section 4 recommended Wilson OR Jeffreys as the two small-n methods of choice; today’s PT #44 closes that thread with Jeffreys’s original 1946 derivation.
Small-sample-inference arc — five papers mapped
| Post | Paper | Year | Contribution |
|---|---|---|---|
| PT #40 | Hanley & Lippman-Hand JAMA | 1983 | Named the 'rule of three' — 3/n upper CI for k=0 |
| PT #41 | Wilson JASA | 1927 | Score-test-inversion interval (60 years pre-Wilson-name) |
| PT #42 | Agresti & Coull AmStat | 1998 | 'Add 2 successes and 2 failures' pedagogic simplification |
| PT #43 | Brown, Cai, DasGupta Stat.Sci | 2001 | Evaluation — recommend Wilson OR Jeffreys for n ≤ 40 |
| PT #44 | Jeffreys ProcRoySocA | 1946 | Invariance derivation of Beta(0.5, 0.5) prior |
Arc closed.BCD 2001 (PT #43) is the evaluation paper — it read the small-sample landscape and recommended Wilson (PT #41) or Jeffreys (today PT #44) for n ≤ 40. Today’s post gives Jeffreys its due primary-source citation.
Provenance and access
Verified working source on 2026-09-13: https://web.archive.org/web/20260906062915/ https://royalsocietypublishing.org/doi/pdf/10.1098/rspa.1946.0056 — Wayback Machine capture from 2026-09-06 of the Royal Society Publishing PDF. 1,590,772 bytes, 9 pages (pp. 453-461), 23,168 characters extracted via pdfminer.six. The live royalsocietypublishing.org PDF returned 403 to WebFetch (pre-2000 articles gated behind JSTOR / institutional access). The paper is also reprinted in H. Jeffreys, “Theory of Probability” (3rd ed., 1961), Oxford University Press, Section 3.10 pp. 179-181.
Not verified from this fetch
Pages 458-461 of the paper (sections beyond the binomial case — the location-scale extension, the Cauchy example, and the discussion of “restricted invariance” for laws not everywhere differentiable) exist in the fetched PDF but are not cross-checked line-by-line in this post. The four-item verbatim scope above covers the direct linkage to today’s Stats #44 Jeffreys interval; anything else claimed about the paper should be re-verified against the primary source before quoting. Jeffreys does NOT use the words “arc-sine” or sin⁻¹√panywhere in the paper — the “arcsine” naming for the Beta(0.5, 0.5) distribution is a later convention.