Comparing Exact, Mid‑p, and Wilson Score Tests for a Single Binomial Proportion
Main Article Content
Abstract
Recent methodological studies have continued to refine and scrutinize exact binomial inference, providing comparisons between exact tests, approximate methods, and mid‑p corrections. This ongoing work reflects that even after many decades, the simple binomial problem continues to yield valuable theoretical insights. A consistent finding across studies is the inadequacy of methods based on the normal approximation. In particular, the exact binomial test, while exact in principle, is known to be highly conservative, with its actual significance level often substantially lower than the nominal level (actual significance leve l), implying a loss of statistical power and thus a reduced ability to detect a true effect when it exists. This study compares four methods for inference on a single binomial proportion: the standard exact binomial test (using the nominal significance level ), the exact binomial test (using the actual significance level), the mid-p test, and the widely recommended Wilson score test. Using Monte Carlo simulation with replications, we compare these tests in terms of empirical Type I error rate and statistical power across a comprehensive grid of sample sizes ( to ) and nine null success probabilities ( ). The evidence strongly indicates that the mid-p test outperforms the standard exact binomial test and the Wilson score test for the analysis of a single binomial proportion, particularly in small-to-moderate sample sizes. It achieves a Type I error rate close to the nominal level while consistently providing higher power. The use of an actual significance threshold, although a valuable theoretical exercise, offers no practical advantage over the simpler and more effective mid-p procedure. We therefore strongly recommend the adoption of the mid-p test in practice.
Article Details

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
You are free to:
- Share — copy and redistribute the material in any medium or format.
- Adapt — remix, transform, and build upon the material.
Under the following terms:
- Attribution — You must give appropriate credit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use.
- NonCommercial — You may not use the material for commercial purposes.
- No additional restrictions — You may not apply legal terms or technological measures that legally restrict others from doing anything the license permits.
References
[1] C. J. Clopper, & E. S. Pearson, "The use of confidence or fiducial limits illustrated in the case of the binomial", Biometrika, vol. 26, no. 4, pp. 404–413, 1934.
[2] A. Agresti & A. Gottard, " Independence in multi-way contingency tables: S. N. Roy's breakthroughs and later developments", Journal of Statistical Planning and Inference, vol. 137, pp. 3216–3226, 2007.
[3] C. R. Mehta & N. R. Patel, "A network algorithm for performing Fisher's exact test in r × c contingency tables", Journal of the American Statistical Association, vol. 78, no. 382, pp. 427–434, 1983.
[4] H. O. Lancaster, "Significance tests in discrete distributions," Journal of the American Statistical Association, vol. 56, no. 294, pp. 223–234, 1961.
[5] E. B. Wilson, "Probable inference, the law of succession, and statistical inference," Journal of the American Statistical Association, vol. 22, no. 158, pp. 209–212, 1927.
[6] M. W. Fagerland, S. Lydersen, & P. Laake, "Recommended tests for association in 2×2 tables," Statistics in Medicine, vol. 34, no. 14, pp. 2159–2175, 2015.
[7] J. T. Hwang & M. C. Yang, "An optimality theory for mid p values in 2 × 2 contingency tables," Statistica Sinica, vol. 11, pp. 807–826, 2001.
[8] A. Agresti & B. A. Coull, "Approximate is better than 'exact' for interval estimation of binomial proportions," The American Statistician, vol. 52, no. 2, pp. 119–126, 1998.