Epistemological scope and limitations of contemporary statistical inference: a systematic review
DOI:
https://doi.org/10.36097/rsan.v1iEspecial_3.4198Keywords:
Statistical inference, epistemology, statistical significance, causal inferenceAbstract
Statistical inference is fundamental to producing scientific knowledge, interpreting evidence, and guiding decision-making under uncertainty. However, its validity depends on study design, underlying assumptions, analytical transparency, and the appropriate interpretation of results. The objective of this study was to qualitatively synthesize the epistemological scope and limitations attributed to statistical inference in recent scientific literature. A qualitative systematic review guided by PRISMA 2020 was conducted. The PubMed and Europe PMC databases were searched to identify publications issued between 2021 and 2025. Thirteen studies were included and analyzed through a thematic, comparative, and interpretive synthesis. The results were organized into five themes: reinterpretation of the p-value and intervals as measures of compatibility; coexistence of frequentist and Bayesian approaches; reproducibility and quality of statistical reporting; uncertainty quantification in artificial intelligence; and the relationship between causal inference and machine learning. The reviewed literature questions dichotomous decisions based exclusively on statistical significance and emphasizes the need to consider effect sizes, assumptions, context, and practical relevance. It also indicates that inferential validity is shaped by institutional practices, transparency, and study design quality. The study concludes that statistical inference remains essential to scientific reasoning but does not provide definitive certainty. Its rigorous use requires making assumptions explicit, calibrating uncertainty, ensuring analytical transparency, and clearly distinguishing among association, prediction, and causality.
Downloads
References
Allen, G. I., Gan, L., & Zheng, L. (2024). Interpretable machine learning for discovery: Statistical challenges and opportunities. Annual Review of Statistics and Its Application, 11, 97-121. https://doi.org/10.1146/annurev-statistics-040120-030919
Altman, D. G. (1994). The scandal of poor medical research. The BMJ, 308(6924), 283-284. https://doi.org/10.1136/bmj.308.6924.283
Amrhein, V., Greenland, S., & McShane, B. (2019). Scientists rise up against statistical significance. Nature, 567(7748), 305-307. https://doi.org/10.1038/d41586-019-00857-9
Benjamin, D. J., Berger, J. O., Johannesson, M., Nosek, B. A., Wagenmakers, E.-J., Berk, R., Bollen, K. A., Brembs, B., Brown, L., Camerer, C., Cesarini, D., Chambers, C. D., Clyde, M., Cook, T. D., De Boeck, P., Dienes, Z., Dreber, A., Easwaran, K., Efferson, C., ... Johnson, V. E. (2018). Redefine statistical significance. Nature Human Behaviour, 2(1), 6-10. https://doi.org/10.1038/s41562-017-0189-z
Boscardin, C. K., Sewell, J. L., Tolsgaard, M. G., & Pusic, M. V. (2024). How to use and report on p-values. Perspectives on Medical Education, 13(1), 250-254. https://doi.org/10.5334/pme.1324
Brand, J. E., Zhou, X., & Xie, Y. (2023). Recent developments in causal inference and machine learning. Annual Review of Sociology, 49, 81-110. https://doi.org/10.1146/annurev-soc-030420-015345
Button, K. S., Ioannidis, J. P. A., Mokrysz, C., Nosek, B. A., Flint, J., Robinson, E. S. J., & Munafo, M. R. (2013). Power failure: Why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience, 14(5), 365-376. https://doi.org/10.1038/nrn3475
Chakraborti, T., Banerji, C. R. S., Marandon, A., Hellon, V., Mitra, R., Lehmann, B., Braeuninger, L., McGough, S., Turkay, C., Frangi, A. F., Bianconi, G., Li, W., Rackham, O., Harbron, C., Parashar, D., & MacArthur, B. D. (2025). Personalized uncertainty quantification in artificial intelligence. Nature Machine Intelligence, 7, 522-530. https://doi.org/10.1038/s42256-025-01024-8
Cobey, K. D., Ebrahimzadeh, S., Page, M. J., Thibault, R. T., Nguyen, P.-Y., Abu-Dalfa, F., et al. (2024). Biomedical researchers' perspectives on the reproducibility of research. PLOS Biology, 22(11), e3002870. https://doi.org/10.1371/journal.pbio.3002870
Fornacon-Wood, I., Mistry, H., Johnson-Hart, C., Faivre-Finn, C., O'Connor, J. P. B., & Price, G. J. (2022). Understanding the differences between Bayesian and frequentist statistics. International Journal of Radiation Oncology, Biology, Physics, 112(5), 1076-1082. https://doi.org/10.1016/j.ijrobp.2021.12.011
Gelman, A., & Loken, E. (2014). The statistical crisis in science. American Scientist, 102(6), 460-465. https://doi.org/10.1511/2014.111.460
Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., & Rubin, D. B. (2013). Bayesian data analysis (3rd ed.). CRC Press.
Goligher, E. C., Heath, A., & Harhay, M. O. (2024). Bayesian statistics for clinical research. The Lancet, 404(10457), 1067-1076. https://doi.org/10.1016/S0140-6736(24)01295-9
Goodman, S. N. (1999a). Toward evidence-based medical statistics. 1: The P value fallacy. Annals of Internal Medicine, 130(12), 995-1004. https://doi.org/10.7326/0003-4819-130-12-199906150-00008
Goodman, S. N. (1999b). Toward evidence-based medical statistics. 2: The Bayes factor. Annals of Internal Medicine, 130(12), 1005-1013. https://doi.org/10.7326/0003-4819-130-12-199906150-00019
Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337-350. https://doi.org/10.1007/s10654-016-0149-3
Homa-Bonell, J. K. (2023). Interpreting P values in 2023. Journal of Patient-Centered Research and Reviews, 10(3), 102-103. https://doi.org/10.17294/2330-0698.2064
Ioannidis, J. P. A. (2005). Why most published research findings are false. PLOS Medicine, 2(8), e124. https://doi.org/10.1371/journal.pmed.0020124
Mansournia, M. A., & Nazemipour, M. (2024). Recommendations for accurate reporting in medical research statistics. The Lancet, 403(10427), 611-612. https://doi.org/10.1016/S0140-6736(24)00139-9
Mansournia, M. A., Nazemipour, M., & Etminan, M. (2023). P-value, compatibility, and S-value. Global Epidemiology, 5, 100085. https://doi.org/10.1016/j.gloepi.2022.100085
McShane, B. B., Gal, D., Gelman, A., Robert, C., & Tackett, J. L. (2019). Abandon statistical significance. The American Statistician, 73(sup1), 235-245. https://doi.org/10.1080/00031305.2018.1527253
Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600-2606. https://doi.org/10.1073/pnas.1708274114
Oliveira, R. I., Orenstein, P., Ramos, T., & Romano, J. V. (2024). Split conformal prediction and non-exchangeable data. Journal of Machine Learning Research, 25(225), 1-38. https://www.jmlr.org/papers/v25/23-1553.html
Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., et al. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, n71. https://doi.org/10.1136/bmj.n71
Pearl, J. (2009). Causality: Models, reasoning, and inference (2nd ed.). Cambridge University Press.
Rafi, Z., & Greenland, S. (2020). Semantic and cognitive tools to aid statistical science: Replace confidence and significance by compatibility and surprise. BMC Medical Research Methodology, 20, 244. https://doi.org/10.1186/s12874-020-01105-9
Sun, E. D., Ma, R., Navarro Negredo, P., Brunet, A., & Zou, J. (2024). TISSUE: Uncertainty-calibrated prediction of single-cell spatial transcriptomics improves downstream analyses. Nature Methods, 21, 444-454. https://doi.org/10.1038/s41592-024-02184-y
Tejada-Lapuerta, A., Bertin, P., Bauer, S., Aliee, H., Bengio, Y., & Theis, F. J. (2025). Causal machine learning for single-cell genomics. Nature Genetics, 57(4), 797-808. https://doi.org/10.1038/s41588-025-02124-2
Wasserstein, R. L., & Lazar, N. A. (2016). The ASA's statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129-133. https://doi.org/10.1080/00031305.2016.1154108
Wasserstein, R. L., Schirm, A. L., & Lazar, N. A. (2019). Moving to a world beyond p < .05. The American Statistician, 73(sup1), 1-19. https://doi.org/10.1080/00031305.2019.1583913
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Lissette Beatriz Villavicencio Cedeño, Yoiler Batista Garcet, Sandra Elizabeth Tenelanda Cudco , Deysi Margoth Guanga Chunata , Darwin Casaliglla Ger

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.









