Epistemological scope and limitations of contemporary statistical inference: a systematic review

Authors

DOI:

https://doi.org/10.36097/rsan.v1iEspecial_3.4198

Keywords:

Statistical inference, epistemology, statistical significance, causal inference

Abstract

Statistical inference is fundamental to producing scientific knowledge, interpreting evidence, and guiding decision-making under uncertainty. However, its validity depends on study design, underlying assumptions, analytical transparency, and the appropriate interpretation of results. The objective of this study was to qualitatively synthesize the epistemological scope and limitations attributed to statistical inference in recent scientific literature. A qualitative systematic review guided by PRISMA 2020 was conducted. The PubMed and Europe PMC databases were searched to identify publications issued between 2021 and 2025. Thirteen studies were included and analyzed through a thematic, comparative, and interpretive synthesis. The results were organized into five themes: reinterpretation of the p-value and intervals as measures of compatibility; coexistence of frequentist and Bayesian approaches; reproducibility and quality of statistical reporting; uncertainty quantification in artificial intelligence; and the relationship between causal inference and machine learning. The reviewed literature questions dichotomous decisions based exclusively on statistical significance and emphasizes the need to consider effect sizes, assumptions, context, and practical relevance. It also indicates that inferential validity is shaped by institutional practices, transparency, and study design quality. The study concludes that statistical inference remains essential to scientific reasoning but does not provide definitive certainty. Its rigorous use requires making assumptions explicit, calibrating uncertainty, ensuring analytical transparency, and clearly distinguishing among association, prediction, and causality.

Downloads

Download data is not yet available.

References

Allen, G. I., Gan, L., & Zheng, L. (2024). Interpretable machine learning for discovery: Statistical challenges and opportunities. Annual Review of Statistics and Its Application, 11, 97-121. https://doi.org/10.1146/annurev-statistics-040120-030919

Altman, D. G. (1994). The scandal of poor medical research. The BMJ, 308(6924), 283-284. https://doi.org/10.1136/bmj.308.6924.283

Amrhein, V., Greenland, S., & McShane, B. (2019). Scientists rise up against statistical significance. Nature, 567(7748), 305-307. https://doi.org/10.1038/d41586-019-00857-9

Benjamin, D. J., Berger, J. O., Johannesson, M., Nosek, B. A., Wagenmakers, E.-J., Berk, R., Bollen, K. A., Brembs, B., Brown, L., Camerer, C., Cesarini, D., Chambers, C. D., Clyde, M., Cook, T. D., De Boeck, P., Dienes, Z., Dreber, A., Easwaran, K., Efferson, C., ... Johnson, V. E. (2018). Redefine statistical significance. Nature Human Behaviour, 2(1), 6-10. https://doi.org/10.1038/s41562-017-0189-z

Boscardin, C. K., Sewell, J. L., Tolsgaard, M. G., & Pusic, M. V. (2024). How to use and report on p-values. Perspectives on Medical Education, 13(1), 250-254. https://doi.org/10.5334/pme.1324

Brand, J. E., Zhou, X., & Xie, Y. (2023). Recent developments in causal inference and machine learning. Annual Review of Sociology, 49, 81-110. https://doi.org/10.1146/annurev-soc-030420-015345

Button, K. S., Ioannidis, J. P. A., Mokrysz, C., Nosek, B. A., Flint, J., Robinson, E. S. J., & Munafo, M. R. (2013). Power failure: Why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience, 14(5), 365-376. https://doi.org/10.1038/nrn3475

Chakraborti, T., Banerji, C. R. S., Marandon, A., Hellon, V., Mitra, R., Lehmann, B., Braeuninger, L., McGough, S., Turkay, C., Frangi, A. F., Bianconi, G., Li, W., Rackham, O., Harbron, C., Parashar, D., & MacArthur, B. D. (2025). Personalized uncertainty quantification in artificial intelligence. Nature Machine Intelligence, 7, 522-530. https://doi.org/10.1038/s42256-025-01024-8

Cobey, K. D., Ebrahimzadeh, S., Page, M. J., Thibault, R. T., Nguyen, P.-Y., Abu-Dalfa, F., et al. (2024). Biomedical researchers' perspectives on the reproducibility of research. PLOS Biology, 22(11), e3002870. https://doi.org/10.1371/journal.pbio.3002870

Fornacon-Wood, I., Mistry, H., Johnson-Hart, C., Faivre-Finn, C., O'Connor, J. P. B., & Price, G. J. (2022). Understanding the differences between Bayesian and frequentist statistics. International Journal of Radiation Oncology, Biology, Physics, 112(5), 1076-1082. https://doi.org/10.1016/j.ijrobp.2021.12.011

Gelman, A., & Loken, E. (2014). The statistical crisis in science. American Scientist, 102(6), 460-465. https://doi.org/10.1511/2014.111.460

Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., & Rubin, D. B. (2013). Bayesian data analysis (3rd ed.). CRC Press.

Goligher, E. C., Heath, A., & Harhay, M. O. (2024). Bayesian statistics for clinical research. The Lancet, 404(10457), 1067-1076. https://doi.org/10.1016/S0140-6736(24)01295-9

Goodman, S. N. (1999a). Toward evidence-based medical statistics. 1: The P value fallacy. Annals of Internal Medicine, 130(12), 995-1004. https://doi.org/10.7326/0003-4819-130-12-199906150-00008

Goodman, S. N. (1999b). Toward evidence-based medical statistics. 2: The Bayes factor. Annals of Internal Medicine, 130(12), 1005-1013. https://doi.org/10.7326/0003-4819-130-12-199906150-00019

Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337-350. https://doi.org/10.1007/s10654-016-0149-3

Homa-Bonell, J. K. (2023). Interpreting P values in 2023. Journal of Patient-Centered Research and Reviews, 10(3), 102-103. https://doi.org/10.17294/2330-0698.2064

Ioannidis, J. P. A. (2005). Why most published research findings are false. PLOS Medicine, 2(8), e124. https://doi.org/10.1371/journal.pmed.0020124

Mansournia, M. A., & Nazemipour, M. (2024). Recommendations for accurate reporting in medical research statistics. The Lancet, 403(10427), 611-612. https://doi.org/10.1016/S0140-6736(24)00139-9

Mansournia, M. A., Nazemipour, M., & Etminan, M. (2023). P-value, compatibility, and S-value. Global Epidemiology, 5, 100085. https://doi.org/10.1016/j.gloepi.2022.100085

McShane, B. B., Gal, D., Gelman, A., Robert, C., & Tackett, J. L. (2019). Abandon statistical significance. The American Statistician, 73(sup1), 235-245. https://doi.org/10.1080/00031305.2018.1527253

Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600-2606. https://doi.org/10.1073/pnas.1708274114

Oliveira, R. I., Orenstein, P., Ramos, T., & Romano, J. V. (2024). Split conformal prediction and non-exchangeable data. Journal of Machine Learning Research, 25(225), 1-38. https://www.jmlr.org/papers/v25/23-1553.html

Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., et al. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, n71. https://doi.org/10.1136/bmj.n71

Pearl, J. (2009). Causality: Models, reasoning, and inference (2nd ed.). Cambridge University Press.

Rafi, Z., & Greenland, S. (2020). Semantic and cognitive tools to aid statistical science: Replace confidence and significance by compatibility and surprise. BMC Medical Research Methodology, 20, 244. https://doi.org/10.1186/s12874-020-01105-9

Sun, E. D., Ma, R., Navarro Negredo, P., Brunet, A., & Zou, J. (2024). TISSUE: Uncertainty-calibrated prediction of single-cell spatial transcriptomics improves downstream analyses. Nature Methods, 21, 444-454. https://doi.org/10.1038/s41592-024-02184-y

Tejada-Lapuerta, A., Bertin, P., Bauer, S., Aliee, H., Bengio, Y., & Theis, F. J. (2025). Causal machine learning for single-cell genomics. Nature Genetics, 57(4), 797-808. https://doi.org/10.1038/s41588-025-02124-2

Wasserstein, R. L., & Lazar, N. A. (2016). The ASA's statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129-133. https://doi.org/10.1080/00031305.2016.1154108

Wasserstein, R. L., Schirm, A. L., & Lazar, N. A. (2019). Moving to a world beyond p < .05. The American Statistician, 73(sup1), 1-19. https://doi.org/10.1080/00031305.2019.1583913

Downloads

Published

2026-08-31

Issue

Section

ARTÍCULOS DE REVISIÓN

How to Cite

Villavicencio Cedeño, L. B., Batista Garcet, Y. ., Tenelanda Cudco , S. E., Guanga Chunata , D. M., & Casaliglla Ger, . D. (2026). Epistemological scope and limitations of contemporary statistical inference: a systematic review. Revista San Gregorio, 1(Especial_3), 131-141. https://doi.org/10.36097/rsan.v1iEspecial_3.4198