Back to Vol. 1 No. 1 (2025): Journal of Interdisclipinary Inquiry Table of Contents
Calibrated Uncertainty in Large Language Models
Uma Maheswara Rao Ulisi
Journal of Interdisclipinary Inquiry (2025) 1 (1).
Issue Section: Research
Keywords: large language models, epistemic calibration, uncertainty quantification, metacognition, AI alignment, Bayesian epistemology, mechanistic interpretability
Abstract
Large language models (LLMs) have demonstrated remarkable proficiency across reasoning-intensive tasks, yet systematic evidence indicates that their expressed confidence is poorly calibrated relative to their actual accuracy. This paper advances a theoretical and empirical framework for understanding epistemic self-awareness in LLMs, drawing on cognitive science, Bayesian epistemology, and mechanistic interpretability research. We argue that confidence miscalibration in LLMs reflects a deeper architectural tension between the distributional statistics of pretraining corpora and the demands of genuine uncertainty quantification. We introduce Generative Epistemic Calibration (GEC) and propose diagnostic protocols grounded in second-order probability theory. A four-part taxonomy of failure modes is developed, distinguishing overconfidence rooted in frequency-based heuristics from underconfidence arising from distributional ambiguity, reflective inconsistency, and prompted overconfidence. The paper concludes with implications for AI safety, human-AI collaboration, and training regimes that incentivize calibrated self-report, yielding testable predictions relevant to clinical decision support, legal reasoning, and scientific hypothesis generation.
Uma Maheswara Rao Ulisi
Published 2026-07-01
Ulisi, U. M. rao . (2026). Calibrated Uncertainty in Large Language Models. The Journal of Interdisciplinary Inquiry, 1(1). The Journal of Interdisciplinary Inquiry. https://doi.org/10.5281/zenodo.21583437
© Copyright 2026 JII.
This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors retain copyright and agree to license their articles with a Creative Commons Attribution (CC BY 4.0) International License.
Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1-3.
Burns, C., Ye, H., Klein, D., & Steinhardt, J. (2022). Discovering latent knowledge in language models without supervision. arXiv preprint arXiv:2212.03827.
Elazar, Y., Kassner, N., Ravfogel, S., Ravichander, A., Hovy, E., Schutze, H., & Goldberg, Y. (2021). Measuring and improving consistency in pretrained language models. Transactions of the Association for Computational Linguistics, 9, 1012-1031.
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., & Olah, C. (2021). A mathematical framework for transformer circuits. Transformer Circuits Thread. https://transformer-circuits.pub/2021/framework/index.html
Flavell, J. H. (1979). Metacognition and cognitive monitoring: A new area of cognitive-developmental inquiry. American Psychologist, 34(10), 906-911.
Fleming, S. M., & Dolan, R. J. (2012). The neural basis of metacognitive ability. Philosophical Transactions of the Royal Society B, 367(1594), 1338-1349.
Gneiting, T., & Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477), 359-378.
Good, I. J. (1962). Subjective probability as the measure of a non-measurable set. In E. Nagel, P. Suppes, & A. Tarski (Eds.), Logic, Methodology and Philosophy of Science (pp. 319-329). Stanford University Press.
Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning (pp. 1321-1330). PMLR.
Jang, M., Ye, S., Yang, S., Shin, J., Han, J., Kim, G., & Seo, M. (2022). Becel: Benchmark for consistency evaluation of language models. In Proceedings of the 29th International Conference on Computational Linguistics (pp. 3818-3834).
Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., & Kaplan, J. (2022). Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221.
Kruger, J., & Dunning, D. (1999). Unskilled and unaware of it: How difficulties in recognizing one’s own incompetence lead to inflated self-assessments. Journal of Personality and Social Psychology, 77(6), 1121-1134.
Nelson, T. O., & Narens, L. (1990). Metamemory: A theoretical framework and new findings. Psychology of Learning and Motivation, 26, 125-173.
Niculescu-Mizil, A., & Caruana, R. (2005). Predicting good probabilities with supervised learning. In Proceedings of the 22nd International Conference on Machine Learning (pp. 625-632).
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., & Carter, S. (2020). Zoom in: An introduction to circuits. Distill, 5(3), e00024-001.
Skyrms, B. (1980). Higher order degrees of belief. In D. H. Mellor (Ed.), Prospects for Pragmatism (pp. 109-137). Cambridge University Press.
Weiser, B., & Schweber, N. (2023, June 22). The ChatGPT lawyer explains himself. The New York Times.
Williamson, J. (2010). In Defence of Objective Bayesianism. Oxford University Press.
No related articles available.
The Journal of Interdisciplinary Inquiry promotes the examination of complex questions through diverse methodological and theoretical lenses, fostering connections across disciplines.