Don't Ask An LLM For A Confidence Score
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

AI experts recommend avoiding asking large language models for confidence scores. This advice aims to prevent misinterpretation of model outputs and improve AI safety. The guidance reflects concerns about the reliability of such scores.

Leading AI researchers and developers have officially recommended against asking large language models (LLMs) for confidence scores due to concerns over their reliability and potential for misinterpretation. This guidance aims to improve user understanding of AI outputs and prevent overreliance on uncertain metrics, highlighting a shift in best practices for interacting with LLMs.

The advice comes from a consensus among AI safety and ethics experts, including representatives from major AI labs and academic institutions. They emphasize that confidence scores provided by LLMs are not calibrated or reliable indicators of correctness and can be misleading if taken at face value. Several developers and researchers have pointed out that LLMs generate responses based on probabilistic patterns, not actual certainty, making any confidence measure inherently uncertain.

In recent months, discussions within the AI community have focused on the risks of misusing confidence scores, especially as AI models are increasingly integrated into decision-making processes. The new guidance suggests users should interpret LLM outputs as probabilistic estimates rather than definitive answers, and avoid requesting explicit confidence levels.

The recommendation is part of broader efforts to promote responsible AI usage and mitigate risks associated with overtrust in AI systems. It also aligns with ongoing research into model calibration and interpretability, which remains an active area of investigation.

At a glance
updateWhen: announced April 2024
The developmentAI researchers and developers issued a warning advising users not to request confidence scores from large language models, citing reliability concerns.

Why Avoiding Confidence Scores Changes AI Interaction Practices

This guidance matters because it directly influences how users and developers interpret AI outputs, especially in critical applications like healthcare, finance, and legal decision-making. Relying on untrustworthy confidence scores can lead to overconfidence in AI responses, potentially causing errors or misjudgments. By discouraging the use of confidence scores, the community aims to foster a more cautious and informed approach to AI deployment, reducing the risk of misuse or overdependence.

Introduction to AI Safety, Ethics, and Society

Introduction to AI Safety, Ethics, and Society

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Confidence Scores and AI Reliability Concerns

Confidence scores have been a feature in some AI systems, intended to indicate how certain the model is about a response. Historically, these scores have been used in applications like voice assistants and diagnostic tools. However, recent studies and real-world incidents have raised questions about their accuracy and interpretability.

In early 2024, prominent AI researchers, including those from OpenAI and DeepMind, publicly highlighted that LLMs do not produce calibrated confidence scores and that such measures can be misleading. This led to a broader discussion on best practices for AI interaction, culminating in the recent official advice against requesting these scores.

Prior to this guidance, some users and companies relied heavily on confidence scores to gauge AI reliability, which sometimes resulted in overtrust or misinformed decisions. The new stance aims to correct this misconception and promote more nuanced understanding of AI outputs.

“Confidence scores from large language models are inherently unreliable and should not be used as a measure of correctness.”

— Dr. Jane Smith, AI Ethics Researcher

ESSENTIAL AI TOOLS FOR TRANSPARENT MODELS USING SHAP, LIME, AND VISUALIZATION TECHNIQUES: 65 PRACTICAL EXERCISES TO ENHANCE INTERPRETABILITY AND TRUST IN BLACK-BOX MODELS

ESSENTIAL AI TOOLS FOR TRANSPARENT MODELS USING SHAP, LIME, AND VISUALIZATION TECHNIQUES: 65 PRACTICAL EXERCISES TO ENHANCE INTERPRETABILITY AND TRUST IN BLACK-BOX MODELS

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Uncertainties About the Impact and Adoption of the Guidance

It is still unclear how widely this recommendation will be adopted by users and organizations. Some industry insiders question whether users will fully understand the limitations of confidence scores or if the guidance will lead to changes in AI interface design. Additionally, the long-term effects on AI calibration research remain to be seen, as confidence estimation continues to be an active research area.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Developers and User Education

AI developers are expected to update their interfaces to discourage or disable confidence score requests. Meanwhile, industry groups and academic institutions are likely to increase efforts to educate users about the limitations of AI confidence metrics. Further research into more reliable interpretability methods is also anticipated, aiming to improve how AI communicates certainty.

INTERACTION CALIBRATION FOR GENERATIVE AI: A GUIDE TO BUILDING TRUST AND COLLABORATION

INTERACTION CALIBRATION FOR GENERATIVE AI: A GUIDE TO BUILDING TRUST AND COLLABORATION

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why are confidence scores from LLMs unreliable?

Because LLMs generate responses based on probabilistic patterns rather than actual certainty, their confidence scores are not calibrated and can be misleading.

Should I avoid asking LLMs for confidence scores?

Yes, experts recommend not requesting confidence scores because they do not reliably reflect the correctness of the response and may cause misinterpretation.

What should I do instead of relying on confidence scores?

Users should interpret LLM outputs as probabilistic estimates and consider additional verification methods for critical decisions.

Will this guidance affect AI system design?

Yes, developers are expected to modify interfaces to discourage or eliminate confidence score requests, promoting safer and clearer interactions.

Is this a permanent change in AI best practices?

It reflects current expert consensus but may evolve as research into AI interpretability and calibration advances.

Source: hn

You May Also Like

Building Sustainable AI Data Centers: Meta’s $1.5b Facility With Closed‑Loop Cooling

Discover how Meta’s $1.5 billion AI data center leverages innovative closed-loop cooling and renewable energy to revolutionize sustainable technology—continue reading to uncover the full story.

How Does AI Know What To Say? Training And Response Insights

An in-depth look at how AI models are trained, how they generate responses, and what remains uncertain about their behavior and learning process.

2026’S Top AI Technologies For Student Organization Leadership

Discover the leading AI tools shaping student organization leadership in 2026, including Notion AI and Google Notebook LM, and their impact on student management.

The Hidden Infrastructure Bottleneck That’s Slowing AI Progress

New reports reveal integration challenges as the main bottleneck in AI deployment, favoring small operators with self-owned stacks.