An AI system’s willingness to answer can depend on internal confidence signals, controlled experiments now suggest. The finding adds to efforts to understand model behaviour by looking inside neural representations rather than judging outputs alone, an approach also used in research that probes how language models internally represent plausible and impossible events.
Published in Nature Machine Intelligence on 7 September, the study examined how language models decide whether to answer questions or abstain. Researchers first measured confidence without offering an abstention option, then tested whether those estimates predicted decisions.
The strongest evidence came from an intervention inside Gemma 3 27B. Increasing internal confidence made abstention less likely; suppressing it had the opposite effect. This moves the finding beyond an association between sounding confident and giving an answer.
Models also adjusted their behaviour when instructed to answer only above specified confidence thresholds. That connects with broader attempts to build AI systems that explicitly acknowledge uncertainty rather than behave as unquestioned authorities. Meanwhile, verbal confidence ratings provided information about abstention beyond the probabilities attached to generated answers.
The results suggest that models use internal assessments alongside decision rules when deciding whether to respond. They do not establish consciousness, human understanding or dependable judgment in every setting.
For developers building autonomous assistants, the practical question is whether this mechanism can support better decisions about seeking clarification or human help. That distinction matters because fluent AI systems can still sound confident when they are wrong. Confidence influences action, but confident action can still be wrong.