Language models can answer familiar questions correctly while struggling to follow the same logic when the facts change. A study published in the KDD 2026 proceedings puts that weakness under closer scrutiny.
Researchers built a diagnostic benchmark using DBpedia, a knowledge graph that connects entities through structured relationships. Its questions require between one and five reasoning steps, allowing the team to examine how performance changes as the chain grows longer.
The tests compare ordinary factual questions with counterfactual versions whose premises contradict familiar knowledge but still support a logically consistent answer. A third category introduces similar entities among the answer options to probe confusion and dependence on prior knowledge.
Across seven language models, the researchers found strong reliance on learned knowledge. Performance deteriorated sharply on counterfactual tasks as reasoning depth increased.
The distinction matters because a correct answer alone cannot reveal whether a model followed the supplied evidence or recalled something it already knew.
The benchmark offers a controlled way to investigate that difference. Its findings concern the models and tasks tested; they do not establish that language models are universally incapable of reasoning.