Between 2012 and 2022, the share of scientific papers involving AI quadrupled across twenty disciplines, from economics to geology. That explosion has brought genuine advances, but two of the field's most careful observers, Princeton computer scientists Arvind Narayanan and Sayash Kapoor, argue in a Nature commentary that the scientific workforce is nowhere near prepared to police the errors these tools introduce. Their case rests on a straightforward problem: AI models are largely opaque, and the standard mechanisms of scientific quality control were never built to look inside them.
The errors are not exotic. Data leakage, where a model encounters information during training that would not be available in a real-world prediction, remains one of the most common and least detected failures in AI-assisted research [1]. Overfitting and model mis-specification follow close behind. The trouble, as Narayanan and Kapoor see it, is not that these flaws are new but that the sheer accessibility of modern AI tools means they are now being committed at scale, by researchers who may have limited understanding of how the underlying models behave. A 2026 review of AI challenges in research and publishing noted that black-box algorithms and undisclosed training data make replication of AI-driven studies significantly harder, compounding a reproducibility problem that predated the current generation of tools.
The evidence from life sciences is particularly stark. A BioSkepsis analysis found that roughly 75 percent of code released alongside published papers could not run without errors, and that GPT-4o fabricated nearly 20 percent of biomedical citations in controlled tests worse, a majority of those fabricated references carried valid DOIs pointing to entirely unrelated papers, creating what the authors described as "a false trail of evidence" that could survive casual verification. A Patterns journal study published in 2026 echoed the concern, documenting how generative AI introduces new categories of risk into established machine-learning workflows, often in ways that standard peer review is not equipped to flag.
Narayanan and Kapoor stop short of calling for restraint. Their argument is more specific: without field-level protocols for validating AI-derived results (clear standards for data auditing, model transparency, and error detection), the volume of AI-assisted publication will outpace the community's ability to separate sound findings from artefacts. The question is no longer whether AI belongs in the laboratory. It does. The question is whether the habits and institutions surrounding its use have caught up with the speed of its adoption.