An AI system built to automate hydrological modelling completed 92.8 percent of 125 tasks, but its performance collapsed when researchers removed the expert workflow guiding its decisions.
The findings, published in the Journal of Hydrology, reveal both the potential of scientific AI agents and their dependence on structured planning.
Called HydroAIM, the system connects language models to an algorithm library, tools and a feedback process that checks execution. Multiple agents coordinate through the Model Context Protocol, following prescribed steps for the modelling work.
Researchers tested the approach across five language models. In a separate evaluation covering 531 river catchments from the CAMELS dataset, it ran without human intervention and achieved median Nash-Sutcliffe efficiency scores of 0.58 for local modelling and 0.73 for global modelling.
Those scores assess modelling performance; the 92.8 percent figure measures successful task execution.
The starkest result came from removing the expert task workflow: the execution success rate fell to zero. The authors identify planning over extended sequences as a central obstacle.
HydroAIM offers evidence that domain expertise, constrained tools and execution feedback can make language models useful for quantitative scientific work.