Artificial intelligence tools are increasingly being used not only to generate scientific research, but also to check the accuracy of existing work, exposing mistakes that have sometimes remained in the scientific record for decades.
One recent chemistry example came from Zhejiang Lab’s 021 science foundation model. During a boiling-point prediction exercise, the model produced results that conflicted with established values in scientific literature and databases. Researchers initially treated the discrepancy as a possible model error, but checking the underlying sources reportedly revealed a decades-old typographical mistake and measurements made with outdated instruments.
The episode illustrates a broader problem with scientific data: an incorrect value can become difficult to challenge once it has been copied from an original paper into reference books and then into digital databases. Researchers have increasingly turned to machine-learning systems to identify such inconsistencies, particularly when large datasets make manual checking difficult. Studies of chemical databases have also documented errors introduced both in original research and during the transfer of information into databases.
The same approach is now being tested directly on scientific papers.
A study by researchers including Federico Bianchi and James Zou developed a GPT-5-based system called the Paper Correctness Checker to identify objective errors in published artificial-intelligence research, including incorrect equations, calculations, figures and tables. The researchers examined 2,500 papers from ICLR, NeurIPS and TMLR and found an average of 4.7 such errors per paper.

The number increased over time at NeurIPS, from an average of 3.8 errors per paper in 2021 to 5.9 in 2025, a rise of 55.3%. Mathematical and formula-related mistakes accounted for 54% of the errors identified by the system.
Human experts examined 316 potential mistakes flagged by the AI and confirmed 263 of them, giving the system a precision rate of 83.2%. It also proposed corrections that researchers judged correct in 75.8% of cases. The researchers stressed that the system is intended to assist human reviewers rather than replace them.
Other research points to the limits of this approach. The SPOT benchmark tested AI systems on 83 published scientific papers containing 91 errors that had been confirmed by authors and were serious enough to result in corrections or retractions. The best-performing model detected only 21.1% of the errors, while its precision was 6.1%, showing how difficult reliable scientific verification remains for current models.
Google is also developing tools aimed at bringing AI into scientific validation. Its experimental Paper Assistant Tool is designed to analyse complete manuscripts, check theoretical results and experiments, and identify possible flaws. Google says it is testing such systems with major scientific conferences including ICML, STOC and NeurIPS.
The emerging picture is therefore more complicated than the familiar warning that AI can invent information. The same technology can also be useful for finding contradictions in human-generated research, particularly where an answer can be checked through calculations, equations or underlying data.
But the evidence does not support treating AI as the final authority on scientific truth. Current systems can miss serious errors and generate incorrect warnings of their own. Their most useful role is narrower: identifying anomalies and questions that scientists can then investigate against the original evidence.
No comments yet. Be the first to share your thoughts.