AI Models Inherit Flawed Research from Journals

AI chatbots’ misinformation problem extends beyond early-stage glitches, with models absorbing weak or ideologically driven research from a deteriorating scientific publishing ecosystem. In Canada, the assumption that publication equals peer review is increasingly at odds with reality, as journals become “slop-ified,” according to Vass Bednar, managing director of the Canadian SHIELD Institute. Large language models, designed to detect patterns rather than adjudicate truth, treat all text equally during pre-training, including fiction and highly peer-reviewed research, as explained by Thor Tronrud, senior machine learning scientist at StarFish Medical.
The Pipeline of Polluted Research
Predatory and pay-to-play journals have long proliferated, with studies packaged to mimic legitimate research. Despite this, assumptions about publication credibility persist in policy circles. Bednar notes that the scientific publishing ecosystem has been eroding for decades, yet the belief that “It’s published. That means it’s peer-reviewed” remains common. Large language models absorb vast amounts of text, reducing everything to probabilistic relationships, which means weak research gains undue weight in AI outputs.
Companies attempt to filter training data using heuristics like domain reputation and citation volume, originally designed for search engines. However, Tronrud asserts that the core challenge is the need for as much text as possible during training. This creates a vulnerability, particularly in Canada, where AI systems are becoming embedded in society. In the United States, politically aligned actors are constructing a parallel ecosystem of journals and think tanks that reinforce one another, appearing credible to models through automated authority signals like .gov domains.
Government-Backed Pseudoscience
Dr. Peter Hotez, Dean of the National School of Tropical Medicine at Baylor College of Medicine, warns that U.S. government officials are assembling “a whole alternative universe of pseudoscience” complete with journals and institutional trappings. Once ingested, this material loses authorship details, and retractions fail to “unbake” a model’s internal weights, dissolving the distinction between gold-standard science and gold-standard pseudoscience. This leads to junk science being laundered through AI summaries, which then propagate into policy memos and news articles, making it harder to challenge once widely cited.
The Sovereignty Challenge
Countries must decide whether to build domestic digital public infrastructure or risk inheriting distortions from foreign models trained on polluted data. Procurement standards and public-interest models could prioritize transparency in data curation, while neglecting this issue may cause nations to inherit distortions from foreign models.
