Skip to main content

A Trap for AI Use in Peer Reviews Sparks Controversy

Conference organizers secretly inserted hidden prompts into papers to catch peer reviewers using artificial intelligence tools, stirring up criticism on social media.

Written byDalmeet Singh ChawlaBrought to you byThe Transmitter
| 3 min read
An AI agent-like robot looks through a computer screen with binary code behind it, representing hidden AI prompts.
Register for free to listen to this article
Listen with Speechify
0:00
3:00

Organizers of a prominent neuroscience conference are facing pushback on social media after adding hidden prompts to their papers to catch peer reviewers who are using generative artificial intelligence (AI) to referee papers.

The 40th Annual Conference on Neural Information Processing Systems (NeurIPS)—which is slated to take place in Sydney, Australia, in December 2026—bans peer reviewers from uploading papers they referee to AI chatbots, as the practice breaches confidentiality. Reviewers can still use AI chatbots for background research purposes, according to the policy outlined in the conference’s handbook.

To enforce the policy and catch illicit AI use in peer review, the event’s organizers have included deliberately concealed instructions for large language models (LLMs) in papers sent out for peer review.

The instructions tell an LLM to use telltale phrases—such as “This work addresses the central challenge” and “The claims of the paper”—in a peer-review report. Some researchers have already been caught trying to sneak secret messages into their papers in a bid to game AI tools into giving them favorable referee reports. Many publishers ban the use of AI in peer review.

Multiple researchers refereeing papers for NeurIPS have taken to social media to express their concerns about the indirect prompt injections inserted into papers.

Continue reading below...

Like this story? Sign up for FREE Newsletter updates:

Latest science news storiesTopic-tailored resources and eventsCustomized newsletter content
Subscribe

“Designing a trap that presumes bad faith corrodes the relationship the whole system depends on,” Sören Auer, a computer scientist at Leibniz University Hannover, wrote on LinkedIn. “You do not build a healthy reviewing culture by treating your reviewers as suspects.”

But others see merits in the approach. A similar prompt-injection effort has caught hundreds of reviewers misusing LLMs in submissions for next week’s 43rd International Conference on Machine Learning (ICML 2026) in Seoul, South Korea, according to Nihar Shah, a computer scientist at Carnegie Mellon University and scientific integrity chair of that conference.

In a statement to The Transmitter, the NeurIPS organizing committee says it can’t discuss injected prompts in detail “without eroding the effectiveness of this intervention.”

Auer told The Transmitter he was assigned eight NeurIPS papers to review. He says he sometimes converts PDF files into Microsoft Word documents when carrying out peer review, which renders some prompts visible.

Auer says he initially rejected the first paper he was reviewing because he thought the prompts had been inserted by the study’s authors. But he removed the flag after discovering hidden prompts in a second paper and seeing researchers discussing this issue on a Reddit thread.

You do not build a healthy reviewing culture by treating your reviewers as suspects.

—Sören Auer

It’s possible that more papers are being rejected because referees don’t know that prompts were inserted by conference organizers, he says. “I personally think it’s not good to prohibit the use of AI,” Auer adds. “We should rather, of course, have a discussion on how to use it.”

The NeurIPS committee has been replying directly to any reviewer who has noticed the hidden prompts, informing them not to penalize individual papers, according to the statement.

Like Auer, Sara Atito, an AI researcher at the University of Surrey, told The Transmitter she spotted the same prompt in all four papers she reviewed for NeurIPS. She says she also found it in the version of her own paper that NeurIPS organizers created before sending the paper out for peer review.

Atito calls hidden prompts a “poor mechanism,” arguing that it may filter out some problematic submissions but won’t solve the bigger problems with peer review. “We put too much blame on reviewers because they are the visible point of failure,” she says.

But Shah says hidden prompts are “viable and feasible.” Shah led a similar effort at ICML 2026 by injecting hidden prompts into all submitted papers.

By doing so, Shah says, he and his team identified hundreds of referees who were using AI when they weren’t supposed to, leading to their reviews being rejected. ICML 2026 desk-rejected just under 500 papers over violations of its LLM review policy—about 2 percent of the total number of submissions the conference received this year.

Researchers expressed “overwhelming support” for the strategy, says Shah, who adds that he shared the methodology with the NeurIPS team. “I have been working on conference peer review for several years, and I have hardly seen such strong support for anything,” he says. “People were really tired of reviewers copy-pasting AI-generated reviews without putting any effort.”

This article was first published at The Transmitter.

Add The Scientist as a preferred source on Google

Add The Scientist as a preferred Google source to see more of our trusted coverage.

Meet the Author

  • Image of Dalmeet Singh Chawla.

    Dalmeet Singh Chawla is a freelance science journalist based in London. His work has been featured in Nature, Science, Slate, Undark, The Economist, New Scientist and Pacific Standard, among other publications.

    View Full Profile

Related Topics

You might also be interested in...
Loading Next Article...
You might also be interested in...
Loading Next Article...
The Scientist Digest cover September 2026
September 2026

Multiplex Microscopy Becomes Easier with Encoded Antibodies

A new system that enables researchers to uniquely tag monoclonal antibodies for use in microscopy could help simplify complex imaging studies.

View this Issue
Rethinking ALS Biomarkers: From Discovery to Clinical Impact

Rethinking ALS Biomarkers: From Discovery to Clinical Impact

Alamar Biosciences logo
Best Practices for qPCR Assay Design and Optimization

Best Practices for qPCR Assay Design and Optimization

Bio-Rad
Beyond the Basics: Strategies for Single-Cell and Spatial Transcriptomics Analysis

Beyond the Basics: Strategies for Single-Cell and Spatial Transcriptomics Analysis

bioxcell
Scientist reviewing cellular and molecular data on a computer in a laboratory.

Building Translation-Ready Biomarkers with Connected Workflows

Danaher Logo

Products

Closeup image of a multi channel pipette dispensing pink liquid into a 96-well plate.

The ASSIST PLUS pipetting robot for affordable workflow automation

Integra Logo
Single cells in suspension

Rapidly isolate primary cells and make uniform single-cell suspensions with Corning® Cell Strainers

Corning logo
Abstract image representing cell membranes linked together.

CellBrite® Steady Membrane Stain: Cell surface staining built for real-time imaging

Biotium
sino biological logo

Monod Bio Licenses AI-designed Protein Technologies to SignalChem Biotech for Custom Discovery Assays