The rules for AI in research are shifting on several fronts at once. Within a single week, OpenAI laid out plans to embed invisible watermarks in AI text to meet the EU AI Act, a new study found AI-written passages in nearly 30% of recent US STEM PhD dissertations, a UK funding network apologized after an AI tool rejected almost half of its grant applications before any human read them, and arXiv capped submissions after a record month. The common thread is a hard question: how should academia detect AI-written work, and how much judgment should it hand to machines?
OpenAI’s text watermarks struggle with math
On October 6, 2026, OpenAI published its approach to text provenance under the EU AI Act. Provenance here means a record of where a piece of text came from. API customers get access to a detector for an invisible watermark, and consumer tools such as ChatGPT will start embedding the signal for users in Europe.
The watermark works by adding a slight statistical bias to how the model picks tokens, the small chunks of text a language model reads and writes. People cannot see the bias, but a detector can measure it. OpenAI’s own benchmarks show how fragile that signal is:
| Test condition | Detection rate |
|---|---|
| Psychology passage, 400 tokens | 94.3% |
| Mathematics passage, 200 tokens | 36.5% |
| 10% of words swapped for synonyms | Fell from 92% to 66% |
| 25% of words swapped for synonyms | Fell to 17% |
Short passages and formula-heavy writing make detection much weaker, and ordinary editing erodes it further. OpenAI also stated plainly what a watermark cannot do: it cannot measure how much a human contributed, confirm whether the content is accurate, or settle legal authorship and responsibility. It only signals that a model was involved.
That limit matters because a flag can easily be misread. A student who used AI to polish a draft and one who outsourced the research itself can produce the same signal. If institutions treat a watermark hit as proof of misconduct, honest students may be punished for using AI as a drafting aid.
Nearly 30% of US STEM dissertations show AI writing
An October 2026 working paper from the National Bureau of Economic Research (NBER) measured how often AI-written text appears in US STEM PhD dissertations. Before 2023, the researchers detected none. After that, the share climbed quickly:
| Period | Dissertations with detectable AI writing |
|---|---|
| 2023 | 1.7% |
| 2024 | 5.5% |
| 2025 | 18.8% |
| Through May 2026 | 29.4% |
Computer science and mathematics led at 36%, with engineering close behind at 35%. Use was substantially higher among students from non-English-speaking backgrounds and in lower-ranked doctoral programs. For international students writing dense technical prose in a second language, leaning on a language model for phrasing appears to be an understandable choice.
The career finding stands out. Graduates whose dissertations contained AI-generated writing were significantly more likely to go into industry than to stay on an academic tenure track. That does not prove AI use changed anyone’s path. Students already headed for industry may simply care less about traditional dissertation norms.

▲ Rising AI writing in PhD dissertations
An AI screener rejected half the grant proposals first
A cybersecurity research network funded by UK Research and Innovation (UKRI) ran its funding call through an automated AI triage tool. The tool scored 179 applications against seven criteria and rejected nearly half of them before any human reviewer looked at them.
Researchers who had spent weeks refining complex proposals pushed back hard. The network apologized and promised to reassess every affected application.
The episode shows where the line for AI in research evaluation likely sits. AI handles surface checks well: formatting, spelling, grammar, phrasing and cross-references. Judging a novel hypothesis or a niche research method requires deep domain understanding that current models do not appear to have. Letting software make funding decisions without expert review risks wasting the effort researchers put into their proposals and devaluing human expertise.
A reviewer declined a paper over an AI detector score
Peer review is facing the same tension. On October 5, 2026, an adjunct professor at Western Sydney University in Australia posted on X about declining a review invitation. The reason given was that the free version of an AI detector had rated the manuscript’s abstract as entirely AI-generated. The post drew more than 700,000 views.
The frustration is easy to understand. Reviewers work unpaid and should not have to spend their time on low-effort automated submissions. But many academics pushed back. AI detectors are notoriously inaccurate and often flag human writing as machine-made. Changing a few sentence structures or word choices can flip a verdict, which turns detection into a cat-and-mouse game. One commenter compared rejecting a paper on a detector’s word to refusing a physics manuscript because the author used a calculator.
Wariness of AI-generated submissions is reasonable. Blind trust in a flawed detection algorithm, however, appears to go too far.
arXiv caps submissions after a 40,000-paper month
On October 1, 2026, the preprint server arXiv put a new rate-limiting policy into effect, capping how many papers individual submitters can post. The trigger was a surge in volume driven by AI writing tools:
| Month | arXiv submissions |
|---|---|
| September 2016 | 9,869 |
| September 2024 | 20,569 |
| September 2026 | 40,363 (record) |
Typing speed and the time needed to synthesize ideas once set a natural ceiling on how much one person could publish. Generative AI removed that ceiling, and single authors can now produce dozens of papers in a month. The load falls on arXiv’s volunteer moderators. In September 2026 alone, staff and moderators faced nearly 9,000 support tickets, and postings were delayed across the site.
Academics and moderators have largely backed the cap. A mathematics professor at Rutgers University argued that authors should pick their best two papers a month rather than flood preprint servers. The limit looks like a sensible stopgap to keep volunteer infrastructure from collapsing.

▲ Submission surge at a preprint server
Mathematicians push back on OpenAI’s 719 manuscripts
On the same day as its watermark brief, OpenAI released a GitHub repository of 719 math manuscripts, grouped into 372 families, written by its internal frontier models. A working group of mathematicians published a statement on Terence Tao’s blog arguing that the release ignored research norms, that no one had asked for a mass of unverified papers, and that dumping them all at once was a show of corporate power rather than scholarship. Whether the backlash reflects legitimate concerns about ethics and verification, or unease about machines taking over core parts of mathematical work, remains an open question.
What researchers and institutions should do now
AI detection tools are not yet reliable judges, and AI cannot replace expert judgment of research quality. Researchers and institutions can act on a few clear points:
- Do not treat a watermark or AI detector result as proof of misconduct on its own. Short and formula-heavy texts are especially hard to judge.
- Remember that a watermark does not define authorship. Human authors stay fully responsible for the accuracy of what they submit.
- Use generative AI for polishing style, checking grammar, formatting and structuring drafts, not for producing unverified scientific claims.
- If English is not your first language, use AI openly as translation and phrasing support, but keep the methods and substance your own.
- Post a small number of strong papers to preprint servers such as arXiv instead of mass-producing drafts.