
Neural networks are generating fictional scientists whose names subsequently appear in real scientific databases as authors of non-existent studies. Researchers from Samsung and the University of Warsaw reached this conclusion in a preprint titled “Ghost Pair: Correlated LLM Names and Their Ghosts in Web and Academic Publications.”
When language models are asked to invent an expert in a specific field, they do not choose random names; instead, they repeatedly produce the same characters. The names Elena Vasquez and Marcus Chen appear more frequently than others. According to the study’s authors, these two have appeared as volcanologists, astronauts, thriller characters, podcast hosts, and co-authors of scientific papers across hundreds of generated documents.
As the authors discovered, each model has its own set of preferred names:
Claude frequently generates Elena Amara Okafor
Gemini — Arisa Torna and Lena Petrova
ChatGPT — Elara Voss and Marcus Chen
Some names “clump together,” appearing jointly as recurring pairs or trios.
The researchers’ most alarming finding concerns the Zenodo scientific repository, administered by CERN. There, the preprint authors discovered over 1,655 entries attributed to fictional names. These publications cite non-existent journals and feature backdated publication dates.
The problem is that Zenodo assigns a genuine DOI—a unique identifier for scientific articles—to every paper, which is then automatically picked up by aggregators. Consequently, these fake publications are already being indexed by Google Scholar and Semantic Scholar without any verification.
“The academic archive is quietly becoming infected,” the study’s authors state. The situation is compounded by the fact that anyone can create a free account on Zenodo. At the peak, 991 fake entries were registered in a single month. Moreover, on ResearchGate, fictitious names are forming entire “research teams” featuring co-authors generated by various neural networks.
Fake publications pollute the body of scientific knowledge: other researchers may cite non-existent works, and academic search algorithms may promote them in search results. At the same time, the specific names used by neural networks can serve as “fingerprints,” making it possible to identify exactly which AI generated a given text.
The preprint has not yet undergone peer review, and its findings have not been independently verified. None of the scientific platforms have publicly responded to the study; consequently, it remains unknown how many of the identified entries repository administrators will deem to be junk or whether they intend to remove them.