Detect Deepfakesby Resemble AI
Deepfake case study · Multi-modal

The AI ‘Ghosts’ Contaminating Academic Publishing - 404…

Researchers reveal that large language models are creating ghost authors like Elena Vasquez and Marcus Chen to populate fake academic papers and websites

Incident date
Aug 2026
Target
Research Gold
Updated Aug 28, 2026 · 1 min read

A preprint research paper from Samsung and the University of Warsaw has exposed the proliferation of non-existent experts being used as ghost authors for AI-generated academic papers, articles, and books. These synthetic personas, including names like Elena Vasquez and Marcus Chen, are being produced by large language models (LLMs) that rely on correlated name priors when tasked with generating experts in specific fields.

What happened

The researchers discovered that LLMs frequently default to the same names and "correlated character ensembles" across different platforms. These ghost authors have been identified in various contexts, including a company called Research Gold, which claimed to provide human medical research that was entirely AI-generated. Furthermore, these fabricated identities have been used to spread misinformation, such as false allegations attributed to a non-existent "Dr. Elena Vasquez" regarding a nursing job.

The scale of this issue is significant; the researchers identified 1,655 ghost-authored records on the CERN-operated repository Zenodo, many of which featured backdated publication dates and legitimate Digital Object Identifiers (DOIs). Because these AI-generated papers are indexed by aggregators like Google Scholar and Semantic Scholar without verification, the scholarly record is becoming increasingly contaminated.

While the researchers suggest that these recurring names could potentially help identify the provenance of AI-generated content—such as noting that the name Elena Vasquez appears more frequently in content generated by Claude Sonnet 4—they warn that the influx of this "slop" into the internet's training data may soon make such detection methods unreliable. As academic repositories and search engines struggle to filter this content, the research highlights a growing infrastructure for large-scale scholarly record contamination.

Sources