To read the original article in full go to : A biologist queries whether AI ‘borrowed’ his work to make major discovery – an expert explains the rules.
Below is a short summary and detailed review of this article written by FutureFactual:
Anthropic Claude Discovers ARTs in Jumbo Phages, Sparking Debate on AI-Led Biological Discovery and Credit
Short summary
The Conversation reports that Anthropic's Claude identified a novel enzyme system in jumbo phages, raising questions about AI-assisted science and credit. Original publisher: The Conversation.
- Claude reportedly identified array-associated reverse transcriptases (ARTs) in jumbo bacteriophages, illustrating AI-driven biological discovery.
- A Danish PhD candidate, Mario Rodríguez Mestre, says his team has studied these enzymes since 2022 and shared drafts with Claude, prompting questions about data provenance.
- Anthropic denies training on user transcripts and says their biology team had no access to unpublished material, but independent verification remains unresolved.
- The piece underscores the need for transparent attribution and data-trail records when AI tools contribute to scientific findings.
Overview
The Conversation reports that the AI system Claude, from Anthropic, identified a novel group of enzymes within jumbo bacteriophages. These enzymes are reverse transcriptases, which copy RNA into DNA, a process central to many diagnostic and research applications. Claude did not just recognize individual enzymes but a broader system that includes the enzyme, a set of repeated DNA sequences, and a potential partner protein. The finding has prompted a broader discussion about AI-assisted discovery and how to attribute human contributions when an AI tool participates in scientific results.
ARTs in jumbo phages and what was found
Reverse transcriptases are enzymes that convert RNA to DNA, a mechanism used widely in biology and biotechnology. The ARTs described by Anthropic were discovered in bacteriophages, viruses that infect bacteria. The claim is that ARTs form part of a larger system involving repeated DNA sequences that produce small RNA molecules, suggesting a functional network beyond a single enzyme. Some evidence for this system came from reanalyzing data published in 2022 by another research group, illustrating how existing data can yield new insights when reexamined. Importantly, the researchers have not yet demonstrated that ARTs are active in copying RNA to DNA within the virus or that the system has a defined biological role in phage biology. The resemblance of ARTs to CRISPR has attracted attention, but the authors stress that this does not imply a gene-editing capability at this stage, but rather a direction for future investigation.
Data provenance and the Mestre dispute
A PhD candidate from the University of Copenhagen, Mario Rodríguez Mestre, has argued that unpublished work from his team on jumbo-phage reverse transcriptases may be relevant to Claude's discovery. Mestre and colleagues have studied enzymes they call “jumbotrons” since 2022 and have reportedly shared their dissertation and manuscript drafts with Claude. The Scientist magazine reported Mestre's claims about sharing material with Claude. Anthropic has publicly denied that Claude was trained on user transcripts or that its biology team accessed Mestre’s unpublished work, but the available reporting cannot independently resolve what data Claude actually used, if any. The US patent application for engineered retrons lists jumbo-phage reverse transcriptases among sequences tied to retron systems, underscoring that related reverse transcriptases have appeared in prior work but not proving that Claude’s ARTs were characterized before the Anthropic announcement.
How scientists verify AI-driven discoveries
The Anthropic report outlines why validating AI-driven discoveries matters. In ten additional search runs, Claude failed to identify the repeat array unless sequences were supplied directly. This highlights a difference between independently discovering a pattern and recognizing a pattern when it is given, challenging the claim of autonomous discovery. The report notes that it has not yet undergone peer review, so the scientific community will likely weigh the findings against independent replication and independent data provenance checks.
Implications for credit, data governance, and responsible use
The article emphasizes that human scientists conducted the laboratory work and that some questions remain unresolved about the activity and function of ARTs in the virus. It also raises broader questions about how to credit original data producers when an AI tool helps interpret or identify connections. The text argues that scientists should document model versions, inputs, and instructions used when AI is essential to methods, and that researchers should adhere to institutional guidelines about sharing confidential information, data privacy, and model-training permissions. It highlights policy nuances in different services offered by AI providers, where inputs and outputs may or may not be included in model training by default, depending on user permissions and service terms. The piece ultimately calls for transparent reporting of human contributions and explicit credit to those who generated and published the underlying data and ideas, alongside notes on what the AI achieved.
Future directions and takeaways for researchers
Anthropic’s expansion into a life sciences group signals ongoing investment in Claude for fundamental biology research. The article recommends researchers consult their university guidelines when sharing unpublished or confidential data with AI tools, verify sources, and keep records of AI usage. It also notes that data processed by AI platforms may be retained or used for training under certain conditions, so researchers should understand data retention and governance policies, including potential model training and data provenance concerns. The piece closes by arguing that ART findings raise compelling biological questions and that resolving them will require further experiments, more transparent data practices, and careful attribution of both human and AI contributions to scientific discoveries.



