Agentic AI can run any genomic analysis. The hard part is trusting it.
Manuel Corpas
Senior Lecturer in Genomics, AI, and Data Science
University of Westminster
Hackathon · King's College London · 18 June 2026
The shift
Two Waves of LLMs in Biology
First wave: information retrieval
Summarising papers, answering pathway questions, extracting structured data from text.
Useful but incremental.
Second wave: autonomous execution
Modern LLMs can write, debug, and execute code.
Connected to file systems, databases, command-line tools, they plan multi-step operations and adapt on intermediate results.
The researcher's role shifts from producing analyses to evaluating them.
Corpas, Fatumo, Guio. Agentic Genomics: From Pipeline Automation to Autonomous Validation. Cell Genomics (in revision), 2026.
The point I want to plant: this is not a chatbot story. The interesting
thing is autonomous tool use with consequence. Once the model is acting,
the rate-limiting step shifts. TIMING: 1.5 min.
Definition
Defining Agentic Genomics
Jointly necessary. Falsifiable via the perturbation test: an agent that ignores perturbed intermediate outputs is not agentic.
Full definition (delivered verbally): "The use of autonomous AI agents,
powered by large language models and operating within domain-constrained
skill libraries, to discover, plan, execute, and iteratively refine
multi-step genomic analyses, where the agent exercises runtime
decision-making over tool selection, parameterisation, error handling,
and output evaluation."
These four conditions are deliberately strict. They exclude workflow
automation (Nextflow, Snakemake, Galaxy, no runtime decisions). They
exclude AutoML (search within a fixed space). They exclude LLM-assisted
scripting (you execute, not the agent). And they exclude general-purpose
biomedical copilots (information retrieval, no multi-step execution
against real data). TIMING: 2 min.
Figure 1 · Cell Genomics (in revision)
The Paradigm Shift: Code Production to Validation
(A) Traditional workflow. Researcher writes code, configures tools, runs pipelines, interprets results. The bottleneck is code production.
(B) Agentic workflow. Researcher describes intent in natural language; an AI agent discovers and executes skills from a modular library; researcher validates results. The bottleneck shifts to validation and judgement.
This is the central diagram of the Cell Genomics Perspective (in revision). The skills shown
in panel B are real ClawBio skills: pharmacogenomics, variant annotation,
ancestry estimation, drug safety, genome QC, PRS calculation, nutri-genomics,
structural variants. Same audience question for the rest of the talk:
what does it take to make panel B trustworthy? TIMING: 1.5 min.
The new bottleneck
The Validation Bottleneck: Silent, Plausible-Looking Failure
AI agents produce results faster than humans can verify them.
AUTOBA
Pipelines omitted critical steps; wrong tool selected for the data type.
Zhou et al., Adv. Sci. 2024.
SINGLE-CELL AGENTS
Incomplete experimental designs; inconsistent recommendations for identical queries.
A skill silently returned "all normal" for 51 drugs on an empty input file.
Independent audit (S. Kornilov, clawbio_bench); ClawBio v0.5.0, Zenodo 2026.
Silent degradation to plausible-looking but incorrect results.
Each of these is from an independently developed system. The convergence is
the point: this is structural, not a one-off bug. The 51-drug ClawBio
incident is mine, surfaced by a community auditor. We discovered it because
the platform is open. That's an argument for transparency. TIMING: 2 min.
The skill library under test
ClawBio
An agent-native skill library for bioinformatics.
Open-source · local-first · reproducible
87
skills
973
GitHub stars
44
contributors
MIT
license
pharmgx-reportervariant-annotationclaw-ancestry-pcagwas-prsscrna-orchestratormendelian-randomisationwes-clinical-report-enmethylation-clock+ 79 more
github.com/ClawBio/ClawBio · the artifact under test in the empirical benchmark that follows.
Hero slide. Establish ClawBio as a real, public, open-source artifact
before the benchmark. The 87 skills span pharmacogenomics, variant
annotation, ancestry/PCA, GWAS/PRS, single-cell, multi-omics, MR,
clinical reporting (EN + ES). The benchmark in the next slides tests
ONE skill: pharmgx-reporter. The framework slide above (tiered
validation) is what ClawBio implements; this slide is the bridge
between abstract framework and concrete empirical test. TIMING: 1 min.
Empirical question
Can a Plain-Text SKILL.md Reach Clinical-Grade?
Domain: pharmacogenomics. Genotype to phenotype to drug recommendation, ground truth from CPIC guidelines.
If specification cannot improve reliability here, where the guideline is fixed and the consequences are clinically measurable, it is unlikely to help in less structured domains.
An empirical test: does specification close the gap?
Frame the question crisply: pharmacogenomics is the right test bed
because the guideline is fixed (CPIC) and the stakes are clinically
measurable. If specification does not help here, it does not help
anywhere. Next slide: the experimental design. TIMING: 1 min.
Experimental design
The first large-scale test of where trust must live
110
CPIC cases
×
9
frontier LLMs
×
3
ancestries
×
3
replicates
=
44,550
scored evaluations
The model reasons
stochastic · unauditable
→
The skill executes
deterministic · auditable
One question across every condition: where does correctness have to live?
Trust is architectural, not a property of the model
Accuracy climbs as correctness is constrained (free-prompt 80.6% → skill-reasoning 95.5% → control 100%), but only executing the skill delivers all four clinical-grade guarantees: deterministic, auditable, model-invariant, population-invariant. The model is not the system; the architecture is.
The counterintuitive result
Giving the model the right guideline made it more dangerous
Retrieval-augmenting the model with the correct CPIC text raised lethal-class errors from 24.6% to 36.6% while raising surface accuracy. Only executing the validated skill as code drives lethal errors toward zero. Correctness must be executed, not reasoned or retrieved.
And it is not equal
On real genomes, accuracy falls by ancestry
72%
European Corpas family
51%
Latin American Peruvian Genome Project
40%
East African Uganda Genome Resource
Curated accuracy of ~96% does not transfer to real diplotypes from over 7,000 individuals, and unguarded interpretation degrades along an ancestry gradient. Executing the skill removes the gradient. Validation is also an equity problem.
Live demo
Ask a real genome, live
Every answer is executed live by the ClawBio pharmgx-reporter v0.2.0 skill on the openly published Corpasome, not generated by a model. conversational.clawbio.ai
Closing thesis
Five Principles for Responsible Agentic Genomics
01
Domain expertise is irreducible
02
Validation proportional to consequence
03
Transparency is non-negotiable
04
Skills testable by design
05
Equity must be engineered
THE CENTRAL THESIS
Agentic genomics shifts the bottleneck from pipeline construction to validation. A plain-text skill specification can satisfy two of three clinical-grade requirements; external multi-site validation is the open work.
Close on the central thesis, not on a pitch. Read the principles, then
the closing line. The question is no longer whether agentic genomics will
be adopted; it is whether the field will establish the standards required
to make it trustworthy before it becomes ubiquitous. TIMING: 2 min,
leaving 5 to 7 min Q&A within the 30-min recorded slot.
Keep in touch
Stay in the ClawBio loop
Subscribe to events
Hackathons & workshops
luma.com/ClawBio
Join the WhatsApp group
Scan with the WhatsApp camera
Build with us: github.com/ClawBio/ClawBio · clawbio.ai
Logistical outro after the thesis. Two ways to stay connected: the Luma
calendar for future hackathons and workshops, and the WhatsApp group for
today's cohort. Leave this slide up during Q&A so people can scan both.