AlphaGenome Atlas: 9 Billion DNA Variants, One Free Lookup Table
Google DeepMind's AlphaGenome Atlas precomputes molecular effect predictions for all 9 billion human DNA single-nucleotide variants, with a free web portal and API.

Google DeepMind released AlphaGenome Atlas on September 8, 2026, a precomputed catalogue of molecular-effect predictions covering all roughly 9 billion single-nucleotide variants (single-letter DNA changes) possible in the human genome. The release adds a new AlphaGenome Variant Impact (AVI) score that collapses thousands of predictions per variant into one ranked number, plus feature attributions and a compendium of over 2,500 DNA sequence motifs. The resource is free for non-commercial research through a web portal and API, with commercial access on Google Cloud listed as coming soon.
What happened
| Detail | Fact |
|---|---|
| Release date | September 8, 2026 |
| Variants covered | ~9 billion single-nucleotide variants |
| Dataset size | 1 petabyte |
| Comparison to AlphaFold DB | More than 30x larger (AlphaFold holds 200M+ protein structures) |
| DNA motifs included | Over 2,500 recurrent short sequences |
| Original AlphaGenome model release | June 2025 |
| Access (non-commercial) | Free web portal and AlphaGenome API |
| Access (commercial) | Coming soon on Google Cloud; model available now via Model Garden |
The Atlas is the bulk output of running DeepMind’s AlphaGenome model across every possible single-nucleotide variant in the human genome and storing the results. AlphaGenome, first published in June 2025, predicts how a DNA change affects molecular processes such as gene expression and RNA splicing (the editing of genetic instructions before they produce proteins). Previously, researchers had to query the model one variant or one genomic region at a time. The Atlas replaces that per-query workflow with a lookup table.
The 1-petabyte dataset is also available as a skill in Google Antigravity, DeepMind’s agent platform.
What is actually inside the Atlas?
DeepMind describes four linked resources packed into the release:
- Molecular effect predictions: Thousands of predictions per variant across hundreds of human and mouse cell types and tissues, covering multiple aspects of gene regulation.
- AVI score: A single impact number per variant. It combines AlphaGenome’s regulatory predictions with AlphaMissense (DeepMind’s model for protein-altering variants), so it covers coding regions (roughly 2% of the genome) and non-coding regions (the other 98%).
- AVI feature attributions: Each score is broken into additive contributions from interpretable categories such as chromatin accessibility (how open the DNA is to proteins), splicing, and conservation. A researcher can see which biological process a variant is predicted to disrupt.
- DNA sequence motifs: A compendium of over 2,500 recurrent short sequences, including transcription factor binding sites (spots where proteins attach to DNA to switch genes on or off), with their genomic locations.
According to DeepMind, the AVI score delivers best-in-class performance across multiple variant pathogenicity and rare disease benchmarks. The full benchmark details are in the accompanying technical report.
Why it matters
Testing 9 billion mutations in a physical laboratory is not feasible. Running a large model on demand for every candidate variant is too slow for genome-scale studies. A precomputed lookup table with attached interpretation removes both constraints at once.
The AVI score is the part most worth watching. Rare disease researchers, drug target teams, and population-scale geneticists currently juggle a patchwork of tools to score variants in coding versus non-coding regions. A single number that spans both, and that can be decomposed into interpretable contributions, reduces a lot of manual reconciliation work.
The non-commercial access model is generous: the portal, API, and GitHub model are all free for academic teams today. The commercial tier on Google Cloud arriving later is the signal that DeepMind intends this to become infrastructure for biotech and pharma pipelines, not just a research curiosity.
For teams building AI-assisted research workflows, the AlphaGenome API is worth noting. Programmatic access to precomputed variant scores fits naturally into automated pipelines. If you are thinking about how to wire external AI APIs into your own systems, our AI integration services page covers how we approach that kind of build.
Our take
The scale here is real. Moving from “query one variant at a time” to “query a precomputed petabyte dataset” is a genuine workflow shift for anyone doing genome-wide association work. The AlphaFold comparison is instructive: that database reshaped structural biology not because the model was new, but because bulk precomputation made it frictionless to use. DeepMind is attempting the same move for variant interpretation.
The AVI score is the most commercially interesting piece. A single ranked number covering 98% of the genome (the non-coding part that most scoring tools still handle poorly) is exactly what a drug discovery pipeline wants. The feature attribution layer is also smarter than it sounds: knowing whether a high-impact score comes from splicing or chromatin changes tells a researcher which wet-lab experiment to run next, rather than just flagging a variant as “bad.”
The honest caveat: “best-in-class on benchmarks” is a claim that needs independent replication. Benchmark performance on curated pathogenicity datasets does not always transfer cleanly to novel rare disease variants. Watch how the research community uses the AVI score over the next six to twelve months before treating it as ground truth.
For business readers tracking AI infrastructure deals, this fits the same pattern we covered in Nvidia’s growing AI portfolio: large labs are racing to turn model outputs into durable data assets, not just model APIs. A petabyte dataset with proprietary scoring is sticky in a way that a model weight is not.
What to do about it
- If you run academic genomics research, register for the free AlphaGenome Atlas portal and test AVI scores against your existing variant shortlists.
- If you are in biotech or pharma, sign up for Google Cloud notifications for the commercial tier so you are not behind when it launches.
- If you build bioinformatics pipelines, evaluate the AlphaGenome API now (free for non-commercial use) to understand latency and query limits before the commercial version ships.
- Check the technical report for the AVI benchmark details before committing it to a production scoring pipeline.
The Atlas is the clearest sign yet that precomputed AI outputs, not on-demand inference, will be the default interface for genome-scale biology.
Frequently asked questions
What is AlphaGenome Atlas?
AlphaGenome Atlas is a precomputed dataset released by Google DeepMind covering molecular-effect predictions for all roughly 9 billion possible single-nucleotide variants in the human genome. It includes a new AVI score, feature attributions, and over 2,500 DNA sequence motifs.
Is AlphaGenome Atlas free to use?
Yes, for non-commercial academic research. The web portal and AlphaGenome API are free. The underlying model is also available on GitHub for academic use. Commercial access on Google Cloud is listed as coming soon.
What is the AVI score in AlphaGenome Atlas?
The AlphaGenome Variant Impact (AVI) score is a single number that ranks DNA variants by predicted impact. It combines AlphaGenome's regulatory predictions with AlphaMissense and covers both coding regions (about 2% of the genome) and non-coding regions (about 98%).
How big is the AlphaGenome Atlas dataset?
The dataset is 1 petabyte, which DeepMind states is more than 30 times larger than the AlphaFold Database, which holds over 200 million protein structure predictions.


