Filling the holes in whole genomes

August 6, 2026

For the past 30 years, “whole-genome sequencing” has been a misnomer. Today the T2T Consortium published a special collection of 12 papers in Cell and Cell Genomics heralding a future of truly complete genomes for humans and nearly any vertebrate 👨‍🔬🐒🐦🐀🦒🐎🫏🐹🐟 (sorry, no salamanders).

This milestone was enabled by the “Q100 Project”: our quest to assemble the complete, diploid genome of a real person, HG002, without errors. Not only did this project produce a new benchmark, it drove improvements in T2T sequencing and assembly technologies. See “A complete diploid human genome benchmark for personalized genomics”.

You might have assumed the genome sequencing problem was solved, because variant callers reached F1 scores of 0.999 against prior HG002 benchmarks, but those results were based on incomplete variant sets, which were constructed by read mapping and missing 15% of the genome. Beware of Goodhart’s law: if you make a list of variants the target, people will get really good at calling those variants. But those variants are not a genome! As a result, we’ve been stuck trying to read genomes under the lamppost of short reads.

To break through this ceiling, we teamed up with the Genome in a Bottle Consortium to create a new “genome benchmark” that represents the actual complete, diploid genome of HG002. Our T2T-HG002v1.1 genome benchmark is nearly perfect across 99.4% of the genome, adding 701.4 Mb of autosomal sequence and both sex chromosomes. And it’s not just in centromeres. Of all 50-kb windows genome-wide, 99.7% contain new bases compared to the prior GIAB v4.2.1 variant benchmark!

Benchmark against the complete HG002 genome and the picture changes. Long-read de novo assembly of “noisy” nanopore reads outperforms state-of-the-art variant calling by an order of magnitude, even when restricted to regions syntenic to GRCh38. We’ve been selling long reads short.

A major strength of T2T genomes is that they cleanly resolve heterozygous variation, complex repeats, and segmental duplications that are lost with short reads, including nearly 400 medically relevant genes, the MHC, and most of the Y chromosome. What have we been missing? Even between the two haplotypes of HG002 we see significant variation, including a megabase-scale inversion of the beta-defensin locus and multiple genes present in one haplotype but not the other, such as DUSP22, CFHR1, CFHR3, GSTT1, and GSTM1.

To properly understand and computationally model these complex regions, we need completely assembled haplotypes. (The diplotype?) That is what the cell’s regulatory machinery sees—shouldn’t our future sequence-to-function models see the same thing?

We are now applying the T2T recipe to hundreds of diverse human genomes as part of the Human Pangenome Project, which promises to expand our understanding of common genomic variation and enable better methods for genome inference (even from short reads). See “HPRC2: A human pangenome reference with near-complete coverage of common genetic variation”.

I hope this new HG002 benchmark will help push the field beyond calling variants and towards calling genomes. We lay this out in a new commentary, “Filling the holes in whole genomes: a vision for personalized genomics from telomere to telomere”.

Big thanks to the T2T and GIAB teams for making the Q100 project a success! A special thanks to Nancy Hansen and Justin Zook for helping to lead this project, as well as the 1KGP, PGP, and HPRC donors for openly releasing their genomic information to everyone’s benefit.

HG002 resources

Papers in the collection

With more to come! Keep an eye on our T2T Consortium page at Cell.