<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://genomeinformatics.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://genomeinformatics.github.io/" rel="alternate" type="text/html" /><updated>2026-09-10T15:43:43+00:00</updated><id>https://genomeinformatics.github.io/feed.xml</id><title type="html">GIS</title><subtitle>Genome Informatics Section</subtitle><entry><title type="html">Filling the holes in whole genomes</title><link href="https://genomeinformatics.github.io/T2Tv2/" rel="alternate" type="text/html" title="Filling the holes in whole genomes" /><published>2026-08-06T00:00:00+00:00</published><updated>2026-08-06T00:00:00+00:00</updated><id>https://genomeinformatics.github.io/T2Tv2</id><content type="html" xml:base="https://genomeinformatics.github.io/T2Tv2/"><![CDATA[<p>For the past 30 years, “whole-genome sequencing” has been a misnomer. Today the T2T Consortium published <a href="https://www.cell.com/consortium/telomere-to-telomere">a special collection of 12 papers in Cell and Cell Genomics</a> heralding a future of truly complete genomes for humans and nearly any vertebrate 👨‍🔬🐒🐦🐀🦒🐎🫏🐹🐟 (sorry, no salamanders).</p>

<!--excerpt-->

<p>This milestone was enabled by the “Q100 Project”: our quest to assemble the complete, diploid genome of a real person, HG002, without errors. Not only did this project produce a new benchmark, it drove improvements in T2T sequencing and assembly technologies. See <a href="https://www.cell.com/cell/fulltext/S0092-8674(26)00703-8">“A complete diploid human genome benchmark for personalized genomics”</a>.</p>

<p>You might have assumed the genome sequencing problem was solved, because variant callers reached F1 scores of 0.999 against prior HG002 benchmarks, but those results were based on incomplete variant sets, which were constructed by read mapping and missing 15% of the genome. Beware of Goodhart’s law: if you make a list of variants the target, people will get really good at calling those variants. But those variants are not a genome! As a result, we’ve been stuck trying to read genomes under the lamppost of short reads.</p>

<p>To break through this ceiling, we teamed up with the Genome in a Bottle Consortium to create a new “genome benchmark” that represents the actual complete, diploid genome of HG002. Our T2T-HG002v1.1 genome benchmark is nearly perfect across 99.4% of the genome, adding 701.4 Mb of autosomal sequence and both sex chromosomes. And it’s not just in centromeres. Of all 50-kb windows genome-wide, 99.7% contain new bases compared to the prior GIAB v4.2.1 variant benchmark!</p>

<p>Benchmark against the complete HG002 genome and the picture changes. Long-read de novo assembly of “noisy” nanopore reads outperforms state-of-the-art variant calling by an order of magnitude, even when restricted to regions syntenic to GRCh38. We’ve been selling long reads short.</p>

<p>A major strength of T2T genomes is that they cleanly resolve heterozygous variation, complex repeats, and segmental duplications that are lost with short reads, including nearly 400 medically relevant genes, the MHC, and most of the Y chromosome. What have we been missing? Even between the two haplotypes of HG002 we see significant variation, including a megabase-scale inversion of the beta-defensin locus and multiple genes present in one haplotype but not the other, such as <em>DUSP22</em>, <em>CFHR1</em>, <em>CFHR3</em>, <em>GSTT1</em>, and <em>GSTM1</em>.</p>

<p>To properly understand and computationally model these complex regions, we need completely assembled haplotypes. (The diplotype?) That is what the cell’s regulatory machinery sees—shouldn’t our future sequence-to-function models see the same thing?</p>

<p>We are now applying the T2T recipe to hundreds of diverse human genomes as part of the Human Pangenome Project, which promises to expand our understanding of common genomic variation and enable better methods for genome inference (even from short reads). See <a href="https://www.biorxiv.org/content/10.64898/2026.07.21.739710">“HPRC2: A human pangenome reference with near-complete coverage of common genetic variation”</a>.</p>

<p>I hope this new HG002 benchmark will help push the field beyond calling <em>variants</em> and towards calling <em>genomes</em>. We lay this out in a new commentary, <a href="https://www.cell.com/cell/fulltext/S0092-8674(26)00817-2">“Filling the holes in whole genomes: a vision for personalized genomics from telomere to telomere”</a>.</p>

<p>Big thanks to the T2T and GIAB teams for making the Q100 project a success! A special thanks to Nancy Hansen and Justin Zook for helping to lead this project, as well as the 1KGP, PGP, and HPRC donors for openly releasing their genomic information to everyone’s benefit.</p>

<h2 id="hg002-resources">HG002 resources</h2>

<ul>
  <li>💿 <a href="https://github.com/marbl/hg002">Sequence data</a></li>
  <li>💾 <a href="https://github.com/marbl/GQC">GQC benchmarking software</a></li>
  <li>🧬 <a href="https://www.telomere2telomere.org/">T2T Consortium page</a></li>
  <li>🧪 <a href="https://www.nist.gov/programs-projects/genome-bottle">GIAB Consortium page</a></li>
</ul>

<h2 id="papers-in-the-collection">Papers in the collection</h2>

<ul>
  <li><a href="https://www.cell.com/cell/fulltext/S0092-8674(26)00817-2">Filling the holes in whole genomes: A vision for personalized genomics from telomere to telomere</a></li>
  <li><a href="https://www.cell.com/cell/fulltext/S0092-8674(26)00703-8">A complete diploid human genome benchmark for personalized genomics</a></li>
  <li><a href="https://www.cell.com/cell/fulltext/S0092-8674(26)00225-4">Complex subtelomeric architectures in a complete rhesus macaque reference genome</a></li>
  <li><a href="https://www.cell.com/cell/fulltext/S0092-8674(26)00816-0">The complete genome of a songbird</a></li>
  <li><a href="https://www.cell.com/cell/fulltext/S0092-8674(26)00815-9">A complete genome for the common marmoset</a></li>
  <li><a href="https://www.cell.com/cell/fulltext/S0092-8674(26)00635-5">Human acrocentric chromosome short-arm de novo mutation and recombination</a></li>
  <li><a href="https://www.cell.com/cell-genomics/fulltext/S2666-979X(26)00143-6">Telomere-to-telomere genome assembly and a pangenome for the rat</a></li>
  <li><a href="https://www.cell.com/cell-genomics/fulltext/S2666-979X(26)00169-2">Phased T2T horse and donkey assemblies from a mule reveal peculiar equid centromere evolution</a></li>
  <li><a href="https://www.cell.com/cell-genomics/fulltext/S2666-979X(26)00186-2">Haplotype-resolved DiMeLo-seq maps centromeric chromatin in a complete diploid human genome</a></li>
  <li><a href="https://www.cell.com/cell-genomics/fulltext/S2666-979X(26)00185-0">Automatic generation of model sequences for complex regions in assembly graphs with TTT</a></li>
  <li><a href="https://www.cell.com/cell-genomics/fulltext/S2666-979X(26)00141-2">Finishing a complete giraffe genome from telomere to telomere with Verkko-Fillet</a></li>
</ul>

<p>With more to come! Keep an eye on our <a href="https://www.cell.com/consortium/telomere-to-telomere">T2T Consortium page</a> at Cell.</p>]]></content><author><name>phillippy</name></author><summary type="html"><![CDATA[For the past 30 years, “whole-genome sequencing” has been a misnomer. Today the T2T Consortium published a special collection of 12 papers in Cell and Cell Genomics heralding a future of truly complete genomes for humans and nearly any vertebrate 👨‍🔬🐒🐦🐀🦒🐎🫏🐹🐟 (sorry, no salamanders).]]></summary></entry><entry><title type="html">We’re moving!</title><link href="https://genomeinformatics.github.io/movingday/" rel="alternate" type="text/html" title="We’re moving!" /><published>2026-04-19T00:00:00+00:00</published><updated>2026-04-19T00:00:00+00:00</updated><id>https://genomeinformatics.github.io/movingday</id><content type="html" xml:base="https://genomeinformatics.github.io/movingday/"><![CDATA[<p>Friday was my last day at NHGRI. After 10 wonderful years, my lab is headed one hour north on I-95 to set up shop at Johns Hopkins University. This is a very bittersweet move for me, as NHGRI has provided an incredibly supportive environment for my research, both in terms of colleagues and resources, and it’s hard to say goodbye. However, I am excited for the opportunity to tackle some new challenges at JHU.</p>

<!--excerpt-->

<p>Joining NHGRI in 2015 is one of the best decisions I ever made. I am so lucky to have been surrounded and supported by inspiring colleagues, whose positive examples have pushed me to be not just a better scientist, but a better person. Add in the amazing research environment of the NIH, and I can’t think of a better home over the past 10 years. I still remember being awestruck during my hiring interview with the great <a href="https://en.wikipedia.org/wiki/Daniel_L._Kastner">Dan Kastner</a>, and I’m grateful that my last week at NHGRI ended with a heartwarming symposium in his honor.</p>

<p>For those that don’t know my history, moving the lab to Baltimore is a bit of a homecoming. My career in genomics started as an undergraduate researcher with Art Delcher at Loyola University Maryland (then Loyola College), just a few blocks north of JHU’s Homewood campus. I’ve bounced around Maryland ever since: TIGR in Rockville → University of Maryland in College Park → NBACC in Frederick → NHGRI in Bethesda. I’m looking forward to coming “home” to Bmore.</p>

<p>Fitting for an <a href="https://doi.org/10.1371/journal.pcbi.0010006">antedisciplinary</a> scientist, my appointments at JHU will include the departments of Computer Science, Biomedical Engineering, and Genetic Medicine, and my lab will join me over the coming months. Stay tuned for more updates to follow, as we aim to make “telomere-to-telomere” genomes the new basis of clinical genomics. If you’ve read our latest preprint on <a href="https://www.biorxiv.org/content/10.1101/2025.09.21.677443">personalized genomes</a>, you know what’s next. If we want to understand the genome, we need to read the <em>whole thing</em>. My lab will continue the work we started at NHGRI, making the sequencing and interpretation of T2T genomes routine.</p>

<p>Farewell NHGRI. You will always be a part of me. Thank you for everything ❤️</p>

<p><img src="/downloads/nhgri2019.jpg" alt="alt text" title="NHGRI staff photo circa 2019" /></p>

<p>NHGRI Symposium 2019</p>]]></content><author><name>phillippy</name></author><summary type="html"><![CDATA[Friday was my last day at NHGRI. After 10 wonderful years, my lab is headed one hour north on I-95 to set up shop at Johns Hopkins University. This is a very bittersweet move for me, as NHGRI has provided an incredibly supportive environment for my research, both in terms of colleagues and resources, and it’s hard to say goodbye. However, I am excited for the opportunity to tackle some new challenges at JHU.]]></summary></entry><entry><title type="html">Choose your reference wisely</title><link href="https://genomeinformatics.github.io/references/" rel="alternate" type="text/html" title="Choose your reference wisely" /><published>2025-10-13T00:00:00+00:00</published><updated>2025-10-13T00:00:00+00:00</updated><id>https://genomeinformatics.github.io/references</id><content type="html" xml:base="https://genomeinformatics.github.io/references/"><![CDATA[<p>In honor of ASHG week, see <a href="https://www.nature.com/articles/s41592-025-02850-9">“Choose your human genome reference wisely”</a> (<a href="https://rdcu.be/eJejg">no paywall</a>), in which Vivien Marx interviewed me on the state of the human reference genome. Vivien is always fun to chat with and I was in a slightly opinionated mood from the start — <em>“The idea of a single reference genome is outdated,” says NIH researcher Adam Phillippy.</em> Some of my other quotes follow with a little added context.</p>

<!--excerpt-->
<p><em>“Conceptually, the pangenome represents all of humankind’s genetic information … Population projects cannot sample each individual in the world, so the idea is to represent the population’s multitude.”</em> This cannot be done with singular references, so enter the <a href="https://humanpangenome.org/">Human Pangenome Reference Consortium</a> (HPRC), where we are working to generate complete haplotypes for hundreds of individuals from around the world. The HPRC just released <a href="https://humanpangenome.org/hprc-data-release-2/">version 2</a> of this growing resource in May 2025, representing near-complete diploid assemblies for &gt;200 individuals.</p>

<p>This resource is a big change from the single reference we have become accustomed to, and we are still coming to grips with how to best leverage it. People often ask when it will be time to “switch” to a pangenome reference or express hesitation about its complexity. <em>When [Phillippy] hears scientists say: “Oh, the pangenome is not for me,” he tells them, “You’re using it.” Illumina’s DRAGEN software already calls variants using graph genomes. Approaches related to graph genomes are, he says, “happening behind the scenes.”</em></p>

<p>This point is often lost. One enormous benefit of building a pangenome is that it improves our general understanding of natural human variation. It’s like the 1000 Genomes Project, but inclusive of ALL variation, not just the variants you can see with short-read variant calling. There is a lot more structural variation in a typical human genome than most people realize, even between the two haplotypes of a single person’s genome, that can have big effects but are rarely captured.</p>

<p>By sampling the pangenome to build good priors on what a typical genome looks like, you can do a much better job of inferring a patient’s genome. In the short term, this means standard variant calling pipelines can acheive improved performance by first mapping to the pangenome, so that all the reads find their best matching haplotype, and then mapping the called variants back onto a common coordinate system like GRCh38. In the long term, <em>“Perhaps, in the future, scientists can depart from the approach of mapping sequencing reads … and accessing data in the context of the reference … I am suggesting we should flip that model, and we should map the metadata to the sequence of the patient, meaning we complete the patient’s genome, and then we take all of that metadata and we annotate it onto the personalized reference.”</em></p>

<p>Each genome is unique and should be treated as such. Analyzing the complete, personalized genome of an individual (yes, with the help of AI) will reduce reference bias and allow for the deep characterization of rare and novel structural variants that are the basis of many genetic diseases. The pangenome resources and genome inference approaches we are building will eventually enable complete, personalized “T2T” genomes for everyone. This is the thesis of “personalized genomes” as we recently described in the <a href="https://www.biorxiv.org/content/10.1101/2025.09.21.677443v1">Q100 project preprint</a>, and we plan to keep working towards this goal until it’s a reality. Stay tuned!</p>]]></content><author><name>phillippy</name></author><summary type="html"><![CDATA[In honor of ASHG week, see “Choose your human genome reference wisely” (no paywall), in which Vivien Marx interviewed me on the state of the human reference genome. Vivien is always fun to chat with and I was in a slightly opinionated mood from the start — “The idea of a single reference genome is outdated,” says NIH researcher Adam Phillippy. Some of my other quotes follow with a little added context.]]></summary></entry><entry><title type="html">The formation and propagation of human Robertsonian chromosomes</title><link href="https://genomeinformatics.github.io/ROBs/" rel="alternate" type="text/html" title="The formation and propagation of human Robertsonian chromosomes" /><published>2025-09-24T00:00:00+00:00</published><updated>2025-09-24T00:00:00+00:00</updated><id>https://genomeinformatics.github.io/ROBs</id><content type="html" xml:base="https://genomeinformatics.github.io/ROBs/"><![CDATA[<p>🚂 The T2T train keeps rolling. Our latest work investigating the structure and cause of Robertsonian chromosomes with the Gerton and Garrison labs <a href="https://www.nature.com/articles/s41586-025-09540-8">is out today in the journal Nature</a>!  What’s a Robertsonian chromosome and why do they matter? More info below the break…</p>

<!--excerpt-->

<p>Let Jen explain the basics in <a href="youtu.be/JmlY5omxQVc">this great video</a> put out by the Stowers Institute press team. You can also find more backstory on this project in the <a href="https://www.stowers.org/news/stowers-scientists-identify-the-fusion-point-of-robertsonian-chromosomes-hinting-at-how-chromosomes-evolve">Stowers Institute press release</a>.</p>

<p>Big congrats to the whole team including Leonardo Gomes de Lima, Andrea Guarracino, Sergey Koren, Tamara Potapova, multiple members of the Genome Informatics Section, the NIH Intramural Sequencing Center, and of course Erik Garrison, Jen Gerton, and their respective labs.</p>

<p>To top it off, here is an article from the Washington Post science section featuring Jen looking at some clear liquid in a tube! 🧪 <a href="https://www.washingtonpost.com/science/2025/09/24/genetic-anomaly-infertility-robertsonian-translocation-chromosome/">A genetic anomaly linked to infertility was a puzzle. Scientists solved it.</a></p>

<p><img src="/downloads/JenG.jpg" alt="alt text" title="Dr. Jennifer Gerton looking at some colorless liquid in a test tube" /></p>

<p>Photo credit: Jill Toyoshiba/Stowers Institute for Medical Research</p>]]></content><author><name>phillippy</name></author><summary type="html"><![CDATA[🚂 The T2T train keeps rolling. Our latest work investigating the structure and cause of Robertsonian chromosomes with the Gerton and Garrison labs is out today in the journal Nature! What’s a Robertsonian chromosome and why do they matter? More info below the break…]]></summary></entry><entry><title type="html">The Q100 preprint!</title><link href="https://genomeinformatics.github.io/Q100/" rel="alternate" type="text/html" title="The Q100 preprint!" /><published>2025-09-22T00:00:00+00:00</published><updated>2025-09-22T00:00:00+00:00</updated><id>https://genomeinformatics.github.io/Q100</id><content type="html" xml:base="https://genomeinformatics.github.io/Q100/"><![CDATA[<p>We are delighted to finally announce a preprint describing the Q100 project, where we finished the HG002 genome to near-perfect accuracy: <a href="https://www.biorxiv.org/content/10.1101/2025.09.21.677443v1">A complete diploid human genome benchmark for personalized genomics</a></p>

<!--excerpt-->

<p>Building benchmarks is hard, unglamorous work, but the impact can be huge. Consider how much the Genome in a Bottle (GIAB) variant benchmarks have shaped the field over the past ~10 years. However, these mapping-based benchmarks omit about 15% of the genome (and not just in centromeres). Assembling and annotating the complete, diploid HG002 genome from “T2T” allowed us to fill in those missing regions and explore the limitations of reference-based variant calling and benchmarking, especially within complex, segmentally duplicated regions that are often heterozygous.</p>

<p><img src="/downloads/HG002Figure2.png" alt="alt text" title="Previously missing benchmark regions overlaid on the complete HG002 genome" /></p>

<p>In our usual fashion, we ran this project entirely in the open, with the first T2T-HG002 assembly released on <a href="https://github.com/marbl/hg002">GitHub</a> back in Nov 2022. What took so long to write the paper? We spent a lot of time checking our work, but we also had to develop new methods for benchmarking against a complete, diploid T2T genome. Enter <a href="https://github.com/marbl/GQC">Genome Quality Checker (GQC)</a> by project leader, <a href="https://genomeinformatics.github.io/people/hansen/">Nancy Hansen</a>. With a complete benchmark and appropriate QC methods now in place, we can measure the accuracy of assemblies, variant callsets, and even raw reads across the entire genome.</p>

<p>This revealed something surprising: long-read de novo assembly methods now outperform reference-based variant calling not just in completeness, but also in overall accuracy, by a substantial margin (10 QV)! Most of this gain comes from “hard to call” regions of the genome, but the result still holds when looking only at regions where the variant caller (e.g. DeepVariant) reports high confidence (GQ &gt;40). We are actively digging into this result, but it hints at plenty of room for improvement in modern variant calling methods beyond the 0.999 F1 scores we’ve come to expect from variant benchmarks.</p>

<p>A huge thanks to Nancy Hansen and the whole Q100 team for shepherding this project over the past 3 years. Along the way, we made new friends who used our methods (e.g. Verkko, Merqury) to assemble near-perfect T2T genomes of their own, including an <a href="https://www.biorxiv.org/content/10.1101/2025.08.01.667781v2">East Asian diploid reference</a>, a <a href="https://www.biorxiv.org/content/10.1101/2025.07.12.664550v2">South Asian diploid reference</a>, a <a href="https://www.biorxiv.org/content/10.1101/2025.08.04.668424v1">rhesus macaque reference</a>, and complete diploid assemblies of the human cell lines <a href="https://www.nature.com/articles/s41467-025-62428-z">RPE-1</a> and <a href="https://www.biorxiv.org/content/10.1101/2025.04.15.648829v1">BJ and IMR-90</a>. Each of these studies includes their own unique twists and insights, and we encourage you to read them as well.</p>

<p>Clinically routine T2T genomes are finally in sight!</p>]]></content><author><name>phillippy</name></author><summary type="html"><![CDATA[We are delighted to finally announce a preprint describing the Q100 project, where we finished the HG002 genome to near-perfect accuracy: A complete diploid human genome benchmark for personalized genomics]]></summary></entry><entry><title type="html">We are looking for postbacs and postdocs!</title><link href="https://genomeinformatics.github.io/jobs2025/" rel="alternate" type="text/html" title="We are looking for postbacs and postdocs!" /><published>2025-05-12T00:00:00+00:00</published><updated>2025-05-12T00:00:00+00:00</updated><id>https://genomeinformatics.github.io/jobs2025</id><content type="html" xml:base="https://genomeinformatics.github.io/jobs2025/"><![CDATA[<p>Join our team and contribute to the development of complete, personalized “telomere-to-telomere” (T2T) genome assemblies and the analysis of previously inaccessible regions of the genome! We are currently accepting applications for <strong>postbaccalaureate and postdoctoral researchers</strong>.</p>

<!--excerpt-->

<p>These positions are under the supervision of <a href="https://www.genome.gov/staff/Adam-M-Phillippy-PhD">Dr. Adam Phillippy</a>, whose <a href="https://genomeinformatics.github.io/">research section</a> develops and applies computational methods for the analysis of massive genomics datasets with a focus on genome sequencing and comparative genomics. Dr. Phillippy and team are leaders in the field of genome assembly and sequence analysis who have developed many widely used bioinformatics tools (e.g., <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC395750/">MUMmer</a>, <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4915045/">Mash</a>, <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411767/">Canu</a>, <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10427740/">Verkko</a>), finished the first <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9186530/">truly complete sequence of a human genome</a>, and recently completed <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC12058530/">T2T reference genomes for the great apes</a>.</p>

<p>The section is seeking applicants with an interest in developing and/or applying computational methods for genome assembly, sequence alignment, variant detection, variant annotation, and information visualization. Current projects include efforts to sequence and explore all corners of the human pangenome (<a href="https://humanpangenome.org/">Human Pangenome Project</a>) and enable personalized genomics for the diagnosis of genetic diseases through collaborations with the <a href="https://www.genome.gov/Current-NHGRI-Clinical-Studies/Undiagnosed-Diseases-Program-UDN">Undiagnosed Disease Program</a> and others.</p>

<p>Perform research at the forefront of genomics in an exciting and supportive environment. Postdocs in the <a href="https://genomeinformatics.github.io/">Genome Informatics Section</a> are supported for up to 5 years and have wide latitude to carry out their own research vision. Postbacs are typically 2-year appointments with more direct supervision. These positions are entirely focused on research and are meant to equip trainees for the next step in their careers. We are best suited to mentor researchers with overlapping interests to our own, and you can find a list of our publications via Dr. Phillippy’s <a href="https://scholar.google.com/citations?hl=en&amp;user=PTTAqsgAAAAJ&amp;view_op=list_works&amp;sortby=pubdate">Google Scholar</a> page. Strong computational skills are a prerequisite and experience with genomic data is a plus. Postdoctoral stipends start around $70k per year and include family health insurance at no additional cost. Postbac stipends start around $40k per year. International postdocs are sponsored for a J-1 visa (while postbacs must be US citizens or permanent residents). More information on the NIH postdoc program is available from the <a href="https://www.training.nih.gov/research-training/">NIH Office of Intramural Training and Education</a>.</p>

<p><strong>To apply:</strong> Interested applicants should submit their CV, a brief statement of interest, and the names of three references to: adam.phillippy@nih.gov</p>

<p>These are in-person positions. The NHGRI Intramural Research Program is located on NIH’s main campus in Bethesda, Maryland, and offers a wide array of training and collaboration opportunities, including access to extensive high-performance computing resources (<a href="https://hpc.nih.gov/">BioWulf</a>), the NIH intramural sequencing center (<a href="https://www.nisc.nih.gov/">NISC</a>), and the NIH Clinical Center.</p>]]></content><author><name>phillippy</name></author><summary type="html"><![CDATA[Join our team and contribute to the development of complete, personalized “telomere-to-telomere” (T2T) genome assemblies and the analysis of previously inaccessible regions of the genome! We are currently accepting applications for postbaccalaureate and postdoctoral researchers.]]></summary></entry><entry><title type="html">Complete sequencing of ape genomes</title><link href="https://genomeinformatics.github.io/ApesComplete/" rel="alternate" type="text/html" title="Complete sequencing of ape genomes" /><published>2025-04-09T00:00:00+00:00</published><updated>2025-04-09T00:00:00+00:00</updated><id>https://genomeinformatics.github.io/ApesComplete</id><content type="html" xml:base="https://genomeinformatics.github.io/ApesComplete/"><![CDATA[<p>Today we published the <em>complete</em> “T2T” genomes of 6 ape species: chimp, bonobo, gorilla, Sumatran orangutan, Bornean orangutan, and siamang gibbon. This landmark resource is the result of a long-running collaboration (5 years of work!) led by myself, Kateryna Makova, and Evan Eichler. The genomes and our initial analyses are now presented in two papers: <a href="https://doi.org/10.1038/s41586-025-08816-3">Complete sequencing of ape genomes</a> published today, and <a href="https://doi.org/10.1038/s41586-024-07473-2">The complete sequence and comparative analysis of ape sex chromosomes</a> published last spring. There is a tremendous amount of data, code, etc. that goes along with this project, which we have organized on the <a href="https://github.com/marbl/Primates">T2T-primates project page</a>. Don’t miss the <a href="https://github.com/marbl/T2T-Browser">T2T Browser Hub</a> which presents all of this data as browser tracks, including expression, methylation, gene annotation, repeat annotation, etc. Comparing these genomes to our own furthers our understanding of human biology and genetic disease, including what makes us uniquely human, and brings us one step closer to understanding the language of the genome. I am excited to see what new discoveries will arise from these genomes!</p>]]></content><author><name>phillippy</name></author><summary type="html"><![CDATA[Today we published the complete “T2T” genomes of 6 ape species: chimp, bonobo, gorilla, Sumatran orangutan, Bornean orangutan, and siamang gibbon. This landmark resource is the result of a long-running collaboration (5 years of work!) led by myself, Kateryna Makova, and Evan Eichler. The genomes and our initial analyses are now presented in two papers: Complete sequencing of ape genomes published today, and The complete sequence and comparative analysis of ape sex chromosomes published last spring. There is a tremendous amount of data, code, etc. that goes along with this project, which we have organized on the T2T-primates project page. Don’t miss the T2T Browser Hub which presents all of this data as browser tracks, including expression, methylation, gene annotation, repeat annotation, etc. Comparing these genomes to our own furthers our understanding of human biology and genetic disease, including what makes us uniquely human, and brings us one step closer to understanding the language of the genome. I am excited to see what new discoveries will arise from these genomes!]]></summary></entry><entry><title type="html">Verkko2 is released!</title><link href="https://genomeinformatics.github.io/Verkko2/" rel="alternate" type="text/html" title="Verkko2 is released!" /><published>2025-01-02T00:00:00+00:00</published><updated>2025-01-02T00:00:00+00:00</updated><id>https://genomeinformatics.github.io/Verkko2</id><content type="html" xml:base="https://genomeinformatics.github.io/Verkko2/"><![CDATA[<p>We are excited to announce that Verkko2 is now available! Not only is it 4x faster than Verkko1, this version adds support for proximity ligation data (e.g. Hi-C, Pore-C) for T2T phasing and scaffolding without the need for trios. Our latest preprint describes the new methods and results: <a href="https://www.biorxiv.org/content/10.1101/2024.12.20.629807v2">“Verkko2: Integrating proximity ligation data with long-read De Bruijn graphs for efficient telomere-to-telomere genome assembly, phasing, and scaffolding”</a>. With these improvements, Verkko2 can now assemble, on average, around 40 out of 46 diploid human chromosomes as T2T scaffolds (and ~20 as T2T contigs), including the most difficult to assemble acrocentric chromosomes. However, these improvements are not limited to human genomes and Verkko2 should work well for any diploid or haploid genome (polyploids are a work in progress). We look forward to enabling many more T2T genomes in 2025!</p>]]></content><author><name>antipov</name></author><summary type="html"><![CDATA[We are excited to announce that Verkko2 is now available! Not only is it 4x faster than Verkko1, this version adds support for proximity ligation data (e.g. Hi-C, Pore-C) for T2T phasing and scaffolding without the need for trios. Our latest preprint describes the new methods and results: “Verkko2: Integrating proximity ligation data with long-read De Bruijn graphs for efficient telomere-to-telomere genome assembly, phasing, and scaffolding”. With these improvements, Verkko2 can now assemble, on average, around 40 out of 46 diploid human chromosomes as T2T scaffolds (and ~20 as T2T contigs), including the most difficult to assemble acrocentric chromosomes. However, these improvements are not limited to human genomes and Verkko2 should work well for any diploid or haploid genome (polyploids are a work in progress). We look forward to enabling many more T2T genomes in 2025!]]></summary></entry><entry><title type="html">The NHGRI Center for Genomics and Data Science Research is hiring!</title><link href="https://genomeinformatics.github.io/jobs2024/" rel="alternate" type="text/html" title="The NHGRI Center for Genomics and Data Science Research is hiring!" /><published>2024-08-13T00:00:00+00:00</published><updated>2024-08-13T00:00:00+00:00</updated><id>https://genomeinformatics.github.io/jobs2024</id><content type="html" xml:base="https://genomeinformatics.github.io/jobs2024/"><![CDATA[<p><strong>Update: This search has closed</strong> Join our team and contribute to the development of complete, personalized “telomere-to-telomere” (T2T) genome assemblies and the analysis of previously inaccessible regions of the genome! We are currently accepting applications for center coordinator, bioinformatics engineer/scientist, and postdoctoral researcher.</p>

<!--excerpt-->

<p>These positions are under the supervision of center director <a href="https://www.genome.gov/staff/Adam-M-Phillippy-PhD">Dr. Adam Phillippy</a>, whose <a href="https://genomeinformatics.github.io/">research section</a> develops and applies computational methods for the analysis of massive genomics datasets with a focus on genome sequencing and comparative genomics. Dr. Phillippy and team are leaders in the field of genome assembly and sequence analysis who have developed many widely used bioinformatics tools (e.g. <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC395750/">MUMmer</a>, <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4915045/">Mash</a>, <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411767/">Canu</a>, <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10427740/">Verkko</a>), finished the first <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9186530/">truly complete sequence of a human genome</a>, and recently completed <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11168930/">T2T reference genomes for the great apes</a>.</p>

<p>The section is seeking applicants with an interest in developing and/or applying computational methods for genome assembly, sequence alignment, variant detection, variant annotation, and information visualization. Current projects include efforts to sequence and explore all corners of the human pangenome (<a href="https://humanpangenome.org/">Human Pangenome Project</a>), complete the genomes of diverse vertebrate species (<a href="https://vertebrategenomesproject.org/">Vertebrate Genomes Project</a>), and enable personalized genomics for the diagnosis of genetic diseases through collaborations with the <a href="https://www.genome.gov/Current-NHGRI-Clinical-Studies/Undiagnosed-Diseases-Program-UDN">Undiagnosed Disease Program</a> and others.
Available positions include:</p>

<ol>
  <li>
    <p><strong>Center Coordinator.</strong> Help build the future of data science at NHGRI! This position will be responsible for coordinating research and outreach activities across the newly established <a href="https://www.genome.gov/about-nhgri/Division-of-Intramural-Research/Center-for-Genomics-and-Data-Science-Research">Center for Genomics and Data Science Research</a>. Duties will include the organization of bioinformatics conferences and other community-engagement activities aimed towards growing the profile of genomic data science at the NIH, as well as providing administrative support for members of the center and our associated research consortia (e.g. <a href="https://sites.google.com/ucsc.edu/t2tworkinggroup">T2T Consortium</a>). There will be additional opportunities to participate in center’s research. Past experience in genomics research and large-scale data management is preferred, but applications from all education levels from BS to PhDs are encouraged. Strong organizational and English skills are a requirement, along with an eagerness to help and enthusiasm for genomic data science! Salary commensurate with experience, roughly in the range of $80–130k.</p>
  </li>
  <li>
    <p><strong>Bioinformatics Engineer/Scientist.</strong> Enjoy the fun of bioinformatics methods development in a low-stress and highly collaborative environment. This position will support algorithm and software development within the <a href="https://genomeinformatics.github.io/">Genome Informatics Section</a>, primarily working with <a href="https://www.genome.gov/staff/Sergey-Koren-PhD">Dr. Sergey Koren</a> and the genome assembly team that has developed Canu and Verkko. A <a href="https://genomeinformatics.github.io/projects/">list of software</a> currently supported by the section is available on our homepage. Previous experience in bioinformatics, especially the problem of genome assembly is preferred, but there is no minimum education requirement. PhD applicants will be considered for the position of Bioinformatics Scientist. Very strong programming and analytical skills are a requirement. Salary commensurate with experience, roughly in the range of $100–150k.</p>
  </li>
  <li>
    <p><strong>Postdoctoral Researcher.</strong> Perform research at the forefront of genomics in an exciting and supportive environment. Postdocs in the <a href="https://genomeinformatics.github.io/">Genome Informatics Section</a> supervised by Dr. Phillippy are supported for up to 7 years and have wide latitude to carry out their own research vision. This position is focused entirely on research and is meant to build independence and equip the trainee for the next step in their career. We are best suited to mentor computational genomics researchers with overlapping interests to our own. You can find a list of our publications on our <a href="https://genomeinformatics.github.io/publications/">lab homepage</a> or via Dr. Phillippy’s <a href="https://scholar.google.com/citations?user=PTTAqsgAAAAJ&amp;hl=en">Google Scholar</a> page. Postdoctoral stipends start between $67–77k per year and include family health insurance at no additional cost. International postdocs are sponsored for a J-1 visa. More information on the NIH postdoc program is available from the <a href="https://www.training.nih.gov/research-training/pd/">NIH Office of Intramural Training and Education</a>.</p>
  </li>
</ol>

<p><strong>To apply:</strong> Interested applicants should submit their CV, a brief statement of interest, and the names of three references to: adam.phillippy@nih.gov</p>

<p>The NHGRI Intramural Research Program is located on NIH’s main campus in Bethesda, Maryland and offers a wide array of training and collaboration opportunities. These are in-person positions with flexible hours and an option for telework once a week. Funding is stable and includes access to extensive high-performance computing resources (<a href="https://hpc.nih.gov/">BioWulf</a>), the NIH intramural sequencing center (<a href="https://www.nisc.nih.gov/">NISC</a>), NHGRI core facilities, and the NIH Clinical Center.</p>

<p><em>The NIH is dedicated to building a diverse community in its training and employment programs.</em></p>]]></content><author><name>phillippy</name></author><summary type="html"><![CDATA[Update: This search has closed Join our team and contribute to the development of complete, personalized “telomere-to-telomere” (T2T) genome assemblies and the analysis of previously inaccessible regions of the genome! We are currently accepting applications for center coordinator, bioinformatics engineer/scientist, and postdoctoral researcher.]]></summary></entry><entry><title type="html">Going ape for T2T</title><link href="https://genomeinformatics.github.io/Apesv2/" rel="alternate" type="text/html" title="Going ape for T2T" /><published>2023-11-27T00:00:00+00:00</published><updated>2023-11-27T00:00:00+00:00</updated><id>https://genomeinformatics.github.io/Apesv2</id><content type="html" xml:base="https://genomeinformatics.github.io/Apesv2/"><![CDATA[<p>Last year we released complete, gapless, “T2T” sex chromosomes for chimp, bonobo, gorilla, Sumatran orangutan, Bornean orangutan, and siamang gibbon. This December we are proud to announce our latest preprint <a href="https://www.biorxiv.org/content/10.1101/2023.11.30.569198v1">“The Complete Sequence and Comparative Analysis of Ape Sex Chromosomes”</a>! Over the past year, we have also finished the autosomes for these genomes! The v2.0 assemblies for these species are now available from our <a href="https://github.com/marbl/Primates">T2T-primates project page</a>, and all of the raw HiFi, ONT, Hi-C, and Illumina sequencing data can be found on <a href="https://www.genomeark.org/t2t-all/">GenomeArk</a>. This has been a Herculean effort involving nearly everyone in the lab and a large swath of the T2T team. It turns out that finishing six genomes is a lot more work than finishing one! A huge thank you to everyone involved, especially Kateryna Makova for spearheading the project.</p>]]></content><author><name>phillippy</name></author><summary type="html"><![CDATA[Last year we released complete, gapless, “T2T” sex chromosomes for chimp, bonobo, gorilla, Sumatran orangutan, Bornean orangutan, and siamang gibbon. This December we are proud to announce our latest preprint “The Complete Sequence and Comparative Analysis of Ape Sex Chromosomes”! Over the past year, we have also finished the autosomes for these genomes! The v2.0 assemblies for these species are now available from our T2T-primates project page, and all of the raw HiFi, ONT, Hi-C, and Illumina sequencing data can be found on GenomeArk. This has been a Herculean effort involving nearly everyone in the lab and a large swath of the T2T team. It turns out that finishing six genomes is a lot more work than finishing one! A huge thank you to everyone involved, especially Kateryna Makova for spearheading the project.]]></summary></entry></feed>