Sequencing.com — Outsmart Your Genes

hg19 vs. hg38: Why Coordinates from an Old DNA Report May Not Match

By Sequencing Team, The team of bioinformaticians, Genetic Health Coaches, and writers at Sequencing.

What is a reference genome?

When scientists analyze your DNA, they compare it against a master template called the reference genome, a complete, curated sequence of the human genome assembled from many donors. Think of it as the world’s most detailed biological atlas: every chromosome mapped, every gene assigned a precise address.

Those addresses called genomic coordinates are typically written as a chromosome number followed by a nucleotide position (for example, chr17:41,244,936). Every variant, gene, and region mentioned in a DNA report is defined by those coordinates against a specific reference genome build. Change the build, and the address changes too, even though the underlying biology hasn’t moved.

hg19 and hg38: two editions of the same atlas

The human reference genome has been revised multiple times as sequencing technology and data quality have improved. Two versions are in widespread clinical use today:

  • hg19 (also called GRCh37) — released in 2009 by the Genome Reference Consortium. The dominant standard in clinical labs throughout the 2010s, and still the basis of millions of older reports.
  • hg38 (also called GRCh38) — released in 2013 and updated regularly since. The current standard for most genomics research and modern clinical sequencing.

Both refer to the same species, the same chromosomes, and largely the same genes. But they are different editions of the atlas, and the street-numbering system changed between them.

What changed between hg19 and hg38?

The differences between the two builds are more than cosmetic. Here are the key improvements hg38 brought:

  • A more complete sequence. hg38 covers approximately 95% of the human genome, compared to roughly 92.5% in hg19. Thousands of previously unresolved gaps were filled in using improved sequencing methods. More complete coverage means fewer regions where a variant might be missed simply because the reference was blank.
  • Corrected sequence errors. Over 8,000 individual nucleotides were corrected between builds. In hg19, some positions contained errors inherited from early sequencing technology. hg38 fixed those, improving the accuracy of variant calling.
  • Better representation of human diversity. hg38 introduced hundreds of alternate haplotype sequences, versions of genomic regions that better reflect the natural genetic variation across different ancestral populations. hg19 had only nine such alternate sequences; hg38 launched with 261. This matters clinically because a variant that looks unusual against a single reference sequence may be a normal variant in certain populations.
  • Decoy sequences to reduce false positives. hg38 added decoy sequences (including the Epstein-Barr virus genome, which infects roughly 90% of people) that absorb misplaced sequencing reads and prevent them from generating spurious variant calls.
  • Rearranged chromosomal coordinates. Because new sequences were added and errors were corrected at specific positions throughout the genome, coordinates in regions affected by these sequence changes may differ between builds.

Why the same variant sits at a different position in each build

Here is the key insight: genomic coordinates are not universal. They are build-specific.

Think of it this way. Imagine a very long street with numbered houses. In the old edition of the map (hg19), a certain house sits at number 41,244,936. Then urban planners resurvey the area and discover that several streets were missing from the original map. They insert them. Now every house past that insertion point has a new number in the new edition (hg38). The house hasn’t moved, but its address has changed.

The Genome Reference Consortium made exactly this kind of correction across the entire genome between builds. Sequences were inserted, corrected, and reorganized at hundreds of locations. Variants downstream of regions affected by these corrections may have shifted coordinates in the new build.

A real-world example: BRCA1

BRCA1, a gene strongly associated with hereditary breast and ovarian cancer risk, sits on chromosome 17. In hg19, it occupies the region chr17:41,196,312–41,277,500. In hg38, the very same gene occupies chr17:43,044,295–43,125,483. That is a shift of approximately 1.85 million base pairs upstream.

Now take a specific pathogenic variant inside BRCA1: the c.68_69delAG mutation (also written 185delAG), a well-documented founder mutation in Ashkenazi Jewish populations.

BuildChromosomePosition
hg19chr1741,276,045
hg38chr1743,124,028

A difference of nearly 1,848,000 positions for the exact same variant, in the exact same gene.

If a person with an hg19-based report looks up position chr17:41,276,045 in a modern database that uses hg38 coordinates, they might land in the completely wrong region of chromosome 17 or find nothing at all.

What this means for an older report

If a DNA report or clinical sequencing test was performed before roughly 2016–2018, there is a reasonable chance its coordinates are in hg19. Many variant databases, including ClinVar, have since migrated to hg38 as their primary build. The mismatch can create real problems:

  • A variant reported as pathogenic in hg19 may not appear when its hg19 position is searched in an hg38-based database.
  • Combining data from two reports that used different builds can produce false positives or miss real variants entirely.
  • Variant databases may return outdated or simply empty results if queried with the wrong build’s coordinates.

The reference build used is almost always stated in the methodology or technical appendix of a clinical report. Look for the notation ‘GRCh37’ or ‘GRCh38’ (or ‘hg19’ and ‘hg38’ used interchangeably). If the build is not labeled, reports generated before 2016 are more likely to be in hg19; those from 2019 onward are more commonly in hg38.

Tools to convert coordinates between builds

When you have coordinates from one build and need to find the equivalent position in the other, two well-regarded tools can help:

  • Broad Institute LiftOver (liftover.broadinstitute.org) — a free, web-based converter. Enter a position in hg19, select hg38 as the target build, and the tool returns the corresponding coordinates. It also accepts entire coordinate files, making it useful for bulk conversions.
  • Franklin by Genoox (franklin.genoox.com) — a clinical variant database that displays variant information in both builds side by side. Useful for confirming a variant’s clinical significance regardless of which set of coordinates you are starting from.

One important caveat: In rare cases involving regions that were substantially resequenced or rearranged between builds, a coordinate may not map cleanly, or may produce ambiguous results. For variants in complex genomic regions, or when a result doesn’t match any known variant, a geneticist or genetic counselor should review the original report alongside the conversion.

References

[1] Zhao T, et al. Closing Human Reference Genome Gaps. G3 Genes|Genomes|Genetics. 2020;10(8):2801-2809.

[2] Guo Y, et al. Improvements and impacts of GRCh38 human reference on high throughput sequencing data analysis. Genomics. 2017;109(2):83-90.

[3] Church DM, et al. Extending reference assembly models. Genome Biol. 2015;16(1):13.

[4] Schneider VA, et al. Evaluation of GRCh38 and de novo haploid genome assemblies. Genome Res. 2017;27(5):849-864.

[5] Li H, et al. Exome variant discrepancies due to reference-genome differences. Am J Hum Genet. 2021;108(7):1239-1250.

[6] Nostos Genomics. Decoding genomes: the evolution from HG19 to HG38. Nostos Genomics. Published February 2024.

Frequently Asked Questions

How do I know which reference genome my report used?

Check the methodology or technical appendix of your report. It will typically state 'GRCh37/hg19' or 'GRCh38/hg38.' If it is not labeled, reports from before 2016–2018 are more likely to be in hg19; reports from 2019 onward are more commonly in hg38. If you are unsure, contact the lab that performed the test, as they can confirm which build was used.

Does the build affect which variants were found, or just where they appear?

Primarily where they appear, but build choice can also affect whether certain variants are detected at all. A small number of variants called under hg19 do not appear under hg38 (and vice versa) because the underlying reference sequence was corrected or reorganized in those regions. For most well-characterized variants, the same variant will be found in both builds; only the coordinates differ.