ClinVar 2026-07: 951 records inside our regions changed classification since 2026-05. What changed
Sample report Support Log in
aimostı Check your file

Learn Explainer

Runs of homozygosity: what identical stretches of DNA record, and why Finns carry more

Every genome has stretches where its two copies match letter for letter, because both came down from one ancestor. How long those stretches are says roughly how long ago that ancestor lived. Finnish genomes carry more of them than most European genomes, and genomes from the north and east of Finland carry the most.

Key takeaways

  • Length dates the shared ancestor. Segments shared across m meioses average 100/m centimorgans, so an ancestor 25 generations back leaves runs of about 2 cM on average[3].
  • Finnish genomes carry more long runs than most European ones, and the youngest settlements in the north-east carry the most: 2.0 percent of the genome in runs over 1 Mb in North Kainuu against 0.9 percent on the south coast[6].
  • The Finnish excess comes mostly from few founders: only 1.65 percent of Finns studied had inbreeding levels like those of first cousins' children[6].
  • The report measures runs of 3 Mb and longer from whole-genome VCF and gVCF files only, and reads them as population history. It does not measure them from chip exports[12].

A run of homozygosity is a stretch of chromosome where the two copies a person inherited, one from each parent, are identical over millions of bases. Long runs were long read as a sign of marriage between relatives; genome-wide data show that runs are universally common, including in outbred people[1]. In four European populations, runs of up to 4 Mb were common in people whose pedigrees showed no shared ancestor on the two sides for at least five generations[2]. Genomes differ less in whether they have runs than in how long the runs are and how much of the genome they add up to.

Two copies, one ancestor

Each of the 22 autosomes comes in two copies, one inherited from each parent. Over most of their length the two differ here and there, at the positions where a person is heterozygous. Inside a run of homozygosity they do not: position after position reads the same letter twice. A long run forms when the same stretch of an ancestor's chromosome reaches a person down both sides of the family tree, so that they inherit one ancestral piece twice[1].

Terms this guide relies on

Run of homozygosity (ROH)
A stretch, usually counted from several hundred thousand bases up, where both copies of a chromosome carry the same letters at every position read.
F(ROH)
The share of the autosomes inside runs longer than a chosen minimum. McQuillan and colleagues named it, and found that it tracked inbreeding measured from family trees in Orkney closely, with a correlation of 0.86[2].
Centimorgan (cM)
A distance on a chromosome measured in recombination: 1 cM is the stretch in which a crossover happens once in a hundred meioses. On average it spans about a million bases[3].
Meiosis
The cell division that makes eggs and sperm, in which each chromosome is cut and rejoined between the parent's two copies.

F(ROH) is only as comparable as its minimum. A study that counts runs from 500 kb adds up far more homozygous genome than one that starts at 3 Mb, so the figures in this guide always come with the length they were counted from.

Length works as a clock

Each meiosis cuts the chromosomes and splices them. A segment that travels from one ancestor down two lines of descent is cut along both, so the further back the ancestor lived, the shorter the piece that arrives intact on both copies. Thompson gives the rule: when two copies are separated by m meioses, the lengths of the segments they share follow an exponential distribution with a mean of 100/m centimorgans. The paper's worked example is a time depth of 25 generations, which is 50 meioses and an average of 2 cM[3].

Figure 1. Average length of a run by generations back to the ancestor both copies came from, from Thompson's formula with m = 2 × generations[3]. Our arithmetic: these are averages of a wide distribution, not the length of any one run. The dashed line is the report's 3 Mb minimum, drawn at 3 cM using the average of about one megabase per centimorgan.

So runs of 3 Mb and up, the ones the report counts, mostly come from ancestors within the last few dozen generations, and runs under 1 Mb from ancestors much of a population shares. Kirin and colleagues put it in one sentence: offspring of cousin marriages have long runs, while the numerous shorter tracts relate to shared ancestry tens and hundreds of generations ago[4].

The averages hide a wide spread, and that limits what one genome can say. McQuillan's paper gives the expected inbreeding coefficient of a first-cousin couple's child as 0.0625 with a standard deviation of 0.0243, because which segments pass down each line is a matter of chance; the groups it compared, sorted by pedigree, overlapped considerably in F(ROH)[2]. A single genome's total is one draw from that spread.

Short runs, long runs and the histories behind them

Pemberton and colleagues mapped runs in 1,839 people from 64 populations with 577,489 markers and found three length classes. Short runs of tens of kilobases probably reflect ancient haplotypes. Intermediate runs of hundreds of kilobases to several megabases probably result from background relatedness in a population of limited size. Long runs of multiple megabases probably result from recent relatedness between a person's parents[5]. Long runs were higher and more variable in most populations from the Middle East, Central and South Asia, Oceania and the Americas than in Africa, Europe and East Asia, and higher where consanguineous marriage is common[5].

Figure 2. The same genome yields different totals depending on where counting starts. The class edges are approximate, drawn from the authors' own descriptions; the starting points are the minimum run lengths each study and the report used[2, 4, 5, 6, 7].

Short runs carry most of the total, even where long runs are frequent. Kirin's survey of worldwide populations found that in every one, including the most inbred, more of the genome sat in runs of 0.5 to 2 Mb than in runs over 2 Mb[4]. A total built mostly from long runs points to recent shared ancestry; one built from many intermediate runs points to a small population, and the two can sit in the same genome.

A 2018 review of how runs are found in microarray and sequence data put the point in one line:

The number and length of ROH reflect individual demographic history

Ceballos and colleagues, Nature Reviews Genetics, 2018[1]

Why Finnish genomes carry more

Finland's inland north and east were settled late, from the 1500s[8], and the youngest sub-isolates in the north-east were founded only 300 to 400 years ago. Jakkula and colleagues read the genetic pattern as multiple bottlenecks from consecutive founder effects[6]: each settlement drew a small sample of genomes out of an older population, and its descendants then married mostly among themselves. The Finnish DNA guide traces the same history through the Y chromosome and the east–west divide.

They measured it in 1,395 people, from ten Finnish subpopulations plus comparison samples from Helsinki and Sweden, counting runs longer than 1 Mb from 231,116 markers. Mean homozygosity rose from 0.9 percent of the genome in the early-settled south coast to 2.0 percent in the North Kainuu isolate. The number and length of runs were highest in the youngest subpopulations, where up to 90 percent of people had at least one run longer than 5 Mb; in a US sample, 9.5 percent did[6].

Martin and colleagues saw the same geography in segments shared between people rather than within one genome, which is the same inheritance seen across two people instead of two copies. In 43,254 Finns, people around Kuusamo, a late-settled area in the north-east, shared on average about 60 Mb in segments longer than 3 cM with others born nearby; around Helsinki, Turku and Tampere the figure was about 5 to 15 Mb[9]. Across countries, two unrelated Finns shared on average 107.0 cM, two Swedes 22.9 cM[9].

~60 Mb

Shared in segments over 3 cM with people born nearby, around Kuusamo[9]

~5 to 15 Mb

The same measure around Helsinki, Turku and Tampere[9]

~63.9 Mb

In runs over 1 Mb, 1000 Genomes Finnish genomes; 40 Faroese averaged ~82.5 Mb[7]

Finland is not the extreme case. In high-coverage whole genomes analysed in 2026, the Finnish samples of the 1000 Genomes Project averaged about 63.9 Mb in runs longer than 1 Mb, and 40 people from the Faroe Islands about 82.5 Mb, more than any group in the study; the Faroese also carried more runs in the 5 to 15 Mb range, which the authors read as a more recent or stronger bottleneck[7]. On chip data, endogamous island communities in Dalmatia and Orkney carried several times more than cosmopolitan samples[2].

Table 1. Published averages of the genome in runs of homozygosity, each counted from its own minimum length
Population sampleStudy and dataRuns counted fromAverage in runs
CEU (HapMap)McQuillan 2008, 300,000-marker chip1.5 Mb8 Mb (F 0.003)
Scotland, 984 peopleMcQuillan 2008, 300,000-marker chip1.5 Mb7 Mb (F 0.003)
Orkney, three or more grandparents from one islandMcQuillan 2008, 300,000-marker chip1.5 Mb28 Mb (F 0.011)
Dalmatian island, all four grandparents from one villageMcQuillan 2008, 300,000-marker chip1.5 Mb35 Mb (F 0.013)
Finland, early-settled south coastJakkula 2008, 231,116 markers1 Mb0.9% of the genome
Finland, North Kainuu isolateJakkula 2008, 231,116 markers1 Mb2.0% of the genome
Finland, 1000 Genomes FINHamid 2026, whole genomes1 Mb~63.9 Mb
Faroe Islands, 40 peopleHamid 2026, whole genomes1 Mb~82.5 Mb

Source: McQuillan et al. 2008[2]; Jakkula et al. 2008[6]; Hamid et al. 2026[7]. Rows with different minimums are not directly comparable.

The link to the Finnish Disease Heritage

A recessive condition appears when both copies of a gene carry a disease-causing change. Inside a run, both copies are the same ancestral copy, so a variant that one founder carried can arrive twice. Martin and colleagues note that long runs are enriched for deleterious variation, and they found that two Finns who carry the same variant from the Finnish disease database share a stretch of genome around it about 30 percent of the time or more, an order of magnitude above the background rate[9].

The same settlement history drew both maps. FinDis records that for most heritage diseases the birthplaces of patients' families concentrate in the late-settlement area populated from the 1500s, and that Northern epilepsy is found in Kainuu, near the eastern border[8], where Jakkula measured the highest homozygosity[6]. The Finnish Disease Heritage guide covers the 39 diseases themselves, and carrier status vs your own risk what one copy of a recessive variant means. A run says nothing about which variants sit inside it, so this module reports no carrier results; the carrier panel reads them directly.

What each kind of file can measure

Most of what is known about runs comes from chips. McQuillan used a 300,000-marker chip, Jakkula 231,116 markers and Pemberton 577,489[2, 6, 5]. The consumer chip exports we measured hold between 563,320 and 955,958 rows[10], so a chip of that size can, in research hands, find runs of a megabase and longer. Density sets the floor: with a panel of 3 million markers, runs as short as 100 kb can be detected reliably, and Han Chinese genomes that showed about 130 Mb of runs on a panel of about 400,000 markers showed about 510 Mb on the dense one[4]. A typical genome differs from the reference at 4.1 to 5.0 million sites[11], and a whole-genome VCF lists them all: by our arithmetic, more than a thousand per megabase on average.

Our report does not measure runs on chips. Its scan counts heterozygous calls in each megabase of a whole-genome call set, and it was built and checked on that kind of file only. The report says so in place of a number. Exomes, targeted panels and low-pass or imputed genomes have the same problem in another shape, and the report refuses them too, naming the evidence it found.

Table 2. Runs of homozygosity in the report, by file type
FileWhat the report doesWhy
Chip export, such as 23andMe or AncestryDNANot measuredThe scan is built for whole-genome call sets and is not run on chip exports
Plain VCF from a whole genomeMeasuredThe variant list carries every heterozygous call
gVCF from a whole genomeMeasuredIts variant calls are counted; reference blocks are skipped
Exome or targeted-panel VCFNot measuredThe stretches between captured regions were never sequenced
Low-pass or imputed genomeNot measuredUncalled heterozygous positions would read as homozygous
Any VCF under 300 variants per megabaseNot measuredToo sparse to tell quiet from unexamined
BAM or CRAM (Deep Read)From the variant file uploaded with itThe aligned-reads slice is not scanned for runs

Source: Aimosti's report as of October 2026; the coverage page lists every module by file type[12].

What the report shows

In the report the module follows the ancestry section, under the heading Runs of homozygosity. The scan splits the autosomes into 1 Mb bins, treats a bin as homozygous when it holds at least 50 variant calls of which at most 2 are heterozygous, joins neighbouring bins, and keeps runs of 3 Mb or more. A bin with too few calls counts as unmeasured, never as homozygous. The total, divided by the 2,875 Mb of autosome the scan counts, is the share the report gives. The Finnish DNA guide draws the scan on an invented example.

The card leads with a band headline and three figures: the share of the autosomes in long homozygous stretches, to one decimal; the number of stretches of 3 Mb or more; and the longest stretch. Below them come the band's description, a short account of what runs mean and a caveat. The report's summary list shows the headline alone, never the percentage. A refused file gets the headline Not measured for this file and the reason.

Table 3. The report's three bands for the share of the autosomes in runs of 3 Mb or more
Share of autosomesHeadlineDescription shown
Under 1%Few long homozygous stretchesUnder one percent of your autosomes sits in a run of this length. This is the ordinary result across most of Europe.
1% to 3%A moderate share, common in FinlandBetween one and three percent. This is the range a Finnish genome commonly falls in, and it reflects the size of the population your ancestors came from rather than anything about your immediate family.
3% and overA larger share than most European genomes carryAbove three percent. Ancestry from a small or long-separated population, ancestors who married within one community, and closer family relationships all produce this pattern, and a genome alone does not say which applies. A clinical genetics service is where that question belongs.

Source: The report's ROH module, content version of 2026-09-03; its sources are the Ceballos, McQuillan and Martin papers cited here[1, 2, 9].

The band edges are round numbers chosen for reading, not clinical thresholds. They also sit on a different scale from the published Finnish figures: Jakkula counted runs from 1 Mb, the report from 3 Mb in whole 1 Mb bins, so the same genome usually shows a smaller share in the report than it would in that study[6]. The caveat on every card names descent from a founder population, marriage within a small community and recent shared ancestry together, and says the measurement does not distinguish between them.

What Aimosti would (and wouldn't) show you

From a whole-genome VCF or gVCF the report counts runs of 3 Mb and longer, gives their share of the autosomes, the number of runs and the longest, and places the share in one of three bands, read as population history. Chip exports, exomes, low-pass or imputed genomes and files under 300 variants per megabase get a sentence saying why nothing was measured.

What we won't claim

We won't read a run-of-homozygosity figure as a statement about anyone's parents or family, give it a health meaning, or treat a band boundary as a clinical threshold. One total cannot separate a founder population from recent shared ancestry, and the report does not try.

Bottom line. Runs of homozygosity are a record of how small and how separate the population behind a genome was, written in lengths. Finnish genomes carry the marks of a few small founding groups, most of all in the north and east. A single total mixes old history with recent history, which is why the report gives it as population history and nothing more.

Questions people ask

What do runs of homozygosity mean?

Stretches where the two copies of a chromosome are identical because both descend from one ancestor. Short runs go back to ancestors a population shares; long runs to ancestors within the last few dozen generations[4, 3].

Can 23andMe or AncestryDNA raw data show runs of homozygosity?

Research studies have measured runs of 1 Mb and longer from chips with 231,000 to 577,000 markers, and consumer exports hold 563,320 to 955,958 rows[6, 5, 10]. Our report does not measure runs from chip exports; it needs a whole-genome VCF or gVCF.

Which files does Aimosti measure runs of homozygosity from?

A whole-genome VCF or gVCF with at least 300 variants per megabase. Chip exports, exomes, targeted panels and low-pass or imputed genomes get a sentence saying why nothing was measured, and with Deep Read the variant file uploaded with the reads is the one scanned[12].

Why do Finns have more runs of homozygosity than other Europeans?

Because Finland was settled by small founding groups, the inland north and east only from the 1500s, and those communities stayed comparatively separate. Runs over 1 Mb covered 0.9 percent of the genome on the south coast and 2.0 percent in North Kainuu, and up to 90 percent of people in the youngest subpopulations had a run over 5 Mb, against 9.5 percent of a US sample[6].

References

  1. Ceballos FC, Joshi PK, Clark DW, Ramsay M, Wilson JF. Runs of homozygosity: windows into population history and trait architecture. Nature Reviews Genetics, 2018. doi:10.1038/nrg.2017.109 Cited for statements in its abstract, the part read in full; the review itself is not open access.
  2. McQuillan R, Leutenegger AL, Abdel-Rahman R, et al. Runs of homozygosity in European populations. American Journal of Human Genetics, 2008. doi:10.1016/j.ajhg.2008.08.007 2,618 people genotyped on the Illumina HumanHap300 chip. Mean F(ROH) with a 1.5 Mb minimum: endogamous Dalmatians 0.013 (35 Mb), endogamous Orcadians 0.011 (28 Mb), CEU 0.003 (8 Mb), Scottish 0.003 (7 Mb).
  3. Thompson EA. Identity by descent: variation in meiosis, across genomes, and in populations. Genetics, 2013. doi:10.1534/genetics.112.148825
  4. Kirin M, McQuillan R, Franklin CS, Campbell H, McKeigue PM, Wilson JF. Genomic runs of homozygosity record population history and consanguinity. PLoS ONE, 2010. doi:10.1371/journal.pone.0013996
  5. Pemberton TJ, Absher D, Feldman MW, Myers RM, Rosenberg NA, Li JZ. Genomic patterns of homozygosity in worldwide human populations. American Journal of Human Genetics, 2012. doi:10.1016/j.ajhg.2012.06.014
  6. Jakkula E, Rehnström K, Varilo T, et al. The genome-wide patterns of variation expose significant substructure in a founder population. American Journal of Human Genetics, 2008. doi:10.1016/j.ajhg.2008.11.005 1,395 people, 231,116 SNPs; runs longer than 1 Mb. Mean homozygosity 0.9 percent in ESS to 2.0 percent in ISC.
  7. Hamid I, Mortensen Ó, Refoyo-Martínez A, et al. Faroese whole genomes provide insight into ancestry and recent selection. eLife, 2026. doi:10.7554/eLife.107428 Average genome in runs longer than 1 Mb: about 82.5 Mb in 40 Faroese genomes, about 63.9 Mb in the 1000 Genomes Finnish (FIN) genomes.
  8. The Finnish Disease Heritage. FinDis, Finnish Disease Database, 2026.
  9. Martin AR, Karczewski KJ, Kerminen S, et al. Haplotype sharing provides insights into fine-scale population history and disease in Finland. American Journal of Human Genetics, 2018. doi:10.1016/j.ajhg.2018.03.003
  10. What your DNA file can read. Aimosti, 2026. Ten chip exports measured on 2026-10-06, holding 563,320 to 955,958 rows.
  11. The 1000 Genomes Project Consortium. A global reference for human genetic variation. Nature, 2015. doi:10.1038/nature15393
  12. Coverage: what each file type can tell you. Aimosti, 2026.

Last reviewed . Every number on this page links to the source it comes from; if one of them has moved, tell us.

This is the kind of answer we give. See what your file says.

See a sample report Get your report · $39