Learn Explainer
Runs of homozygosity: what identical stretches of DNA record, and why Finns carry more
Every genome has stretches where its two copies match letter for letter, because both came down from one ancestor. How long those stretches are says roughly how long ago that ancestor lived. Finnish genomes carry more of them than most European genomes, and genomes from the north and east of Finland carry the most.
Key takeaways
- Length dates the shared ancestor. Segments shared across m meioses average 100/m centimorgans, so an ancestor 25 generations back leaves runs of about 2 cM on average[3].
- Finnish genomes carry more long runs than most European ones, and the youngest settlements in the north-east carry the most: 2.0 percent of the genome in runs over 1 Mb in North Kainuu against 0.9 percent on the south coast[6].
- The Finnish excess comes mostly from few founders: only 1.65 percent of Finns studied had inbreeding levels like those of first cousins' children[6].
- The report measures runs of 3 Mb and longer from whole-genome VCF and gVCF files only, and reads them as population history. It does not measure them from chip exports[12].
A run of homozygosity is a stretch of chromosome where the two copies a person inherited, one from each parent, are identical over millions of bases. Long runs were long read as a sign of marriage between relatives; genome-wide data show that runs are universally common, including in outbred people[1]. In four European populations, runs of up to 4 Mb were common in people whose pedigrees showed no shared ancestor on the two sides for at least five generations[2]. Genomes differ less in whether they have runs than in how long the runs are and how much of the genome they add up to.
Two copies, one ancestor
Each of the 22 autosomes comes in two copies, one inherited from each parent. Over most of their length the two differ here and there, at the positions where a person is heterozygous. Inside a run of homozygosity they do not: position after position reads the same letter twice. A long run forms when the same stretch of an ancestor's chromosome reaches a person down both sides of the family tree, so that they inherit one ancestral piece twice[1].
Terms this guide relies on
- Run of homozygosity (ROH)
- A stretch, usually counted from several hundred thousand bases up, where both copies of a chromosome carry the same letters at every position read.
- F(ROH)
- The share of the autosomes inside runs longer than a chosen minimum. McQuillan and colleagues named it, and found that it tracked inbreeding measured from family trees in Orkney closely, with a correlation of 0.86[2].
- Centimorgan (cM)
- A distance on a chromosome measured in recombination: 1 cM is the stretch in which a crossover happens once in a hundred meioses. On average it spans about a million bases[3].
- Meiosis
- The cell division that makes eggs and sperm, in which each chromosome is cut and rejoined between the parent's two copies.
F(ROH) is only as comparable as its minimum. A study that counts runs from 500 kb adds up far more homozygous genome than one that starts at 3 Mb, so the figures in this guide always come with the length they were counted from.
Length works as a clock
Each meiosis cuts the chromosomes and splices them. A segment that travels from one ancestor down two lines of descent is cut along both, so the further back the ancestor lived, the shorter the piece that arrives intact on both copies. Thompson gives the rule: when two copies are separated by m meioses, the lengths of the segments they share follow an exponential distribution with a mean of 100/m centimorgans. The paper's worked example is a time depth of 25 generations, which is 50 meioses and an average of 2 cM[3].
So runs of 3 Mb and up, the ones the report counts, mostly come from ancestors within the last few dozen generations, and runs under 1 Mb from ancestors much of a population shares. Kirin and colleagues put it in one sentence: offspring of cousin marriages have long runs, while the numerous shorter tracts relate to shared ancestry tens and hundreds of generations ago[4].
The averages hide a wide spread, and that limits what one genome can say. McQuillan's paper gives the expected inbreeding coefficient of a first-cousin couple's child as 0.0625 with a standard deviation of 0.0243, because which segments pass down each line is a matter of chance; the groups it compared, sorted by pedigree, overlapped considerably in F(ROH)[2]. A single genome's total is one draw from that spread.
Short runs, long runs and the histories behind them
Pemberton and colleagues mapped runs in 1,839 people from 64 populations with 577,489 markers and found three length classes. Short runs of tens of kilobases probably reflect ancient haplotypes. Intermediate runs of hundreds of kilobases to several megabases probably result from background relatedness in a population of limited size. Long runs of multiple megabases probably result from recent relatedness between a person's parents[5]. Long runs were higher and more variable in most populations from the Middle East, Central and South Asia, Oceania and the Americas than in Africa, Europe and East Asia, and higher where consanguineous marriage is common[5].
Short runs carry most of the total, even where long runs are frequent. Kirin's survey of worldwide populations found that in every one, including the most inbred, more of the genome sat in runs of 0.5 to 2 Mb than in runs over 2 Mb[4]. A total built mostly from long runs points to recent shared ancestry; one built from many intermediate runs points to a small population, and the two can sit in the same genome.
A 2018 review of how runs are found in microarray and sequence data put the point in one line:
The number and length of ROH reflect individual demographic history
Why Finnish genomes carry more
Finland's inland north and east were settled late, from the 1500s[8], and the youngest sub-isolates in the north-east were founded only 300 to 400 years ago. Jakkula and colleagues read the genetic pattern as multiple bottlenecks from consecutive founder effects[6]: each settlement drew a small sample of genomes out of an older population, and its descendants then married mostly among themselves. The Finnish DNA guide traces the same history through the Y chromosome and the east–west divide.
They measured it in 1,395 people, from ten Finnish subpopulations plus comparison samples from Helsinki and Sweden, counting runs longer than 1 Mb from 231,116 markers. Mean homozygosity rose from 0.9 percent of the genome in the early-settled south coast to 2.0 percent in the North Kainuu isolate. The number and length of runs were highest in the youngest subpopulations, where up to 90 percent of people had at least one run longer than 5 Mb; in a US sample, 9.5 percent did[6].
Martin and colleagues saw the same geography in segments shared between people rather than within one genome, which is the same inheritance seen across two people instead of two copies. In 43,254 Finns, people around Kuusamo, a late-settled area in the north-east, shared on average about 60 Mb in segments longer than 3 cM with others born nearby; around Helsinki, Turku and Tampere the figure was about 5 to 15 Mb[9]. Across countries, two unrelated Finns shared on average 107.0 cM, two Swedes 22.9 cM[9].
~60 Mb
Shared in segments over 3 cM with people born nearby, around Kuusamo[9]
~5 to 15 Mb
The same measure around Helsinki, Turku and Tampere[9]
~63.9 Mb
In runs over 1 Mb, 1000 Genomes Finnish genomes; 40 Faroese averaged ~82.5 Mb[7]
Finland is not the extreme case. In high-coverage whole genomes analysed in 2026, the Finnish samples of the 1000 Genomes Project averaged about 63.9 Mb in runs longer than 1 Mb, and 40 people from the Faroe Islands about 82.5 Mb, more than any group in the study; the Faroese also carried more runs in the 5 to 15 Mb range, which the authors read as a more recent or stronger bottleneck[7]. On chip data, endogamous island communities in Dalmatia and Orkney carried several times more than cosmopolitan samples[2].
| Population sample | Study and data | Runs counted from | Average in runs |
|---|---|---|---|
| CEU (HapMap) | McQuillan 2008, 300,000-marker chip | 1.5 Mb | 8 Mb (F 0.003) |
| Scotland, 984 people | McQuillan 2008, 300,000-marker chip | 1.5 Mb | 7 Mb (F 0.003) |
| Orkney, three or more grandparents from one island | McQuillan 2008, 300,000-marker chip | 1.5 Mb | 28 Mb (F 0.011) |
| Dalmatian island, all four grandparents from one village | McQuillan 2008, 300,000-marker chip | 1.5 Mb | 35 Mb (F 0.013) |
| Finland, early-settled south coast | Jakkula 2008, 231,116 markers | 1 Mb | 0.9% of the genome |
| Finland, North Kainuu isolate | Jakkula 2008, 231,116 markers | 1 Mb | 2.0% of the genome |
| Finland, 1000 Genomes FIN | Hamid 2026, whole genomes | 1 Mb | ~63.9 Mb |
| Faroe Islands, 40 people | Hamid 2026, whole genomes | 1 Mb | ~82.5 Mb |
Source: McQuillan et al. 2008[2]; Jakkula et al. 2008[6]; Hamid et al. 2026[7]. Rows with different minimums are not directly comparable.
The link to the Finnish Disease Heritage
A recessive condition appears when both copies of a gene carry a disease-causing change. Inside a run, both copies are the same ancestral copy, so a variant that one founder carried can arrive twice. Martin and colleagues note that long runs are enriched for deleterious variation, and they found that two Finns who carry the same variant from the Finnish disease database share a stretch of genome around it about 30 percent of the time or more, an order of magnitude above the background rate[9].
The same settlement history drew both maps. FinDis records that for most heritage diseases the birthplaces of patients' families concentrate in the late-settlement area populated from the 1500s, and that Northern epilepsy is found in Kainuu, near the eastern border[8], where Jakkula measured the highest homozygosity[6]. The Finnish Disease Heritage guide covers the 39 diseases themselves, and carrier status vs your own risk what one copy of a recessive variant means. A run says nothing about which variants sit inside it, so this module reports no carrier results; the carrier panel reads them directly.
What each kind of file can measure
Most of what is known about runs comes from chips. McQuillan used a 300,000-marker chip, Jakkula 231,116 markers and Pemberton 577,489[2, 6, 5]. The consumer chip exports we measured hold between 563,320 and 955,958 rows[10], so a chip of that size can, in research hands, find runs of a megabase and longer. Density sets the floor: with a panel of 3 million markers, runs as short as 100 kb can be detected reliably, and Han Chinese genomes that showed about 130 Mb of runs on a panel of about 400,000 markers showed about 510 Mb on the dense one[4]. A typical genome differs from the reference at 4.1 to 5.0 million sites[11], and a whole-genome VCF lists them all: by our arithmetic, more than a thousand per megabase on average.
Our report does not measure runs on chips. Its scan counts heterozygous calls in each megabase of a whole-genome call set, and it was built and checked on that kind of file only. The report says so in place of a number. Exomes, targeted panels and low-pass or imputed genomes have the same problem in another shape, and the report refuses them too, naming the evidence it found.
| File | What the report does | Why |
|---|---|---|
| Chip export, such as 23andMe or AncestryDNA | Not measured | The scan is built for whole-genome call sets and is not run on chip exports |
| Plain VCF from a whole genome | Measured | The variant list carries every heterozygous call |
| gVCF from a whole genome | Measured | Its variant calls are counted; reference blocks are skipped |
| Exome or targeted-panel VCF | Not measured | The stretches between captured regions were never sequenced |
| Low-pass or imputed genome | Not measured | Uncalled heterozygous positions would read as homozygous |
| Any VCF under 300 variants per megabase | Not measured | Too sparse to tell quiet from unexamined |
| BAM or CRAM (Deep Read) | From the variant file uploaded with it | The aligned-reads slice is not scanned for runs |
Source: Aimosti's report as of October 2026; the coverage page lists every module by file type[12].
What the report shows
In the report the module follows the ancestry section, under the heading Runs of homozygosity. The scan splits the autosomes into 1 Mb bins, treats a bin as homozygous when it holds at least 50 variant calls of which at most 2 are heterozygous, joins neighbouring bins, and keeps runs of 3 Mb or more. A bin with too few calls counts as unmeasured, never as homozygous. The total, divided by the 2,875 Mb of autosome the scan counts, is the share the report gives. The Finnish DNA guide draws the scan on an invented example.
The card leads with a band headline and three figures: the share of the autosomes in long homozygous stretches, to one decimal; the number of stretches of 3 Mb or more; and the longest stretch. Below them come the band's description, a short account of what runs mean and a caveat. The report's summary list shows the headline alone, never the percentage. A refused file gets the headline Not measured for this file and the reason.
| Share of autosomes | Headline | Description shown |
|---|---|---|
| Under 1% | Few long homozygous stretches | Under one percent of your autosomes sits in a run of this length. This is the ordinary result across most of Europe. |
| 1% to 3% | A moderate share, common in Finland | Between one and three percent. This is the range a Finnish genome commonly falls in, and it reflects the size of the population your ancestors came from rather than anything about your immediate family. |
| 3% and over | A larger share than most European genomes carry | Above three percent. Ancestry from a small or long-separated population, ancestors who married within one community, and closer family relationships all produce this pattern, and a genome alone does not say which applies. A clinical genetics service is where that question belongs. |
Source: The report's ROH module, content version of 2026-09-03; its sources are the Ceballos, McQuillan and Martin papers cited here[1, 2, 9].
The band edges are round numbers chosen for reading, not clinical thresholds. They also sit on a different scale from the published Finnish figures: Jakkula counted runs from 1 Mb, the report from 3 Mb in whole 1 Mb bins, so the same genome usually shows a smaller share in the report than it would in that study[6]. The caveat on every card names descent from a founder population, marriage within a small community and recent shared ancestry together, and says the measurement does not distinguish between them.
What Aimosti would (and wouldn't) show you
From a whole-genome VCF or gVCF the report counts runs of 3 Mb and longer, gives their share of the autosomes, the number of runs and the longest, and places the share in one of three bands, read as population history. Chip exports, exomes, low-pass or imputed genomes and files under 300 variants per megabase get a sentence saying why nothing was measured.
What we won't claim
We won't read a run-of-homozygosity figure as a statement about anyone's parents or family, give it a health meaning, or treat a band boundary as a clinical threshold. One total cannot separate a founder population from recent shared ancestry, and the report does not try.
Bottom line. Runs of homozygosity are a record of how small and how separate the population behind a genome was, written in lengths. Finnish genomes carry the marks of a few small founding groups, most of all in the north and east. A single total mixes old history with recent history, which is why the report gives it as population history and nothing more.
Questions people ask
What do runs of homozygosity mean?
Stretches where the two copies of a chromosome are identical because both descend from one ancestor. Short runs go back to ancestors a population shares; long runs to ancestors within the last few dozen generations[4, 3].
Can 23andMe or AncestryDNA raw data show runs of homozygosity?
Research studies have measured runs of 1 Mb and longer from chips with 231,000 to 577,000 markers, and consumer exports hold 563,320 to 955,958 rows[6, 5, 10]. Our report does not measure runs from chip exports; it needs a whole-genome VCF or gVCF.
Which files does Aimosti measure runs of homozygosity from?
A whole-genome VCF or gVCF with at least 300 variants per megabase. Chip exports, exomes, targeted panels and low-pass or imputed genomes get a sentence saying why nothing was measured, and with Deep Read the variant file uploaded with the reads is the one scanned[12].
Why do Finns have more runs of homozygosity than other Europeans?
Because Finland was settled by small founding groups, the inland north and east only from the 1500s, and those communities stayed comparatively separate. Runs over 1 Mb covered 0.9 percent of the genome on the south coast and 2.0 percent in North Kainuu, and up to 90 percent of people in the youngest subpopulations had a run over 5 Mb, against 9.5 percent of a US sample[6].
References
- Ceballos FC, Joshi PK, Clark DW, Ramsay M, Wilson JF. Runs of homozygosity: windows into population history and trait architecture. Nature Reviews Genetics, 2018. doi:10.1038/nrg.2017.109 Cited for statements in its abstract, the part read in full; the review itself is not open access.
- McQuillan R, Leutenegger AL, Abdel-Rahman R, et al. Runs of homozygosity in European populations. American Journal of Human Genetics, 2008. doi:10.1016/j.ajhg.2008.08.007 2,618 people genotyped on the Illumina HumanHap300 chip. Mean F(ROH) with a 1.5 Mb minimum: endogamous Dalmatians 0.013 (35 Mb), endogamous Orcadians 0.011 (28 Mb), CEU 0.003 (8 Mb), Scottish 0.003 (7 Mb).
- Thompson EA. Identity by descent: variation in meiosis, across genomes, and in populations. Genetics, 2013. doi:10.1534/genetics.112.148825
- Kirin M, McQuillan R, Franklin CS, Campbell H, McKeigue PM, Wilson JF. Genomic runs of homozygosity record population history and consanguinity. PLoS ONE, 2010. doi:10.1371/journal.pone.0013996
- Pemberton TJ, Absher D, Feldman MW, Myers RM, Rosenberg NA, Li JZ. Genomic patterns of homozygosity in worldwide human populations. American Journal of Human Genetics, 2012. doi:10.1016/j.ajhg.2012.06.014
- Jakkula E, Rehnström K, Varilo T, et al. The genome-wide patterns of variation expose significant substructure in a founder population. American Journal of Human Genetics, 2008. doi:10.1016/j.ajhg.2008.11.005 1,395 people, 231,116 SNPs; runs longer than 1 Mb. Mean homozygosity 0.9 percent in ESS to 2.0 percent in ISC.
- Hamid I, Mortensen Ó, Refoyo-Martínez A, et al. Faroese whole genomes provide insight into ancestry and recent selection. eLife, 2026. doi:10.7554/eLife.107428 Average genome in runs longer than 1 Mb: about 82.5 Mb in 40 Faroese genomes, about 63.9 Mb in the 1000 Genomes Finnish (FIN) genomes.
- The Finnish Disease Heritage. FinDis, Finnish Disease Database, 2026.
- Martin AR, Karczewski KJ, Kerminen S, et al. Haplotype sharing provides insights into fine-scale population history and disease in Finland. American Journal of Human Genetics, 2018. doi:10.1016/j.ajhg.2018.03.003
- What your DNA file can read. Aimosti, 2026. Ten chip exports measured on 2026-10-06, holding 563,320 to 955,958 rows.
- The 1000 Genomes Project Consortium. A global reference for human genetic variation. Nature, 2015. doi:10.1038/nature15393
- Coverage: what each file type can tell you. Aimosti, 2026.
Last reviewed . Every number on this page links to the source it comes from; if one of them has moved, tell us.