Learn Explainer
Blood type from DNA: what a raw data file can and cannot say
Most ABO blood groups come down to three letters in one gene, so a genome file can usually name A, B or AB. Group O is harder, because the reference genome is itself an O, and the Rh factor harder still, because Rh-negative usually means a missing gene. This is what the letters say, how often they match a blood test, and what our report reads from each file.
Key takeaways
- The common O allele is a single missing G at position 261 of ABO (rs8176719); A and B differ at rs8176746, rs8176747 and a few other positions[1, 3].
- The reference genome carries the O deletion, so a group O genome has no line for it in a VCF, and that silence looks like a position never read[2].
- A whole-genome typing algorithm matched serology for ABO in all 200 genomes of a blinded test[7].
- RhD-negative is usually a deletion of the whole RHD gene, which sits beside a near-identical twin, RHCE[13, 14].
- In Finland 41 percent of people are group A, 33 percent O, 18 percent B and 8 percent AB, and 14 percent RhD-negative[15].
Your ABO group is decided by one gene, ABO, which makes an enzyme that adds a sugar to a molecule on red cells called the H antigen. The A version of the enzyme adds one sugar, the B version another, and an O allele makes no working enzyme at all. Fumiichiro Yamamoto's group found the reason in 1990: the A and B alleles differ by a few single-letter changes that alter four amino acids, and the common O allele has lost one letter, which shifts the reading frame and produces an entirely different, inactive protein[1]. Three positions in a DNA file carry most of that information: two tell A from B, and one tells a working allele from the O deletion.
Three positions in one gene
The O deletion is a G missing at position 261 of the gene's coding sequence, written c.261delG, and its dbSNP name is rs8176719[2]. Most O alleles carry it. O01 is the A1 sequence with that one letter gone; O02 has the same deletion plus other changes; and a rarer O allele, O03, keeps the G and is switched off instead by a change at codon 268[3].
A and B differ at codons 176, 235, 266 and 268. The A enzyme has arginine, glycine, leucine and glycine there; the B enzyme has glycine, serine, methionine and alanine[3]. Two of those changes, c.796C>A (rs8176746) and c.803G>C (rs8176747), sit seven letters apart. Our report reads both, and UK Biobank's blood-type field reads the first[4, 5, 6].
One detail decides how all of this looks in a file. The reference genome GRCh38 has the O01 letters at these positions, so at rs8176719 the reference already lacks the G[2]. In a VCF, which lists only where a genome differs from the reference, a working A or B allele therefore appears as an insertion, and the O allele appears as nothing at all. The gene also lies on the reverse strand of chromosome 9, so the letters a VCF shows are the complements of the gene's own: c.796C>A is a G to T change in the file, and c.803G>C is C to G[4, 5].
| rsID | Change in the gene | In a GRCh38 VCF | What the change means |
|---|---|---|---|
| rs8176719 | c.261delG (O deletion) | insertion of C at chr9:133,257,521 | insertion: a working A or B copy |
| rs8176746 | c.796C>A | G>T at chr9:133,255,935 | T is a B copy |
| rs8176747 | c.803G>C | C>G at chr9:133,255,928 | G is a B copy |
Source: dbSNP and Ensembl, GRCh38, read 10 October 2026[2, 4, 5].
From letters to a group
Reading a group from the three positions is counting. Each copy of the insertion at rs8176719 is a working allele; each T at rs8176746, confirmed by a G at rs8176747, makes one of them a B, and the rest are A. A and B both show when a person has one of each, and O shows only when neither is present, so the counts map onto a group[1].
| Working copies (insertion at rs8176719) | B copies (rs8176746 and rs8176747) | Likely group | Allele pair |
|---|---|---|---|
| 0 | 0 | O | O/O |
| 1 | 0 | A | A/O |
| 2 | 0 | A | A/A |
| 1 | 1 | B | B/O |
| 2 | 2 | B | B/B |
| 2 | 1 | AB | A/B |
Source: The counting rules follow from the allele structure in Yamamoto's papers[1, 3]; they are the rules our report applies.
The counting hides an assumption. The O position sits 1,586 bases from the B positions, and a typical sequencing fragment is about 300 bases long[7], so the file does not say whether the B change sits on the same copy as a working allele or on the same copy as an O deletion. The counting assumes the common arrangement, B changes on a working copy, and that is right for the alleles above. Lane and colleagues ran into the same problem when they built a whole-genome typing algorithm. In its second test, six of the ten wrong calls across all antigens came from not knowing which changes shared a chromosome, and their final algorithm infers the ABO arrangement from known allele frequencies[7].
VCF
#CHROM POS ID REF ALT GT
9 133255928 rs8176747 C G 0/1
9 133255935 rs8176746 G T 0/1
9 133257521 rs8176719 T TC 0/1Read as a worked example, those three records are one working copy, which is a B, and one O deletion: likely group B, allele pair B/O. With 1/1 at the insertion and the same B markers there would be two working copies, one of them B, and the likely group would be AB.
A1, A2 and the alleles three markers miss
Group A comes in two common subgroups. The A2 allele has lost a single C near the end of the gene, which moves the stop signal and adds an extra stretch to the enzyme; putting that deletion into an A1 enzyme cut its activity sharply in Yamamoto's experiments[9, 3]. Yamamoto's papers number it 1060delC and current reference sequences call it c.1061del (rs56392308); the letter sits in a run of C's, so both names describe the same deletion[10]. None of the three markers above can see it.
Finland carries both the B change and the A2 deletion more often than the rest of Europe. In gnomAD's 5,270 Finnish genomes the B change is on 13.6 percent of chromosomes and the A2 deletion on 11.5 percent, against 8.0 and 6.7 percent in non-Finnish Europeans, and the O deletion is correspondingly rarer[11, 12, 10]. By our arithmetic from the same counts, about a third of the working A copies in gnomAD's Finns are A2, against about a quarter in other Europeans: an estimate that assumes every A2 deletion sits on an A copy.
Beyond A2 there is a long tail: A3, Ax, weak A and weak B alleles, and cis-AB and B(A), single alleles whose hybrid enzyme makes both antigens[3]. O03, the O allele without the deletion, reads as a working A copy to a three-marker method. Each is uncommon; together they are why a genotype prediction carries a small error rate even from a perfect file.
RhD is a missing gene
The plus or minus after a blood group is the RhD antigen, and its genetics work differently. The Rh antigens come from two neighbouring genes, RHD and RHCE. In white populations about 40 percent of haplotypes carry a deletion of the whole RHD gene, and in Europeans that deletion is the commonest reason a person is RhD-negative[13, 14]. If 40 percent of haplotypes carried it, about 16 percent of people would have two copies; that is our arithmetic, and the Blood Service's figures put 14 percent of Finns in the RhD-negative groups[15].
That structure is hard for genotype data. A chip export or a VCF lists letters at positions, and when the gene is gone there are no RHD letters to list. RHD and RHCE are also similar enough that probes or reads meant for one pick up the other. The team that built a blood-donor genotyping array found its first version 'did not detect the RHD copy number accurately', and fixed it by tiling RHD with 114 probe sets at the places least like RHCE[14]. Whole-genome reads can see the deletion as a stretch with no coverage: read-depth analysis called the D antigen correctly in all 110 genomes of one study and in all 200 of its blinded test[7].
How often the genotype matches a blood test
The comparisons use serology, the antibody test a blood service runs, as the standard. The clearest is the bloodTyper study, whose algorithm was improved twice against genomes from the MedSeq Project, then tested blind on 200 genomes from INTERVAL, a study of UK blood donors, sequenced at about 15x. The improved version matched serology for ABO in 98 percent of 90 people, and the final version in all 200[7]. ABO subgroups and hybrid alleles were not tested[7].
200 of 200
genomes whose ABO group a whole-genome algorithm matched to serology in a blinded test[7]
7
of 23 array or software errors in a study of about 8,000 donors that involved ABO[14]
Arrays are the harder case. The donor-typing array of Gleadall and colleagues agreed with clinical typing in 99.91 percent of 89,371 red-cell antigen comparisons in 7,984 donors, yet seven of the 23 discordances caused by array or software errors involved ABO, and the authors write that antibody-based typing of every donation remains international best practice[14]. A pancreatic cancer study that inferred groups from two array SNPs, rs505922 and rs8176746, matched people's own reported groups 92 percent of the time. Its authors judged that within the error of self-report, since in the same cohorts reported groups had matched laboratory typing in only 91 percent[17].
| Study | Data | Compared with | ABO agreement |
|---|---|---|---|
| Lane et al. 2018, final algorithm | 200 whole genomes, about 15x | Serology | 200 of 200 |
| Lane et al. 2018, improved algorithm | 90 whole genomes, 30x | Serology | 98% |
| Giollo et al. 2015, BOOGIE | 71 Personal Genome Project genomes | Serology, as Lane et al. report it | 94% |
| Wolpin et al. 2010 | Two SNPs on a genotyping array, 187 people | Self-reported group | 92% |
Source: Lane and colleagues report the Giollo figure for ABO with its sample size[7, 16, 17].
Finland's groups, counted two ways
| Group | RhD-positive | RhD-negative | All |
|---|---|---|---|
| A | 35 | 6 | 41 |
| O | 28 | 5 | 33 |
| B | 16 | 2 | 18 |
| AB | 7 | 1 | 8 |
| All groups | 86 | 14 | 100 |
Source: Finnish Red Cross Blood Service, page updated 7 July 2025; the totals are our sums of its eight figures[15].
A+ is the commonest group in Finland and AB- the rarest[15]. The gnomAD genomes give a second count of the same thing. Of 5,270 Finnish genomes, 27.8 percent have no copy of the insertion at rs8176719, the genotype that means O by the three-marker method; in 26,684 Finnish exomes the figure is 29.2 percent[11]. Both are below the Blood Service's 33 percent group O. The two sources count different people in different ways, and neither explains the gap. O alleles without the deletion, such as O03, would push the genotype count down, but we found no measurement of how common they are in Finland.
What our report reads, file by file
The report's ABO card reads the three positions in the first table and applies the counting rules above. It labels every group as likely and not a blood test, and gives no group when the markers disagree, when there are more B copies than working copies, or when a rare third letter turns up at a B position.
| File | What the card shows | Why |
|---|---|---|
| Chip export (23andMe, AncestryDNA, MyHeritage and others) | Group not determined: the markers were not examined | Our chip reader cannot use the insertion, and treats rs8176747, a C to G change, as strand-ambiguous |
| Plain VCF | Likely A, B or AB when the file has the insertion; otherwise group not determined, with a note that the reading points to O | The O deletion matches the reference, so it never has a line of its own |
| gVCF | The same as a plain VCF | The card does not use reference blocks here |
| BAM or CRAM (Deep Read) | No ABO reading from the reads | Deep Read covers pharmacogenes, deletions and HLA; ABO comes from the VCF or gVCF |
Source: The report's ABO module as of 10 October 2026. The VCF rows follow from the reference genome carrying the O deletion[2].
One limit follows from how our report reads files. A group O genome has no line at rs8176719 in any VCF, so the card never confirms O and shows a group not determined; in gnomAD's Finns that is the 27.8 percent with no insertion[11]. The card finds the three positions by their place on chromosome 9 as well as by their dbSNP identifiers, so a VCF whose ID column is empty, written ., is read the same as one that names them.
A1 against A2, the rarer subgroups, cis-AB and RhD are not read from any file, and an O03 allele reads as a working A copy[3]. The couple outlook works out a child's possible groups only when both partners' cards name a group, so an undetermined group on either side means no ABO chances are shown. Our guide to VCF, gVCF, BAM and CRAM explains what each file keeps, chip exports have other limits we have measured, and the free file check names what you have before anything is uploaded.
What Aimosti would (and wouldn't) show you
From a VCF or gVCF that records the functional allele at rs8176719, the report gives a likely A, B or AB, labelled as an inference and never as a blood test. A genome without that allele matches the reference, which is O, and the card shows a group not determined. From a chip export the markers are not examined, and a Deep Read of a BAM or CRAM does not read ABO. RhD, A1 against A2 and the rarer subgroups are not read from any file.
What we won't claim
We won't present a DNA-derived blood group as a blood test, say anything about transfusion, donation or compatibility, or confirm a group O that the file only implies. A genome that matches the reference at the O deletion is headed group not determined, with a note that the reading points to O.
Bottom line. A genome file names A, B and AB well, and the best whole-genome method matched a blood test in all 200 people of its blinded test. Group O is the case a variant file cannot prove, because the O deletion is what the reference already says. RhD is mostly a missing gene, which a list of single positions cannot show.
Questions people ask
Can DNA tell your blood type?
Usually, for ABO: a whole-genome typing algorithm matched serology for ABO in all 200 genomes of a blinded test[7]. Rare subgroups, hybrid alleles and O alleles without the usual deletion can still mislead a simple method[3]. RhD is harder, because RhD-negative usually means the whole RHD gene is missing[13].
Can I find my blood type in 23andMe or AncestryDNA raw data?
Our report does not give an ABO group from any chip export: our chip reader cannot use the O-deletion insertion and treats one of the two B positions as strand-ambiguous. Some studies have used rs505922 instead, a marker near ABO that usually travels with the O deletion[17]. UK Biobank used it only when rs8176719 had no result[6], and in the Finnish sample of the 1000 Genomes Project its r squared with the deletion is 0.89, so it is a good proxy rather than the deletion itself[18].
My whole-genome report says group not determined. Does that mean O?
It often does, and the card says the reading points to O, but the file cannot prove it. The reference genome carries the O deletion, so a group O genome has no line at rs8176719, and a VCF cannot show whether a missing line was read as the reference or not read at all[2, 8].
Why is my remembered blood type different from the DNA result?
Either can be wrong. In one pair of cohorts, people's own reported groups matched laboratory typing 91 percent of the time[17], and the three-marker method misses A2, weak subgroups, cis-AB and O alleles without the deletion[3]. Only serology measures the antigens on the red cells themselves.
References
- Yamamoto F, Clausen H, White T, Marken J, Hakomori S. Molecular genetic basis of the histo-blood group ABO system. Nature, 1990. doi:10.1038/345229a0
- rs8176719 (ABO 261delG). NCBI dbSNP. GRCh38 chr9:133257521; the reference sequence lacks the G, so the functional allele is an insertion of C on the forward strand.
- Yamamoto F. Molecular genetics and genomics of the ABO blood group system. Annals of Blood, 2021. doi:10.21037/aob-20-71
- Variant rs8176746 (ABO c.796C>A, p.Leu266Met). Ensembl. GRCh38 chr9:133255935 G>T on the forward strand; NM_020469.4:c.796C>A. Read 10 October 2026.
- Variant rs8176747 (ABO c.803G>C, p.Gly268Ala). Ensembl. GRCh38 chr9:133255928 C>G on the forward strand; NM_020469.4:c.803G>C. Read 10 October 2026.
- Data-Field 23165: Blood type. UK Biobank Showcase. Imputed from rs505922, rs8176719 and rs8176746; rs505922 T used for type O when rs8176719 had no result.
- Lane WJ, Westhoff CM, Gleadall NS, et al. Automated typing of red blood cell and platelet antigens: a whole-genome sequencing study. The Lancet Haematology, 2018. doi:10.1016/S2352-3026(18)30053-X
- The Variant Call Format Specification, VCFv4.3. GA4GH / hts-specs, 2025.
- Yamamoto F, McNeill PD, Hakomori S. Human histo-blood group A2 transferase coded by A2 allele, one of the A subtypes, is characterized by a single base deletion in the coding sequence, which results in an additional domain at the carboxyl terminal. Biochemical and Biophysical Research Communications, 1992. doi:10.1016/s0006-291x(05)81502-5
- gnomAD v4.1: variant 9-133255669-CG-C (rs56392308, ABO c.1061del). Genome Aggregation Database, 2024. Genomes, read 10 October 2026: Finnish 1,222 of 10,602 alleles; non-Finnish European 4,564 of 67,988.
- gnomAD v4.1: variant 9-133257521-T-TC (rs8176719). Genome Aggregation Database, 2024. Genomes, read 10 October 2026: Finnish 5,023 of 10,540 alleles, 1,216 homozygotes; non-Finnish European 24,770 of 67,920. Exomes, Finnish: 24,720 of 53,368, 5,815 homozygotes.
- gnomAD v4.1: variant 9-133255935-G-T (rs8176746). Genome Aggregation Database, 2024. Genomes, read 10 October 2026: Finnish 1,438 of 10,580 alleles; non-Finnish European 5,403 of 67,940.
- Wagner FF, Flegel WA. RHD gene deletion occurred in the Rhesus box. Blood, 2000. doi:10.1182/blood.V95.12.3662
- Gleadall NS, Veldhuisen B, Gollub J, et al. Development and validation of a universal blood donor genotyping platform: a multinational prospective study. Blood Advances, 2020. doi:10.1182/bloodadvances.2020001894
- Tietoa veriryhmistä (blood groups in Finland). Finnish Red Cross Blood Service, 2025. Last updated 7 July 2025, read 10 October 2026: A+ 35%, O+ 28%, B+ 16%, AB+ 7%, A- 6%, O- 5%, B- 2%, AB- 1%.
- Giollo M, Minervini G, Scalzotto M, Leonardi E, Ferrari C, Tosatto SC. BOOGIE: predicting blood groups from high throughput sequencing data. PLoS One, 2015. doi:10.1371/journal.pone.0124579
- Wolpin BM, Kraft P, Gross M, et al. Pancreatic cancer risk and ABO blood group alleles: results from the pancreatic cancer cohort consortium. Cancer Research, 2010. doi:10.1158/0008-5472.CAN-09-2993
- Linkage disequilibrium between rs505922 and rs8176719, 1000 Genomes phase 3. Ensembl REST API. r2 0.885 in the Finnish sample (FIN) and 0.958 in CEU, read 10 October 2026.
Last reviewed . Every number on this page links to the source it comes from; if one of them has moved, tell us.
Free · in your browser
Find out what your own file can read.
The free check opens your DNA file on your device and shows which format you have and what a report could read from it. The file is never uploaded.