Under the hood
The Prothrombin row in every 23andMe file, and why we weren't reading it
Every 23andMe export we tested carries a call for Prothrombin G20210A, one of the two common inherited clotting variants. Until 28 September our reports said it was not examined. This is how that happened and what we changed.
Prothrombin G20210A (rs1799963, in the F2 gene) makes the body produce slightly more prothrombin, which modestly raises the chance of a blood clot. Two to five percent of Americans of European ancestry carry it. It is one of the few clinically meaningful variants that a consumer genotyping chip types reliably, so a service that re-reads chip files has no excuse for dropping it.
Two ways to find a row
A raw-data file from 23andMe is a long text table: an identifier, a chromosome, a position and two letters for each probe. Our report is built on a panel of positions written against GRCh38, the current human reference assembly. 23andMe files are written against GRCh37, the version before it. To find a panel position in a customer's file we look for its rsID first, and if that is missing we look for its coordinate.
23andMe does not label every probe with an rsID. Some carry the company's own identifiers, which start with an i. Prothrombin is one of them: in every 23andMe export we have (versions 3, 4 and 5) it appears as i3002432, at GRCh37 chromosome 11, position 46,761,055. The rsID lookup finds nothing, so everything rests on the coordinate.
Why the coordinate lookup was switched off
In August we found that the coordinate fallback had been comparing GRCh38 positions against GRCh37 files. The same number points at a different place in the two assemblies. On the Y chromosome this meant a man's paternal-line markers were never found, and the note we showed him offered "this is expected if you are biologically female" as the first explanation.
It was worse elsewhere. For any panel rsID missing from a file, the fallback could land on an unrelated SNP at the same number, and the only thing that stopped a made-up genotype was whether that SNP's letters happened not to fit. So we stopped translating coordinates on GRCh37 files, except on the Y chromosome, which goes through a table we had verified, and on mitochondrial DNA, which is numbered the same way in both assemblies.
That refusal was right, and it is exactly why the Prothrombin row went unread. The file offered a GRCh37 coordinate and we would not translate one.
A translation we would accept
The fix is a short table of GRCh37 positions for panel sites that a real export ships under a vendor identifier. We did not guess what belongs in it. We looked up all 1,005 panel rsIDs in Ensembl's GRCh37 service, then checked each 23andMe file for a call at that GRCh37 position under some other identifier. Four sites turned up: Prothrombin, present in all four files, and three sites used by polygenic scores.
A matching coordinate is not proof that two rows describe the same place. For each candidate we took 41 bases centred on the position from each assembly and required the 20 bases on either side to be identical. A single matching base is a one-in-four coincidence; forty matching bases are not. Three candidates passed. The fourth, rs9411377 (23andMe's i708994), sits in a run of repeated A's whose flanks differ between the assemblies, so we left it out. If we cannot prove a mapping, we don't use it.
A second bug on the way
While measuring, we found positions with more than one row. 23andMe includes insertion and deletion probes, reported as I and D, alongside the ordinary probe at the same coordinate, and our reader had been keeping whichever row came last. It now sets the I/D rows aside, and if the rows that remain still disagree, it leaves the position unread. Before changing the rule we checked that none of the sites we read by coordinate sat on a disagreeing row in any of our test files.
What changed on real files
We ran our 14 real chip exports through the reader before and after the change and compared every reading. Two things moved: Prothrombin is now read on all four 23andMe files, and two polygenic-score sites are now read on the version 4 file. Nothing else changed, which is what a narrow fix should look like.
Reports issued before 28 September still show Prothrombin as not examined. If you have one, write to us and we will run it again.
What we take from it
"Not examined" is an honest label, and for seven weeks it hid a gap that was ours, not the file's. A refusal deserves the same checking as a finding. When the report says it could not read something, the question to ask is whether the file actually lacked it.
What Aimosti would (and wouldn't) show you
On a 23andMe file the report now reads the Prothrombin row and shows the result on its card, with the source. Where we still cannot prove that a position in the file is the one on our panel, the card says "not examined" rather than guessing.
What we won't claim
We won't read a position through a coordinate translation we have not checked base by base, and we won't present a Prothrombin result as a diagnosis. It is a common risk factor, and most people who carry one copy never have a clot.
Bottom line. The call was in the file the whole time. A rule we added to stop wrong answers also stopped this right one, and the fix had to meet the same standard of proof as the rule.
Related: Clinical findings from a chip file. Restated from: MedlinePlus Genetics: Prothrombin thrombophilia · dbSNP: rs1799963 · Genome Reference Consortium: human assemblies · Ensembl REST API (GRCh37 and GRCh38 lookups).