what the file says about me
Part one ended with a deflation: 23andMe never read my DNA. What I own is 643,161 multiple-choice answers, inferred from brightness measurements at a curated shortlist of positions, 0.021 percent of a genome. The interesting question was never what the file contains. It is how far 643,161 answers actually reach.
Because they reach absurdly far. From that file, 23andMe told me my eye color, painted an ancestry onto every chromosome I have, named two ancestors' lineages, and ruled on milk. This post checks all of it in public: five famous positions against what my body actually does, the ancestry painting segment by segment, a second and mechanically different answer to where I am from, the two lines of inheritance that never shuffle, and the letters in my export that were never measured at all. It ends with the part I chose not to publish.
01 · five positions, checked
One honest thing first. 23andMe's reports do not all hang off single positions; the company's own methodology paper describes health reports built as statistical models over many variants. But a handful of positions are famous enough to carry a trait's headline, each with its own literature, and all five of the classics are in my file. For each one: my verbatim row, what the literature associates with each allele (each version of the letter that position can carry), what 23andMe's report concluded, and what my body says.
Start with milk, at the position part one used for the probe demo. My row is rs4988235 2 136608646 AG. The association was nailed down in 2002, in Finnish families: a variant upstream of the lactase gene that completely associated with biochemically verified lactase non-persistence, with non-persistence behaving as a recessive condition. Recessive non-persistence means persistence is dominant: one copy of the persistence allele is enough to keep digesting milk as an adult. In my file's plus-strand alphabet the persistence allele is A, which is also the allele ClinVar records for the lactase persistence trait. My AG carries one A. Prediction: probably lactase persistent. The report says "Likely tolerant." The fridge agrees; milk has never bothered me.
The population caveat matters here more than anywhere. That mapping comes from European-ancestry cohorts, and lactase persistence evolved more than once: East African populations carry three different persistence variants, which arose on different inherited backgrounds, and a distinct compound allele reached high prevalence in Middle Eastern populations, in both cases without the European variant. Read rs4988235 alone in someone with those ancestries and you can call a milk-drinker intolerant. Single positions do not travel between populations.
Earwax next, and this one is almost embarrassingly clean. My row is rs17822931 16 48258198 CC, a coding change in the gene ABCC11. The 2006 paper is blunt in a way genetics papers rarely get to be: this SNP is responsible for determination of earwax type, with the dry type recessive. In plus-strand letters, the T allele produces the amino-acid change that goes with dry earwax, so dry needs TT and my CC is two wet-type alleles. Report: "Likely wet earwax." Confirmed, at the cost of some dignity. The same paper maps the dry allele's steep geographic gradient, most frequent in East Asian populations, which is why the base rates of each answer depend on where your ancestors lived.
Third, bitter taste. My row is rs713598 7 141673345 CG, in the bitter receptor gene TAS2R38. The 2003 discovery found that haplotypes of three coding changes in this one gene completely explain the famous bimodal distribution of PTC taste sensitivity, accounting for 55 to 85 percent of the variance. At this site, the G allele encodes the taster form of the receptor, tasting behaves largely dominant, and ClinVar lists G as the allele associated with phenylthiocarbamide tasting. My CG carries one taster copy; the report says "Likely can taste." And here I have to break the pattern: I have never chewed a PTC strip, so for this trait the report is a prediction I have not tested. I am told the experience is memorable. Worth knowing: both alleles are common worldwide, close to 50/50, an unusually balanced polymorphism, and the single site is a proxy for the full five-haplotype system.
Fourth, eyes. My row is rs12913832 15 28365618 AG, and the biology is a small plot twist: the variant is not in a pigment gene. It sits in an intron of HERC2, a stretch of the gene that does not code for protein, in a conserved regulatory element that dials down the promoter of the neighboring pigment gene OCA2. The effect size is unusually large for a complex trait: in a European-descent cohort, people with two blue-associated copies had a 1 percent probability of brown eyes, people with two brown-associated copies an 80 percent probability, and this single SNP predicted eye color better than the previous multi-marker approach. Translated to my file's strand, G associates with blue, A with brown, and blue behaves essentially recessive. My AG makes brown or hazel likely; the report says "Likely brown or hazel eyes"; the mirror says brown. The hedge is still load-bearing: eye color is a quantitative, multifactorial trait, heterozygotes like me span green, hazel, and brown, and the tidy percentages are cohort numbers, not personal guarantees.
Fifth, skin, which I will treat with the care it deserves. My row is rs1426654 15 48426484 AA, a coding change in SLC24A5 found through, of all things, a zebrafish pigmentation mutant. The derived A allele, the newer of the two versions, is nearly fixed in European populations and correlates with lighter skin pigmentation in admixed populations, while the ancestral allele predominates in African and East Asian populations. My AA is two derived copies, and the report concludes "Likely lighter skin." Note the phrasing: lighter, relative to average. Pigmentation is highly polygenic, this locus explains a share of the variation and not the whole, and a relative statement is all one association supports.
The whole exercise, compacted:
| trait | my row | allele mapping (plus strand) | 23andMe report | checked against my body |
|---|---|---|---|---|
| milk | rs4988235 AG | A goes with lactase persistence | "Likely tolerant" | agrees |
| earwax | rs17822931 CC | T goes with dry, so CC is wet | "Likely wet earwax" | agrees |
| bitter taste | rs713598 CG | G encodes the taster form | "Likely can taste" | untested |
| eye color | rs12913832 AG | G with blue, A with brown | "Likely brown or hazel eyes" | brown |
| skin | rs1426654 AA | A with lighter pigmentation | "Likely lighter skin" | agrees |
Five for five on direction, four of them checked against my actual body. Before that impresses you, remember the sample is rigged twice over: these are the best-replicated single positions in the entire trait literature, selected after the fact, and every direction I quoted is an association measured in a specific population, not a mechanism reading my cells. 23andMe's own pages put it plainly: the raw data is suitable only for informational use and not for medical, diagnostic, or other use. So the figure below inverts the order. For each trait, answer for yourself first, then open my row. Your answer is a measurement. The row only shifts odds, and if the two ever disagree for you, the measurement wins.
02 · the painting
The trait rows are parlor tricks next to the ancestry report. Mine says: 73.5 percent European, 44.6 points of it Spanish and Portuguese and 14.4 Italian, then 22.7 percent Indigenous American, 1.4 percent Sub-Saharan African, 1.4 percent Western Asian and North African, and 1.0 percent the model declined to assign. I grew up on the Río de la Plata, and this is the region's founding story rendered as a pie chart: Spanish and Italian immigration layered over Indigenous American lineages, with smaller African and Levantine threads. Nothing in it surprised my family. What surprised me is where the numbers come from.
A caution before the map below. Those labels are not countries I have papers from. They are names of reference populations: the model compares windows of my genome against 45 populations, built from a panel of more than 14,000 people of known ancestry, and reports whichever ones fit. So the markers show where the reference panels sit, not where anyone in my family was born.
The same numbers as a tree, the way the report states them:
| region | percent |
|---|---|
| European | 73.5 |
| · Spanish & Portuguese | 44.6 |
| · Italian | 14.4 |
| · Greek & Balkan | 1.7 |
| · Broadly Southern European | 7.3 |
| · French & German | 2.8 |
| · Broadly Northwestern European | 1.2 |
| · Broadly European | 1.5 |
| Indigenous American | 22.7 |
| Sub-Saharan African | 1.4 |
| · Angolan & Congolese | 1.2 |
| · Senegambian & Guinean | 0.2 |
| Western Asian & North African | 1.4 |
| · North African | 0.7 |
| · Peninsular Arab | 0.6 |
| · Broadly Western Asian & North African | 0.1 |
| Unassigned | 1.0 |
Because the percentages are the summary, not the output. The real output is segments. The export includes a file that assigns an ancestry to stretches of each chromosome copy: this arm of my chromosome 9's second copy reads Southern European (filed under European's color in the painting below), this 36-megabase span of my X reads Indigenous American, and so on, with borders at specific build-37 coordinates. The percentages are just these lengths, added up. Which means my ancestry is not a property of me so much as a property of my chromosomes, one painted stretch at a time, and the painting is the report worth looking at.
How does a pile of multiple-choice answers become painted chromosomes? 23andMe's pipeline is published, and it is a lovely piece of applied statistics. First the genotypes are phased, with an in-house adaptation of the program EAGLE, into two haplotypes per chromosome (more on phasing in section 05, it is its own confession). Each haplotype is chopped into windows of about 300 markers, about 1,800 windows genome-wide. A string-kernel support vector machine then classifies each window against 45 reference populations, drawn from a panel of more than 14,000 people of known ancestry. Raw window calls are noisy, so an autoregressive pair hidden Markov model corrects phasing switch errors, smooths the estimates, and computes confidence scores. Finally, each window is reported at the lowest branch of the ancestry tree that clears a chosen precision threshold: a window the model cannot confidently call Spanish and Portuguese can still be called Southern European, or just European.
That last step is the part most people never see: the painting has a confidence knob. The site lets you view your composition at thresholds from 50 percent (speculative) to 90 percent (conservative), and the numbers move as you turn it. My export happens to ship all five settings, so the figure below has the knob on it, and you can turn it yourself.
Turn it to 90 and the story changes shape. Spanish and Portuguese falls from 44.68 percent to 13.66. Italian falls from 14.03 to 1.69. Nothing about my chromosomes moved, and the genome did not shrink: what leaves those two labels mostly surfaces one level up, as Broadly Southern European and Broadly European, while the share the model declines to place anywhere goes from 0.52 percent to 12.38. That is the trade the knob makes. Fifty percent buys coverage, so nearly every window gets a specific country-shaped name and some of those names are wrong. Ninety percent buys confidence, so what survives is the part the model would defend, and an eighth of me becomes a shrug. The percentages I quoted at the top of this section are the 50 percent reading, the speculative end, which is the right setting for a picture and the wrong one for a conviction.
Below is that export, all of it: both copies of my 22 autosomes plus X, drawn to scale, every segment colored by its top-level assignment, at whichever threshold you pick. The legend percentages are 23andMe's own computed lengths at that threshold rather than sums of the drawn segments, and the report PDF is a separate, earlier run of the same pipeline, which is why its 73.5 and 22.7 sit close to the 50 percent column without matching it. Hover anything: every border you see is the model changing its mind about a reference population at a specific coordinate, on one specific copy.
My favorite detail is the X. I have one X chromosome, inherited from my mother, and its painting is European for most of its length, with a 36-megabase stretch in the middle assigned Indigenous American. One of my maternal lines runs through that segment. The file cannot say which, and I have no records that reach that far; a statistical model, a reference panel, and a stretch of chromosome are the entire archive.
03 · where the file thinks I am from
The painting answers "where are you from" by holding my genome against reference panels. The same export answers it again, by a completely different route.
That second answer is Country Matches. 23andMe keeps reference individuals who reported that all four of their grandparents were born in the same country, then looks for identical pieces of DNA I have in common with those individuals. No model of my ancestry is involved: it is shared segments plus other people's paperwork about their own grandparents. Results come as confidence levels rather than percentages: at least 80 percent confident is "highly likely", 50 to 80 is "likely", 30 to 50 is "possible", under 30 is not reported.
Mine, in the current run: Argentina at 0.8 and Uruguay at 0.8, both highly likely. Italy at 0.7 and Paraguay at 0.7, both likely. Which is my family, roughly as my family tells it.
The previous run said something else. Version 3.0 named five countries, Brazil at 0.5 among them, with twelve Brazilian states under it. Version 3.1 dropped Brazil outright, and with it the largest single count in the dataset: Minas Gerais, 75 grandparents across 48 relatives. A feature update deleted a country from my ancestry, and nothing about me changed. The reading has a version number.
Under the countries sit subregions, where the file gets specific enough to be uncomfortable. Buenos Aires, 23 grandparents across 13 matched relatives. Asunción, 21 across 9. Sicily, 20 across 8. Then Entre Ríos, Piedmont, Córdoba, Apulia, Tacuarembó, a list that reads like an itinerary. Those are not my ancestors. They are counts over other customers' reported grandparents, and the reference individuals are anonymized, so they name nobody.
Then there are Genetic Groups: the same matching, labelled differently. Instead of countries, 23andMe identifies groups of people with significant genetic similarity to each other, then labels those groups from the self-reported information about the individuals in them, such as their birth locations, languages, or cultural affiliations. Mine come back as Guaraní at 0.976, Tupí-Guaraní speakers at 0.948, and Southern Brazilian Highlands at 0.886, all three highly likely.
Be exact about that, because it is easy to say wrongly. It is a statistical association between my DNA and a reference group whose name came from other people's self-reports. It is not a claim that I am Guaraní, and I am not making one. What it is, is the 22.7 percent from section 02 getting more specific: a label the painting could only render as "Indigenous American" resolves, by a different mechanism, toward a particular group in a particular part of the continent, and it lines up with the Paraguay match. That is all it does, and it is enough.
The export dates all of this too. The Ancestry Timeline reads both the number and size of the segments from a population and their distribution across the chromosomes, then estimates how many generations back that ancestry entered the line. The current version predicts Italian at 3 generations, Spanish and Portuguese at 3, French and German at 5, Indigenous American at 6, Angolan and Congolese at 6. The previous version put Spanish and Portuguese at 5, Italian at 5, and Indigenous American at 8. A version bump moved my Iberian estimate by two generations.
23andMe attaches a disclaimer here, and it is aimed directly at me. The algorithm assumes that each ancestry was inherited from a single ancestor, and the company states that for individuals from highly admixed populations, including Latinos and African Americans, that assumption may not be true, and in those cases the estimated number of generations should be used as a place to start rather than a definitive answer. I am from the Río de la Plata. I am the case the caveat names, so those numbers are a starting point and I hold them at arm's length.
Two smaller notes. 23andMe converts generations to dates at an average generation time of 30 years, using the birth date in your profile, which is why I am leaving this in generations and printing no years. And the timeline does not use the X chromosome, because the X is inherited differently from chromosomes 1 through 22. So my favorite detail in the whole export, that 36-megabase stretch on my one X, played no part in the dating. The painting used my X. The timeline did not.
The second answer in full: the countries, the subregion counts behind them, and the three groups.
04 · two unbroken lines
Almost everything in the painting above gets reshuffled every generation. Two pieces of the file do not. Mitochondrial DNA passes from mother to child, in both sexes, and the Y chromosome passes from father to son. No mixing means the variants they carry accumulate in strict lines of descent, and families of shared variants get names: haplogroups. Paternal ones are written as a major branch letter plus a representative marker; maternal ones as nested letters and numbers.
Mine are H6b on the maternal side and R-CTS6889 on the paternal, both read from the same file as everything else: the mitochondrial and Y rows, the single-letter ones from part one. About H6b I will say exactly as much as I could verify: H is the predominant western Eurasian maternal haplogroup, comprising nearly half of the European mitochondrial DNA pool, it is fragmented into many named subclades, and H6b is one of those subclades, with its own distinct phylogeographic pattern. Anything more specific I could not pin to a source I trust, and 23andMe's own page for H6b currently returns an error, so this is where I stop. An unbroken line of mothers, reaching back into the nearly half of Europe's maternal ancestry labeled H, ending at my grandmother, my mother, me.
The paternal line comes with a better story, which is precisely why I hold it more loosely. 23andMe's page for R-CTS6889 places it under R-PF7589, and the haplogroup's long-form name, R1b1a2a1b1, files it deep inside R1b, a paternal family common across western Europe. Among 23andMe research participants it is commonly found in the United Kingdom and Ireland, and the page connects the branch to the lineage of Niall of the Nine Hostages, a King of Tara in northwestern Ireland in the late 4th century C.E., with 2 to 3 million living paternal-line descendants claimed. So the company's telling is that my father's father's fathers trace toward Ireland, medieval king optional. It is a good story. I can verify the branch assignment; the king I leave in the marketing.
The file also counts my Neanderthal inheritance, and the method is one honest sentence: the report screens over 2,000 genetic variants of known Neanderthal origin scattered across the genome and compares your count to other customers. I carry fewer Neanderthal variants than 80 percent of customers. Make of that what you will; I have decided it explains nothing about me and keep mentioning it anyway.
05 · the letters they never measured
Everything so far ran on the 643,161 measured positions. The export contains more than that, and this is the part I find genuinely strange: files full of letters nobody measured.
The first kind is phasing. Part one made the point that a row like AG is an unordered pair; the file knows I carry an A and a G, not which parent contributed which. But the export also ships a phased file, and its accompanying README (vendor boilerplate that ships with every export) describes its contents as genetic information that "has been statistically estimated to have been inherited on the same chromosome together from one parent." Estimated how? Population-based phasing: new haplotypes are old haplotypes reshaped by mutation and recombination, so software can pool information across many genotyped people and estimate which arrangement of my letters matches the haplotype patterns humans actually carry. That is the same EAGLE-adapted phasing the ancestry painting stands on, and 23andMe notes that having a genotyped parent in the system upgrades its quality. The first rows of my phased file:
# rsid chrom position copy1 copy2
rs3131972 1 752721 A G
rs114525117 1 759036 G G
rs12127425 1 794332 G G
Look at the first row. rs3131972 is one of the first five rows in part one's file figure, an unordered AG there. Here it has been split: A on one copy, G on the other. No new measurement happened between those two files. A model assigned my letters to copies because that assignment is the statistically probable one, and the README says so in its own words: phasing "is a statistical process and may contain errors." Whether copy 1 came from my mother or my father, it does not claim to know.
How much does "may contain errors" cover? My export happens to answer that, because it ships two phasing runs over the same genotypes, one computed in 2022 and one in 2025. So I compared them:
node compare-phasing.js phased_2022.txt phased_2025.txt
shared positions 534,444
heterozygous in both 94,747← the only ones phasing can get wrong
switch points 3,012← where the two runs part ways
genotype calls changed 330
That middle number needs care, and getting it wrong is easy. The labels copy 1 and copy 2 are arbitrary, so simply counting positions where the two files differ is meaningless: it came out 47,995 against 46,752, a coin flip, exactly as it should. What matters is where a run changes its mind partway along a chromosome. That is called a switch, and between two switches the runs agree perfectly. There are 3,012 of those points across my genome, roughly one every megabase, with a median of 31 heterozygous positions between them. Note also what this is not: it is a disagreement rate, not an error rate. It says the two runs differ, not which one is right.
Three thousand parting-of-ways is also why the ancestry pipeline carries a hidden Markov model whose stated job is to correct phasing switch errors. The painting in section 02 is built on top of exactly this, which is worth remembering when a border between two ancestries looks precise.
The second kind goes further. The export includes imputed data: positions the chip never assayed at all, filled in statistically. I counted mine:
for f in imputed_genotype_data_r6/*.bcf; do bcftools view -H "$f"; done | wc -l
84031008
84,031,008 rows, 130 times the number of positions actually measured. Genotype imputation uses haplotype reference panels to estimate untyped markers from typed ones, and it is completely standard in genomics, used routinely to boost power and combine studies. 23andMe's own analysis pipeline works with imputed dosages internally. The imputed files' README is admirably direct about what they are: data that "contains genetic information that is not directly assayed by our technology but has been statistically estimated based on the raw genetic data," which "has not undergone any quality review and, as such, is suitable only for research, educational, and informational use and not for medical or other use."
Sit with that. The company shipped me 84 million genotypes it never measured, correctly labeled as statistical estimates, because in this field that is a normal and useful thing to do. It is also the cleanest possible summary of this whole series. The chip measured brightness at 643,161 positions; everything else was inference. Called from intensities, phased by population statistics, classified by a support vector machine, smoothed by a hidden Markov model, thresholded by a knob, and extrapolated to positions never touched. The file does not contain my genome. It contains an argument about my genome, and the argument is very good.
06 · what I left out
Everything above was chosen. Here is what I deliberately did not publish, in one paragraph: the positions in my file with medical literature attached, because the header's own admission that "only a subset of markers have been individually validated for accuracy" deserves to be taken seriously in exactly the cases where being wrong costs something; the family tree and DNA relatives features, which name other living people; and anything that would resolve my relatives' genomes, because my file is not only mine. Half of it is my mother's and half my father's, and publishing my rows publishes half of each of theirs. Earwax cost me nothing to share. Their genomes were not mine to spend.
That choice is the note I want to end the series on. The download arrives as one archive, the famous trait rows sitting in the same folders as the positions you may not want in a search index, and nothing in the interface distinguishes the two. If you pull your own file, and I think you should, decide what you are publishing before you publish it, on purpose, the way you would with any other dataset that happens to describe your family. The file says a lot about me. It says it in probabilities, painted in segments, at a confidence threshold I got to pick. I published the parts of myself where a probability is a fine thing to be.