Heritability

What heritability measures

Heritability is a statistic about groups. Take a large number of people who differ in height, mood, drinking, test scores. The question is what share of that variation tracks the fact that they carry different genes. The answer describes the spread of the population. Individual people do not have heritabilities.

The three sources

The standard model splits that variation three ways. Genes: different DNA. Shared environment: whatever makes siblings in the same family alike — parents, income, neighbourhood, school, era. Non-shared environment: whatever is unique to each person, including randomness in development and error in the measurement itself. Heritability is the share from the first, and the three shares add to one by construction. This is a simplifying model rather than an inventory of causes: it assumes genetic effects add up, and that genes and environments are dealt out independently of each other. Where those assumptions fail, the three components absorb things that do not belong to them.

The findings in outline

Genes account for a substantial share of variation in adult psychological traits: a third to a half for personality and mood, more for intelligence and severe psychiatric conditions. Non-shared environment accounts for a comparably large share. Shared environment accounts for very little. Adopted siblings raised in the same house end up barely more alike in adult personality or intelligence than two strangers, a result that has survived forty years of attempts to overturn it.

Reading the numbers

The baseline is the spread of the trait across the whole population, which is set to 1, and every figure in this field is a proportion of it. Two people drawn at random correlate at 0, so zero means as unalike as strangers. This is also why the estimates move when the population changes: the denominator is the population’s own variability, not any absolute scale.

These are shares of variance, not shares of visible difference, so halving a correlation does not halve how alike two people are. Adult men vary with a standard deviation of about 7cm, so two drawn at random differ by about 8cm on average; identical twins correlate near 0.9 and still differ by roughly 2.5cm; a pair at 0.5 differ by about 5.5cm, seven tenths of the way to strangers rather than half. The top of the scale is compressed like this throughout, and nothing reaches 1 in any case, since no pair can correlate above the reliability of the instrument used on them — about 0.8 for a depression questionnaire.

How the designs work

Separating the three sources requires families in which the amount of genetic overlap is known and fixed.

Every design yields a correlation, computed across many pairs, and its complement. The correlation measures how much of the population’s spread the two members of a pair share. One minus it is the share of variance their resemblance leaves unexplained — what separates them. Both are made of the same three components, and each design is a read on one side or the other.

Start with identical twins raised together. On the resemblance side, they share all their genes and their home, so anything making them alike could be either, and their correlation cannot be attributed to one source. On the difference side, nothing separates them but individual experience. So one minus their correlation is non-shared environment, with nothing else mixed in.

Adopted siblings are the mirror image. On the resemblance side, the only thing that can be making them alike is the home, so their correlation is a direct read on shared environment. On the difference side they are confounded. Two people with no genetic overlap are as genetically unalike as strangers, so one minus their correlation contains genes and individual experience together, with no way to separate the two.

Fixing genetic overlap decides how much genes contribute to resemblance, and whatever the overlap does not explain shows up on the difference side instead. Each design is therefore clean on one side and confounded on the other.

The main designs:

  • Identical twins raised together — the difference side gives non-shared environment.
  • Adopted siblings raised together — the resemblance side gives shared environment.
  • Identical against fraternal twins — home and individual experience act equally in both groups, so both drop out of the comparison. Genetic overlap does not: complete in one group, half in the other. The gap between the two correlations gives genes.
  • Identical twins raised apart — genes shared, home not, so their resemblance gives heritability in a single step with no subtraction. The most direct design and the rarest, though not an assumption-free one: it holds only where placement was unrelated to the family the twins came from, and where the pairs were separated early and stayed out of contact. Many historical cases fail one of these.
  • Siblings compared by measured DNA — the genetic difference between them comes from the random draw at conception, which takes ancestry and family-level circumstance out of the estimate. Siblings still differ in age, sex and experience, so this controls for family background rather than holding everything but DNA fixed, and it costs much larger samples.

No single kind of relative pair identifies all three on its own. The classical design gets all three by combining identical and fraternal pairs, at the cost of assuming the model holds. Running several designs with different weaknesses is how those assumptions are checked against one another.

The arithmetic

Take ten pairs of identical twins and ten pairs of fraternal twins raised by their own parents, and score everyone on a depression questionnaire. Use a continuous score rather than a diagnosis, since concordance between diagnosed pairs must first be converted to a liability scale. Suppose identical pairs correlate at 0.5 and fraternal pairs at 0.3.

Shared environment is constant across the groups, so the extra 0.2 must be genetic. It covers only the step from half-shared to fully-shared genes, half the distance, so genes come to 0.4. On the resemblance side, identical twins correlated at 0.5 and genes account for 0.4, leaving 0.1 for shared environment. On the difference side, one minus their correlation is also 0.5, and since they share genes and home, all of it is non-shared environment.

Both correlations move the result in obvious ways. Anything widening the gap between them raises the estimate for genes, and anything raising the fraternal correlation alone raises shared environment.

The 0.1 was inferred by subtraction rather than measured. Adopted siblings read shared environment directly, off their resemblance side. If the twin arithmetic holds, ten such pairs should correlate near 0.1. Suppose they come in at 0.2.

Where the designs go wrong

The twin method rests on three assumptions: about the environment, about how genetic effects combine, and about how much of their DNA fraternal twins actually share.

The first is that identical and fraternal twins are treated equally similarly, which is not quite true: identical twins are dressed alike, confused for each other, and share more friends. One of the better tests uses twins who are wrong about their own zygosity, and they resemble each other according to their actual type rather than the believed one. The assumption is imperfect but does not appear badly violated. Where it fails, it raises the identical correlation and inflates genes.

The same logic requires that non-shared environment acts equally in both groups, and there is a known reason it might not. About two-thirds of identical pairs share a placenta, which means both competition for a single blood supply and a more nearly common prenatal environment. The first adds difference to identical pairs specifically and would lower their correlation; the second adds resemblance and would raise it. Comparisons of shared-placenta against separate-placenta identical pairs come out in both directions depending on the trait, and the effects are modest. This is a known violation of unsettled sign rather than a correction that can be applied.

The second assumption, that genetic effects simply add up, is more of a problem. Where particular combinations of variants matter, fraternal twins inherit less than half the resemblance identical twins get, so the gap between the correlations widens for reasons that are not additive and doubling it overstates genes. Shared environment then comes out negative whenever the identical correlation is more than twice the fraternal one, which is where reported estimates of zero shared environment often come from.

The third is that fraternal twins share half their genes. That is the expected figure under random mating, and people do not mate randomly. Spouses correlate at roughly 0.5 for political and religious attitudes, 0.4 for education, 0.3 for intelligence, 0.2 for height, and near 0.1 for personality. When parents resemble each other on the variants behind a trait, their children’s genetic resemblance on that trait exceeds the half the model assumes, while identical twins stay at one whatever the parents do. Siblings still share half their DNA on average; it is the trait-relevant share that rises, and sibling genetic correlations above 0.5 have been measured directly for educational attainment. The gap between the correlations narrows for a reason that has nothing to do with the size of the genetic effect, so genes are understated and shared environment overstated. Part of the apparent shared environment for education and political attitudes is likely this.

These biases do not all point one way. Differential treatment and non-additive effects inflate the genetic estimate, assortative mating deflates it, and prenatal differences can push either way. That is part of why the estimates hold up better than the length of the list suggests, and why no single correction fixes them.

The adoption method is biased in two directions at once. Screening makes adoptive homes more uniform than homes in general, which understates how much homes matter and mechanically inflates the heritability estimated in such samples, for the same reason that any restriction of environmental range does. Non-random placement reintroduces genetic similarity, which overstates the home’s contribution. And ten pairs could not reliably distinguish 0.1 from 0.2 in any case.

The discrepancy therefore does not show which estimate is closer to the truth, and with ten pairs it would mostly be sampling noise in any case. What matters is whether larger studies using designs with different weaknesses converge. For depression they roughly do: small shared environment, moderate genes, large non-shared environment.

Why the home explains so little

A near-zero estimate for shared environment means homes rarely make one person differ from another. That is a claim about differences between people, and a much narrower one than a claim about whether homes matter. Nearly every home supplies the basics, and something everyone gets cannot explain why people differ. Where they are absent — neglect, malnutrition, lead, no schooling — the damage is severe. Shared environment behaves like a floor rather than a dial, and its effects fade as people leave home.

The two causes are also confounded: the environments people meet are not dealt out independently of their genes. Gene-environment correlation works three ways.

Passive: parents hand down a household along with their DNA, so the bookshelves and the variants arrive together. Evocative: a child’s disposition draws different responses from parents, teachers and peers, so the environment is partly a reply to the person in it. Active: people seek out and are sorted into settings that suit their existing inclinations.

None of the three is modelled explicitly, so each is absorbed somewhere it does not belong. Active and evocative correlation often add resemblance that is then attributed to genes, since the environments involved track a person’s own dispositions; passive correlation can raise the apparent effect of the home instead. Where the covariance actually lands depends on the design and on whether the environment in question is shared, so these are tendencies rather than rules. Either way the components include environmental effects that genes set in motion, which is a different claim from environments having no effect.

The confounding also runs the other way, and that direction can be measured. Parents’ genes shape the household they build, so some of what looks genetic is the home: the alleles a parent fails to pass on are physically absent from the child, and they still predict that child’s educational outcomes. That prediction can only be running through the home the parent made.

Shared environment does persist for religious denomination, language, political affiliation, and whether a person ever starts smoking or drinking. Religion divides in two. Which faith you belong to is almost entirely familial and barely heritable. How religious you are is heritable at around 0.3 to 0.4 in adults, but not in children, where the same measures are dominated by the home; the genetic share only emerges as people age.

Why individual circumstance explains so much

The non-shared term is large partly because it is a residual, collecting everything the other two do not account for. Some of what collects there is an artefact of how the estimate is made.

Measurement error lands there, and not in trivial amounts. Studies of personality that use several raters and several occasions, rather than one self-report questionnaire, push heritability from roughly 0.45 towards 0.6 or higher. That increase comes straight out of the non-shared estimate.

Unmodelled gene-environment interaction lands there too. The standard model assumes a genetic effect is the same size in every environment. Where it is not, the mismatch is absorbed into this term rather than reported. Correlation is about which environments people meet; interaction is about what a given environment does to different people.

Much of what is left is genuinely individual, though the residual goes on absorbing whatever else the model has not specified. Some is developmental randomness: identical twins have different fingerprints, shaped by random variation in the womb rather than by genes or upbringing. Some is prenatal and lasting, since the smaller identical twin at birth tends to score lower decades later. Life events matter too, though decades of searching for specific replicable ones — birth order, peer groups, differential parenting — have produced small and inconsistent effects.

Some of the remainder is probably chance. It is also consistent with a very large number of causes too small for any study to detect. The two possibilities are difficult to tell apart.

How heritability changes with age

For intelligence, heritability rises from roughly 0.2 in early childhood to around 0.7 in adulthood. The shared-environment component fades as people leave home. At the same time, early genetic differences are amplified rather than replaced, so small initial differences compound into larger ones and much of the rise reflects amplification rather than genes switching on. Active and evocative gene-environment correlation — people selecting, and being sorted into, environments that suit their existing inclinations — are the usual proposed mechanisms, though the amplification itself is much better established than any particular account of what drives it.

A childhood test score is therefore partly a measurement of the household, which is why it predicts adult ability only moderately. Early gains from intervention often fade on cognitive tests, plausibly because they were made in the component that was going to wash out. Several programs nonetheless show durable effects on schooling and earnings long after the test gains go.

The numbers

Adults unless noted: height 0.85; schizophrenia and bipolar disorder near 0.8; autism around 0.8; ADHD around 0.7; adult intelligence 0.6–0.8; body weight 0.6–0.8; addiction 0.5–0.6; personality 0.4–0.5 by questionnaire and higher when better measured; depression and anxiety 0.3–0.4; childhood intelligence 0.2–0.4; denomination and language near zero; political attitudes 0.3–0.5, while party identification is far more context-dependent, running from almost entirely familial in some samples to substantially heritable in others.

These are twin estimates. Methods that add up measured DNA typically recover a third to a half of them, and closer to 0.2 or 0.3 for intelligence and ADHD. Rare variants and imperfect coverage of the genome account for part of the gap. Part of the rest is that the two methods are not measuring quite the same quantity: twin estimates absorb non-additive effects and gene-environment correlation, both of which the DNA-based methods largely miss, so the twin figure is the more inclusive number as well as the larger one. The remainder is unresolved.

Each figure applies to one population living in one range of environments, and moves when that range moves. In a uniformly well-fed population, variation in height is mostly genetic; introduce famine unevenly and heritability drops, because a new source of variation has been added. Where environments vary a lot, they explain more.

What the genetic influence is made of

Genetic influence on intelligence or depression is spread across thousands of common variants, each shifting the trait by an amount too small to detect on its own, with rare variants of larger effect in a minority of cases. There is no single gene for intelligence or depression; ordinary variation in these traits is polygenic.

Adding the common variants up produces a polygenic score. The best of these predict something like 12 to 15 percent of variation in educational attainment, and around 40 percent for height. That is real for ranking a population and weak for any individual, since people with the same score differ enormously. The association also falls by roughly 40 to 50 percent when scores are compared between siblings rather than between strangers, and the variance explained falls further still, since it goes roughly with the square of the association. Part of what a score captures in the general population is ancestry, family circumstance and assortative mating rather than the causal effect of the variants themselves. Scores also transfer poorly to ancestries other than the one they were trained on.

Heritability is not family risk

The most common practical error is reading heritability as a rate of family transmission. The two can point in opposite directions: religion is transmitted through families about as strongly as anything is and is barely heritable, while depression is moderately heritable and most children of depressed parents never develop it.

You inherit half your DNA from each parent, drawn at random, so you get an unpredictable sample of a parent’s variants rather than a copy. Full siblings vary widely around the half they share on average, and a highly heritable trait can fail to run in a family. The same shuffling produces regression to the mean: two very tall parents have tall children, but less extreme, because their own extremity came partly from a combination that will not be reassembled. Assortative mating pushes the other way, since people pair with partners like themselves, which makes traits run in families more strongly than random mating would predict.

Risk to a relative depends on the base rate as much as on the strength of the genetic effect. Schizophrenia is among the most heritable conditions known, and about one percent of people develop it. A child with one affected parent has roughly a ten percent chance. That is a tenfold relative increase and a ninety percent chance of nothing. Most people who develop schizophrenia have no affected parent, because the population contains far more unaffected parents.

For depression, at a lifetime prevalence near fifteen to twenty percent, a depressed parent multiplies risk by two or three. Both figures come from base rates and observed recurrence in families. Heritability is not an input to that calculation.

Three misreadings

Heritable does not mean unchangeable. Phenylketonuria is genetic in the full sense and is prevented by diet.

Heritability does not divide up a person. Everyone is entirely genetic and entirely environmental. The number exists only for a population.

A heritability estimated within a group applies only within that group. Identical seed sown in rich and poor soil produces plants whose height varies almost entirely by genotype within each plot, while the whole gap between the plots comes from the soil.

Summary

Genes and individual circumstance each account for a large share of why adults differ. The family home accounts for little, outside deprivation and a few traits like religion and language. Family risk is a separate question, answered from base rates. And a large part of what makes two people with the same DNA and the same upbringing turn out differently remains unidentified, and may be beyond identifying.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *