Contents
compute$relatedness
Sample relatedness: KING kinship and pruning to unrelated samples
compute$relatedness finds related samples in a PolyGeniusStudy and chooses which to
keep so that no two retained samples are related. It has two steps, run in order:
- compute$relatedness$kinship() estimates KING-robust kinship with PLINK 2 and returns the related pairs as a sparse sample-by-sample matrix, for
data$sample.pairs. - compute$relatedness$prune() reads that matrix and returns a logical vector with one element per sample:
TRUEfor the samples to keep.
Neither step changes data. Storing the matrix and dropping samples are left to the
caller, so every change to the cohort is visible in the analysis script.
How kinship is computed
PolyGenius does not estimate kinship itself. For each genotype fileset it runs PLINK 2
--make-king-table, which applies the between-family KING-robust estimator to every
sample pair over autosomal markers. A fileset split over several files is merged first,
because KING needs all of a pair's markers at once. Only pairs at or above the cut point
are written out, so the result grows with the number of related pairs rather than with
the square of the sample count.
Kinship is scaled so that a duplicate sample scores 0.5. degree selects a
conventional cut point, each the geometric mean of two adjacent degrees:
degree |
Cut point | Pairs reported |
|---|---|---|
0 |
0.354 | duplicates and monozygotic twins |
1 |
0.177 | the above, plus parent-offspring and full siblings |
2 |
0.0884 | the above, plus half siblings, grandparent-grandchild and avuncular pairs |
3 |
0.0442 | the above, plus first cousins |
Kinship is evaluated within one fileset only. A participant genotyped in two filesets is not detected as a duplicate.
A pair with no marker called in both samples has no estimate. PLINK reports it as NaN.
kinship() leaves it out of the matrix, keeps it in the pairs artifact, and warns.
How pruning chooses samples
prune() treats every pair with kinship strictly above the cut point as related, and
removes samples until no related pair remains. It uses the greedy algorithm of PLINK 2
--king-cutoff, implemented in R so that a stored matrix can be pruned without running
PLINK again. At each step:
- If a sample has exactly one remaining relative, that relative is removed.
- Otherwise, the sample with the most remaining relatives is removed.
Ties go to the sample that comes first in the study. For the same pairs and sample
order, prune() keeps exactly the samples PLINK keeps. Like PLINK's, the result is not
guaranteed to be the largest possible set, and a dropped sample may have no retained
relative left.
The choice ignores phenotypes on purpose. Preferring to keep cases would enrich for familial cases, which for a heritable outcome inflates a score's apparent effect.
Before you rely on the result
- Ancestry. KING-robust uses no allele frequencies. It is unbiased for pairs from the same population, including in a multi-ancestry cohort. It is biased for admixed relatives, in either direction: a parent and an admixed child can come out low, and admixed siblings high. Unrelated pairs from different populations come out strongly negative.
- Resolution. Relationships of the same degree estimate the same kinship. Half siblings, grandparent-grandchild and avuncular pairs all sit near 0.125. Parent-offspring and full siblings both sit near 0.25. Only
ibs0separates them: it is near zero for a parent and child. - Order. Resolve relatedness before computing principal components. Prune, run
compute$populationStructure()on the retained samples, then project the rest.
Examples
data$sample.pairs$kinship <- compute$relatedness$kinship(data, degree = 2)
data$samples$unrelated <- compute$relatedness$prune(data, degree = 2)
data <- data[unrelated, ]See Also
compute$relatedness$kinship, compute$relatedness$prune
Other compute-relatedness:
compute-relatedness-kinship,
compute-relatedness-prune