PolyGenius
Contents

compute$relatedness

Sample relatedness: KING kinship and pruning to unrelated samples

compute$relatedness finds related samples in a PolyGeniusStudy and chooses which to keep so that no two retained samples are related. It has two steps, run in order:

  • compute$relatedness$kinship() estimates KING-robust kinship with PLINK 2 and returns the related pairs as a sparse sample-by-sample matrix, for data$sample.pairs.
  • compute$relatedness$prune() reads that matrix and returns a logical vector with one element per sample: TRUE for the samples to keep.

Neither step changes data. Storing the matrix and dropping samples are left to the caller, so every change to the cohort is visible in the analysis script.

How kinship is computed

PolyGenius does not estimate kinship itself. For each genotype fileset it runs PLINK 2 --make-king-table, which applies the between-family KING-robust estimator to every sample pair over autosomal markers. A fileset split over several files is merged first, because KING needs all of a pair's markers at once. Only pairs at or above the cut point are written out, so the result grows with the number of related pairs rather than with the square of the sample count.

Kinship is scaled so that a duplicate sample scores 0.5. degree selects a conventional cut point, each the geometric mean of two adjacent degrees:

degree Cut point Pairs reported
0 0.354 duplicates and monozygotic twins
1 0.177 the above, plus parent-offspring and full siblings
2 0.0884 the above, plus half siblings, grandparent-grandchild and avuncular pairs
3 0.0442 the above, plus first cousins

Kinship is evaluated within one fileset only. A participant genotyped in two filesets is not detected as a duplicate.

A pair with no marker called in both samples has no estimate. PLINK reports it as NaN. kinship() leaves it out of the matrix, keeps it in the pairs artifact, and warns.

How pruning chooses samples

prune() treats every pair with kinship strictly above the cut point as related, and removes samples until no related pair remains. It uses the greedy algorithm of PLINK 2 --king-cutoff, implemented in R so that a stored matrix can be pruned without running PLINK again. At each step:

  1. If a sample has exactly one remaining relative, that relative is removed.
  2. Otherwise, the sample with the most remaining relatives is removed.

Ties go to the sample that comes first in the study. For the same pairs and sample order, prune() keeps exactly the samples PLINK keeps. Like PLINK's, the result is not guaranteed to be the largest possible set, and a dropped sample may have no retained relative left.

The choice ignores phenotypes on purpose. Preferring to keep cases would enrich for familial cases, which for a heritable outcome inflates a score's apparent effect.

Before you rely on the result

  • Ancestry. KING-robust uses no allele frequencies. It is unbiased for pairs from the same population, including in a multi-ancestry cohort. It is biased for admixed relatives, in either direction: a parent and an admixed child can come out low, and admixed siblings high. Unrelated pairs from different populations come out strongly negative.
  • Resolution. Relationships of the same degree estimate the same kinship. Half siblings, grandparent-grandchild and avuncular pairs all sit near 0.125. Parent-offspring and full siblings both sit near 0.25. Only ibs0 separates them: it is near zero for a parent and child.
  • Order. Resolve relatedness before computing principal components. Prune, run compute$populationStructure() on the retained samples, then project the rest.

Examples

data$sample.pairs$kinship <- compute$relatedness$kinship(data, degree = 2)
data$samples$unrelated    <- compute$relatedness$prune(data, degree = 2)
data <- data[unrelated, ]

See Also

Aliases: compute-relatedness, compute$relatedness