Computing Scores and Covariates
Scores, relatedness, population structure and the other quantities compute derives
compute$* derives quantities from your genotypes. PLINK2 does the genotype work; PolyGenius decides what to ask it for and puts the answer somewhere useful.
One rule before anything else: you assign every result. Nothing in compute$* writes into your cohort object for you.
Scores
compute$scores() returns a score per sample per model, which you assign into a layer.
Rows follow the object's sample order and carry no names, so join scores to anything else through the object, never by position after a sort.
The first question after scoring is coverage: how many of the model's variants your cohort actually had. It is worth reporting in a methods section, and it is the fastest way to notice a build or chromosome-naming mismatch.
What the number is: a sum over the variants that were scored. Its scale therefore depends on how many variants that was, which is why you standardise before comparing or reporting.
Coverage differs from what you asked for because variants are matched on position and alleles in either orientation. Do not pre-flip your betas or swap alleles to match your genotypes — that produces a double flip and a silent sign error.
A missing genotype contributes zero rather than the cohort mean, so a person's score never depends on who else is in the file. The cost is that samples with a lot of missingness get systematically lower sums, so filter on missingness before scoring rather than after.
A genome split across files gives one genome-wide score — provided those files are registered as one genotype rather than as several.
You can restrict scoring to common variants with a minor-allele-frequency threshold. It uses your own cohort's frequencies, so if you plan to compare cohorts later, keep that in mind (chapter 07.5).
For a large model library, ask for a size estimate before submitting the job rather than after it fails.
data$scores$X <- compute$scores(data)
data$scores$X.scaled <- scale(data$scores$X)How that size estimate is produced — batching, memory knobs, the fate-table schema, and the execution regimes PolyGenius chooses between — is covered in the scoring engine.
Relatedness
compute$relatedness$kinship() estimates pairwise relatedness and stores it as a sample-by-sample matrix, with the supporting pair-level detail alongside it.
compute$relatedness$prune() returns a decision, not a subset — a logical vector marking who to keep. Applying it stays your explicit step, because a function that silently shrank your cohort would change every reported sample size.
Read the counts back from the diagnostics. Both steps use the same cut point, but kinship reports pairs at or above it while pruning treats only pairs strictly above it as related, as PLINK does. A pair sitting exactly on the cut point is counted by one and not the other.
Reuse is guarded: a stored kinship matrix computed at a coarser degree is refused rather than reused, because the pairs in between were never computed.
What KING can and cannot tell you: it is unbiased for pairs from the same population, even in a multi-ancestry cohort, but biased for admixed relatives, and several relationship types are indistinguishable at the same kinship value.
Pruning uses the greedy algorithm of PLINK 2's --king-cutoff, so it keeps the same samples PLINK would. It is also phenotype-blind. That is deliberate — preferentially retaining cases would enrich for familial cases and inflate the apparent effect of a score.
data$sample.pairs$kinship <- compute$relatedness$kinship(data, degree = 2)
data$samples$unrelated <- compute$relatedness$prune(data, degree = 2)
data <- data[unrelated, ]Population structure
compute$populationStructure() either computes components in your own sample or projects your samples onto a reference panel. One argument switches between them.
Columns are PC1, PC2, and so on. Do not confuse them with compute$embedding$*, which reduces a score matrix and produces PCA1, PCA2 — those are not population-structure covariates.
A projected coordinate is a per-allele average, and when a genome is split across files the parts are combined once rather than averaged per file.
A correlation check cannot validate this. Projecting a dataset onto its own components is near-perfectly correlated whether the scale is right or badly wrong; only an absolute comparison distinguishes them.
Cross-cohort work has to project every cohort onto the same panel with the same variant set, or the components mean different things in each cohort.
Two limits to plan around: in-sample components need roughly 50 samples or more, and the computation is randomised, so two identical calls differ slightly.
A sample with no scoreable variant lands at zero rather than NA — a cluster at the origin that can look like average ancestry.
data$samples$PCA <- compute$populationStructure(data, npcs = 5)Similarity, embeddings and genome signals
compute$similarity$* compares samples to each other or models to each other, and the result belongs in the matching sample-by-sample or model-by-model slot. Restricting to a subset returns a full-sized matrix with NA outside the selection.
compute$embedding$* reduces a score or similarity matrix to a few dimensions, with columns named after the method.
compute$genome$* is different in kind: it reads model tables and association summaries rather than genotypes, returns a standalone positioned object, and refuses per-sample columns so the result stays shareable. Drawing it is chapter 12.2.
One standing habit
Read the PLINK log before theorising. Allele flips, dropped variants and "no variants remaining" appear there and nowhere else.