Contents
compute$populationStructure
Population structure from genotypes, in-sample or projected onto a panel
Computes sample principal components from a study's genotypes with PLINK2. Without a
reference.panel the components are the study's own in-sample axes; with one, the
samples are projected onto that panel's axes instead.
Usage
compute$populationStructure(
data,
npcs = 5,
variants = NULL,
reference.panel = NULL
)Arguments
| Argument | Description |
|---|---|
data | A PolyGeniusStudy. |
npcs | Integer scalar, default 5. Number of principal components to compute, or to project onto. |
variants | A file path, a data.frame, or NULL (default). The marker set the PCA is computed over -- a BED1-style text file with one variant per row (chromosome, start, end), or an in-memory table with those columns. NULL resolves the common20k variant space through workspace$catalogs$variantSpaces, and aborts if that resource cannot be reached. |
reference.panel | Character scalar naming a super-population reference panel, or NULL (default). NULL computes in-sample PCA; a name projects the samples onto that panel and aborts when no panel matches both the name and the genotypes' build. See [ReferencePanels](/reference/referencepanels/) for the supported names. |
Value
For a PolyGeniusStudy, a numeric matrix with one row per sample in data$sample.names
order (rows carry no names) and npcs columns, PC1..PCnpcs. It carries a
PolyGeniusProvenance() readable with provenance(), whose $misc holds eigenvalues
-- a named list, one numeric vector per genotype, since an in-sample eigendecomposition
is per fileset -- for in-sample PCA. variants, the marker specification as supplied or
defaulted, is recorded on either branch; a projection adds panel.id, build and
variant.space.digest.
The matrix is reduction()-tagged, so data$samples$PCA <- ... files it as an embedding
rather than a table member; storing it is the caller's job.
Details
In-sample PCA has no scale correction across a fileset or study boundary. --pca
runs once per genotype fileset and each run's axes are its own, so PC1 of one fileset is
not the same axis as PC1 of another and stacking their embeddings would combine
incomparable coordinates. A study spanning more than one fileset therefore aborts when
unconditionally. The error names both remedies: merge the filesets into one genotype, or
supply reference.panel so every sample is projected onto the same axes. Projection in
turn requires every genotype in the study to be on one build, and aborts otherwise.
Scale. In-sample PCA returns PLINK's eigenvector coordinates; projection returns
PLINK's per-allele average projection (its default SCORE_AVG). Both are valid
stratification covariates, but they are not on the same numeric scale: projecting a
dataset onto its own in-sample PCA reproduces each component only up to a per-component
positive multiple and an arbitrary sign. Downstream regression absorbs both, so the
stratification adjustment is unaffected.
Examples
# In-sample components, stored as a sample embedding for use as covariates
data$samples$PCA <- compute$populationStructure(data, npcs = 10)
data$samples$PCA <- compute$populationStructure(data, npcs = 10, reference.panel = "EUR")