PolyGenius
Contents

compute$populationStructure

Population structure from genotypes, in-sample or projected onto a panel

Computes sample principal components from a study's genotypes with PLINK2. Without a reference.panel the components are the study's own in-sample axes; with one, the samples are projected onto that panel's axes instead.

Usage

compute$populationStructure(
  data,
  npcs = 5,
  variants = NULL,
  reference.panel = NULL
)

Arguments

ArgumentDescription
dataA PolyGeniusStudy.
npcsInteger scalar, default 5. Number of principal components to compute, or to project onto.
variantsA file path, a data.frame, or NULL (default). The marker set the PCA is computed over -- a BED1-style text file with one variant per row (chromosome, start, end), or an in-memory table with those columns. NULL resolves the common20k variant space through workspace$catalogs$variantSpaces, and aborts if that resource cannot be reached.
reference.panelCharacter scalar naming a super-population reference panel, or NULL (default). NULL computes in-sample PCA; a name projects the samples onto that panel and aborts when no panel matches both the name and the genotypes' build. See [ReferencePanels](/reference/referencepanels/) for the supported names.

Value

For a PolyGeniusStudy, a numeric matrix with one row per sample in data$sample.names order (rows carry no names) and npcs columns, PC1..PCnpcs. It carries a PolyGeniusProvenance() readable with provenance(), whose $misc holds eigenvalues -- a named list, one numeric vector per genotype, since an in-sample eigendecomposition is per fileset -- for in-sample PCA. variants, the marker specification as supplied or defaulted, is recorded on either branch; a projection adds panel.id, build and variant.space.digest. The matrix is reduction()-tagged, so data$samples$PCA <- ... files it as an embedding rather than a table member; storing it is the caller's job.

Details

In-sample PCA has no scale correction across a fileset or study boundary. --pca runs once per genotype fileset and each run's axes are its own, so PC1 of one fileset is not the same axis as PC1 of another and stacking their embeddings would combine incomparable coordinates. A study spanning more than one fileset therefore aborts when unconditionally. The error names both remedies: merge the filesets into one genotype, or supply reference.panel so every sample is projected onto the same axes. Projection in turn requires every genotype in the study to be on one build, and aborts otherwise.

Scale. In-sample PCA returns PLINK's eigenvector coordinates; projection returns PLINK's per-allele average projection (its default SCORE_AVG). Both are valid stratification covariates, but they are not on the same numeric scale: projecting a dataset onto its own in-sample PCA reproduces each component only up to a per-component positive multiple and an arbitrary sign. Downstream regression absorbs both, so the stratification adjustment is unaffected.

Examples

# In-sample components, stored as a sample embedding for use as covariates
data$samples$PCA <- compute$populationStructure(data, npcs = 10)

data$samples$PCA <- compute$populationStructure(data, npcs = 10, reference.panel = "EUR")

See Also

Aliases: compute-population-structure, compute.population.structure, compute$populationStructure