PolyGenius
Contents

compute

PolyGenius data transformation engine

Namespaced entry point for every transformation that reads or writes a PolyGeniusStudy's sample, model or score axes. Its components:

  • compute$scores(): Calculate PRS scores from genotypes
  • compute$scores.plan(): A-priori resource estimate (peak RAM, batch count, strategy) for compute$scores() -- never a score matrix; optionally sweeps a memory budget
  • compute$standardize(): Build a per-model centered/scaled score layer, tagged with its (center, scale, scale.source)
  • compute$populationStructure(): Compute principal components for population stratification correction
  • compute$relatedness$kinship(): Sparse KING kinship over related sample pairs
  • compute$relatedness$prune(): Samples to retain after removing relatives
  • compute$similarity$samples(): Compute sample-sample similarity matrices
  • compute$similarity$models(): Compute model-model similarity matrices
  • compute$embedding$samples(): Compute low-dimensional embeddings for samples
  • compute$embedding$models(): Compute low-dimensional embeddings for models
  • compute$genome$concordance(): Per-variant cross-trait effect-direction concordance (anchored)
  • compute$genome$convergence(): Per-variant convergent outcome-signal (Sigma|attribution|)
  • compute$genome$cumulativeWeight(): Per-variant cumulative |beta| across the model library
  • compute$genome$attribution(): Per-(variant,model) attribution (weight x single-variant effect)

Usage

compute

Details

A typical compute workflow:

  1. Calculate PRS scores:
compute$scores(data, maf.thr = 0.01)

  1. Compute population structure:
compute$populationStructure(data,
                            npcs = 10,
                            reference.panel = "EUR")

  1. Resolve relatedness (before PCA -- components are distorted by families):
data$sample.pairs$kinship  <- compute$relatedness$kinship(data, degree = 2)
data$samples$unrelated <- compute$relatedness$prune(data, degree = 2)
data <- data[unrelated, ]

  1. Compute similarity and embeddings:
# Sample similarity
sim.samples <- compute$similarity$samples(data, X, method = "pearson")
# Model similarity
sim.models <- compute$similarity$models(data, X, method = "snp.jaccard")
# PCA embedding
pca.emb <- compute$embedding$samples(data, X, method = "pca", n.components = 10)
# UMAP embedding
umap.emb <- compute$embedding$samples(data, X, method = "umap")

Single-variant association (GWAS) lives in associate$singleVariant(), not here.