PolyGenius
Contents

compute$embedding

Low-dimensional embeddings of a study's samples or models

compute$embedding$samples() embeds samples and compute$embedding$models() embeds models, from any matrix-shaped study slot or from a pre-computed similarity matrix (PCA, MDS, UMAP, t-SNE, PHATE). Both return a full-sized coordinate matrix and write nothing into data.

Usage

compute$embedding$samples(
  data,
  layer,
  method = "pca",
  similarity = NULL,
  n.components = 2,
  samples = NULL,
  ...
)

compute$embedding$models(
  data,
  layer,
  method = "pca",
  similarity = NULL,
  n.components = 2,
  models = NULL,
  ...
)

Arguments

ArgumentDescription
dataA PolyGeniusStudy. Read-only.
layerUnquoted slot expression, resolved by data$fetch() -- e.g. scores, samples$PCA, sample.pairs$kinship. Required unless similarity is supplied, and ignored when it is.
methodOne of "pca" (default), "mds", "umap", "tsne", "phate".
similarityA square numeric matrix over every sample (model), or NULL (default). Replaces layer when supplied. For method = "mds" a matrix whose mean diagonal exceeds 0.5 is read as a similarity and converted to 1 - similarity. UMAP and t-SNE then take it as a precomputed distance; PCA and PHATE embed it as an ordinary matrix.
n.componentsPositive integer scalar, default 2. Number of dimensions requested. For "pca" it must not exceed min(n.rows - 1, n.cols) of the matrix being embedded: PCA truncates to that many components, and the call then fails while naming the columns.
samplesCharacter vector of sample names, numeric or logical index, tidyselect call, a predicate over data's own columns (e.g. age > 65), or NULL (default) for every sample. Excluded samples are NA rows in the returned matrix.
...Method-specific arguments. PCA: center (TRUE), scale. (FALSE), na.impute (TRUE). MDS: metric (TRUE). UMAP: n.neighbors (15), min.dist (0.1), metric ("euclidean"), n.epochs (NULL). t-SNE: perplexity (30), theta (0.5), max.iter (1000). PHATE: k (5), a (NULL), t ("auto"), gamma (1).
modelsAs samples, on the model axis. NULL (default) embeds every model.

Value

A numeric matrix, n.samples x n.components for compute$embedding$samples() and n.models x n.components for compute$embedding$models(). Rows carry no names: row i is the study's sample (model) i, and is NA wherever samples/models excluded it. Columns are named <METHOD>1..<METHOD>k.

The matrix carries a PolyGeniusProvenance() readable with provenance(). Its $misc holds method, the delivered n.components, the layer read (or "pre-computed similarity"), whether a subset was applied, n.samples (or n.models), and the method's own extras -- PCA sdev / var.explained / var.pct / n.missing / dropped.cols, metric MDS eigenvalues / GOF, non-metric MDS stress, t-SNE the applied perplexity and final.cost. Its $params records every supplied argument as the expression the caller wrote, so a layer or a similarity matrix is named rather than retained.

The sample-side matrix is reduction()-tagged, so data$samples$<name> <- ... files it as an embedding rather than a table member; the model-side matrix carries no such tag. Storing the result is the caller's job either way.

Details

Columns are named <METHOD><k> -- PCA1, UMAP1, MDS1, TSNE1, PHATE1 -- and are deliberately distinct from compute$populationStructure()'s PC1..PCn: PC1 is genotype structure, PCA1 is a reduction of a study slot, and the two are not interchangeable as covariates.

For compute$embedding$models(), a layer that is shaped like a score matrix (n.samples x n.models) is transposed so models are the embedded rows. The shape test needs the two axis lengths to differ, so a study with exactly as many models as samples is embedded untransposed.

"pca" requires irlba, "umap" requires uwot, "tsne" requires Rtsne and "phate" requires phateR; each aborts when its package is missing. Non-metric MDS (metric = FALSE) falls back to metric MDS with a warning when MASS is missing. PCA drops all-NA columns and mean-imputes the remaining missing values, warning as it does. t-SNE lowers perplexity to floor((n - 1) / 3), with a warning, when the requested value is too large for the row count.

The call aborts when n.components is below 1, when layer is absent and similarity was not supplied, when layer does not resolve to a matrix or data.frame, when similarity is not square or does not cover every sample (model), or when a samples/models expression resolves to fetched columns rather than a subset.

Examples

# PCA of the score profiles, kept as a sample embedding
data$samples$PCA <- compute$embedding$samples(data, scores, method = "pca", n.components = 5)

# UMAP with a non-default neighbourhood
umap.emb <- compute$embedding$samples(data, scores, method = "umap", n.neighbors = 30)

# MDS from a pre-computed similarity matrix
sim <- compute$similarity$samples(data, X, method = "pearson")
mds.emb <- compute$embedding$samples(data, similarity = sim, method = "mds")

# Restrict the computation to a subset; excluded samples come back as NA rows
pca.subset <- compute$embedding$samples(data, scores, method = "pca", samples = age > 65)

# Model-side embedding from the same score layer (transposed internally)
pca.models <- compute$embedding$models(data, scores, method = "pca")

See Also

Aliases: compute-embedding, compute.embedding.samples, compute.embedding.models, compute$embedding, compute$embedding$samples, compute$embedding$models