Contents
compute$embedding
Low-dimensional embeddings of a study's samples or models
compute$embedding$samples() embeds samples and compute$embedding$models() embeds
models, from any matrix-shaped study slot or from a pre-computed similarity matrix
(PCA, MDS, UMAP, t-SNE, PHATE). Both return a full-sized coordinate matrix and write
nothing into data.
Usage
compute$embedding$samples(
data,
layer,
method = "pca",
similarity = NULL,
n.components = 2,
samples = NULL,
...
)
compute$embedding$models(
data,
layer,
method = "pca",
similarity = NULL,
n.components = 2,
models = NULL,
...
)Arguments
| Argument | Description |
|---|---|
data | A PolyGeniusStudy. Read-only. |
layer | Unquoted slot expression, resolved by data$fetch() -- e.g. scores, samples$PCA, sample.pairs$kinship. Required unless similarity is supplied, and ignored when it is. |
method | One of "pca" (default), "mds", "umap", "tsne", "phate". |
similarity | A square numeric matrix over every sample (model), or NULL (default). Replaces layer when supplied. For method = "mds" a matrix whose mean diagonal exceeds 0.5 is read as a similarity and converted to 1 - similarity. UMAP and t-SNE then take it as a precomputed distance; PCA and PHATE embed it as an ordinary matrix. |
n.components | Positive integer scalar, default 2. Number of dimensions requested. For "pca" it must not exceed min(n.rows - 1, n.cols) of the matrix being embedded: PCA truncates to that many components, and the call then fails while naming the columns. |
samples | Character vector of sample names, numeric or logical index, tidyselect call, a predicate over data's own columns (e.g. age > 65), or NULL (default) for every sample. Excluded samples are NA rows in the returned matrix. |
... | Method-specific arguments. PCA: center (TRUE), scale. (FALSE), na.impute (TRUE). MDS: metric (TRUE). UMAP: n.neighbors (15), min.dist (0.1), metric ("euclidean"), n.epochs (NULL). t-SNE: perplexity (30), theta (0.5), max.iter (1000). PHATE: k (5), a (NULL), t ("auto"), gamma (1). |
models | As samples, on the model axis. NULL (default) embeds every model. |
Value
A numeric matrix, n.samples x n.components for
compute$embedding$samples() and n.models x n.components for
compute$embedding$models(). Rows carry no names: row i is the study's sample (model)
i, and is NA wherever samples/models excluded it. Columns are named
<METHOD>1..<METHOD>k.
The matrix carries a PolyGeniusProvenance() readable with provenance(). Its $misc
holds method, the delivered n.components, the layer read (or
"pre-computed similarity"), whether a subset was applied, n.samples (or
n.models), and the method's own extras -- PCA sdev / var.explained / var.pct /
n.missing / dropped.cols, metric MDS eigenvalues / GOF, non-metric MDS
stress, t-SNE the applied perplexity and final.cost. Its $params records every
supplied argument as the expression the caller wrote, so a layer or a similarity matrix is
named rather than retained.
The sample-side matrix is reduction()-tagged, so data$samples$<name> <- ... files
it as an embedding rather than a table member; the model-side matrix carries no such
tag. Storing the result is the caller's job either way.
Details
Columns are named <METHOD><k> -- PCA1, UMAP1, MDS1, TSNE1, PHATE1 -- and are
deliberately distinct from compute$populationStructure()'s PC1..PCn: PC1 is
genotype structure, PCA1 is a reduction of a study slot, and the two are not
interchangeable as covariates.
For compute$embedding$models(), a layer that is shaped like a score matrix
(n.samples x n.models) is transposed so models are the embedded rows. The shape test
needs the two axis lengths to differ, so a study with exactly as many models as samples
is embedded untransposed.
"pca" requires irlba, "umap" requires uwot, "tsne" requires Rtsne and
"phate" requires phateR; each aborts when its package is missing. Non-metric MDS
(metric = FALSE) falls back to metric MDS with a warning when MASS is missing. PCA
drops all-NA columns and mean-imputes the remaining missing values, warning as it does.
t-SNE lowers perplexity to floor((n - 1) / 3), with a warning, when the requested
value is too large for the row count.
The call aborts when n.components is below 1, when layer is absent and similarity
was not supplied, when layer does not resolve to a matrix or data.frame, when
similarity is not square or does not cover every sample (model), or when a
samples/models expression resolves to fetched columns rather than a subset.
Examples
# PCA of the score profiles, kept as a sample embedding
data$samples$PCA <- compute$embedding$samples(data, scores, method = "pca", n.components = 5)
# UMAP with a non-default neighbourhood
umap.emb <- compute$embedding$samples(data, scores, method = "umap", n.neighbors = 30)
# MDS from a pre-computed similarity matrix
sim <- compute$similarity$samples(data, X, method = "pearson")
mds.emb <- compute$embedding$samples(data, similarity = sim, method = "mds")
# Restrict the computation to a subset; excluded samples come back as NA rows
pca.subset <- compute$embedding$samples(data, scores, method = "pca", samples = age > 65)
# Model-side embedding from the same score layer (transposed internally)
pca.models <- compute$embedding$models(data, scores, method = "pca")