Contents
compute$similarity
Similarity and distance matrices between samples or between models
compute$similarity$samples() compares samples by their PRS score profiles;
compute$similarity$models() compares models by SNP overlap or by their score vectors.
Both return a full-sized square matrix in the study's own axis order, whether or not
the computation was subsetted.
Usage
compute$similarity$samples(
data,
layer,
method = c("pearson", "spearman", "euclidean", "manhattan", "cosine"),
samples,
use = "pairwise.complete.obs",
...
)
compute$similarity$models(
data,
layer,
method = c("score.pearson", "score.spearman", "snp.jaccard", "snp.weighted.overlap"),
models,
use = "pairwise.complete.obs",
...
)Arguments
| Argument | Description |
|---|---|
data | A PolyGeniusStudy. Read-only. |
layer | Unquoted score-layer name, e.g. X; a character scalar naming the layer also works. Required for compute$similarity$samples() and for the score.* model methods, and must name a key in data$scores.keys. Unused by snp.jaccard and snp.weighted.overlap. |
method | Metric to compute. For compute$similarity$samples() one of "pearson" (default), "spearman", "euclidean", "manhattan", "cosine"; for compute$similarity$models() one of "score.pearson" (default), "score.spearman", "snp.jaccard", "snp.weighted.overlap". "euclidean" and "manhattan" return distances, where a low value means similar; every other method returns a similarity. |
samples | Optional subset of samples: a character vector of names, a numeric or logical index, a tidyselect call, or a predicate over data's own sample table (e.g. age > 65). Omit it to use every sample. Only the selected samples enter the computation; the rest are NA rows and columns of the returned matrix. |
use | Character scalar, default "pairwise.complete.obs". Missing-value handling passed to stats::cor(); ignored by the distance and snp.* methods. |
... | Unused. Reserved for future method-specific arguments. |
models | Optional subset of models: a character vector of names, a numeric or logical index, a tidyselect call, or a predicate over data's own model table (e.g. gwas$trait == "AD"). Omit it to use every model. The rest are NA rows and columns of the returned matrix. |
Value
A numeric square matrix: n.samples x n.samples for
compute$similarity$samples(), n.models x n.models for
compute$similarity$models(). Rows and columns carry no names: row and column i are
the study's sample (model) i, and are NA wherever samples/models excluded it. The
matrix carries a PolyGeniusProvenance() readable with
provenance(), whose $misc holds data.name, layer, method, whether a subset
was applied, the computed and total axis counts, and use for the correlation
methods.
Details
Neither call writes into data. Store the result yourself, on the matching pair plane
-- data$sample.pairs$<name> or data$model.pairs$<name>. That is why a subsetted
computation still comes back full-sized: those planes validate a matrix against the
full axis length and reject a shrunken one.
coop is used for Pearson correlation and for cosine similarity when it is installed,
base R otherwise.
Each sample is taken as its vector of scores across models, and the metric is applied between those vectors: correlation for scale-invariant profile shape, cosine for direction alone, Euclidean or Manhattan for absolute separation.
A samples/models expression that selects nothing aborts, as does one that resolves
to fetched columns rather than to a subset.
snp.jaccard and snp.weighted.overlap read each model's own variant table through
data$library and key a variant on its canonical chr:position:allele-pair identity,
so a model that recorded the allele pair swapped still matches. snp.jaccard divides
the number of shared keys by the number of keys in the union; snp.weighted.overlap
sums abs(beta.i * beta.j) over the shared keys and divides by
sqrt(sum(beta.i^2) * sum(beta.j^2)). Two rows of one model that land on the same
canonical key are summed first, so a within-model duplicate raises that model's own
normalizer. Both methods force the diagonal to 1, and set a pair of models built on
different genome builds to NA rather than 0, since their variant keys are not
comparable. Both also abort before computing anything when the estimated cost is over
the package ceiling: compare fewer models, or narrow their variant sets.
score.pearson and score.spearman correlate the models' score columns in
data$scores[[layer]] across samples, so they need layer and the snp.* methods do
not.
Functions
compute(similarity.samples): Sample-sample similarity from score profilescompute(similarity.models): Model-model similarity from SNP overlap or score vectors
Examples
# Correlation between samples' score profiles
data$sample.pairs$cor <- compute$similarity$samples(data, X, method = "pearson")
# A distance instead -- low values mean similar
data$sample.pairs$dist <- compute$similarity$samples(data, X, method = "euclidean")
# Subset by name, or by a predicate over the sample table
data$sample.pairs$pair <- compute$similarity$samples(data, X, samples = c("subj_1", "subj_2"))
data$sample.pairs$elderly <- compute$similarity$samples(data, X, samples = age > 65)
```
```r
# SNP overlap -- no score layer needed
data$model.pairs$jaccard <- compute$similarity$models(data, method = "snp.jaccard")
data$model.pairs$weighted <- compute$similarity$models(data, method = "snp.weighted.overlap")
# Correlation of the models' score vectors
data$model.pairs$score.cor <- compute$similarity$models(data, X, method = "score.pearson")
# Subset by name, or by a predicate over the model table
data$model.pairs$best <- compute$similarity$models(data, X, models = c("LDpred2", "CT_5e8"))
data$model.pairs$AD <- compute$similarity$models(data, X, models = gwas$trait == "AD")