PolyGenius
Contents

compute$similarity

Similarity and distance matrices between samples or between models

compute$similarity$samples() compares samples by their PRS score profiles; compute$similarity$models() compares models by SNP overlap or by their score vectors. Both return a full-sized square matrix in the study's own axis order, whether or not the computation was subsetted.

Usage

compute$similarity$samples(
  data,
  layer,
  method = c("pearson", "spearman", "euclidean", "manhattan", "cosine"),
  samples,
  use = "pairwise.complete.obs",
  ...
)

compute$similarity$models(
  data,
  layer,
  method = c("score.pearson", "score.spearman", "snp.jaccard", "snp.weighted.overlap"),
  models,
  use = "pairwise.complete.obs",
  ...
)

Arguments

ArgumentDescription
dataA PolyGeniusStudy. Read-only.
layerUnquoted score-layer name, e.g. X; a character scalar naming the layer also works. Required for compute$similarity$samples() and for the score.* model methods, and must name a key in data$scores.keys. Unused by snp.jaccard and snp.weighted.overlap.
methodMetric to compute. For compute$similarity$samples() one of "pearson" (default), "spearman", "euclidean", "manhattan", "cosine"; for compute$similarity$models() one of "score.pearson" (default), "score.spearman", "snp.jaccard", "snp.weighted.overlap". "euclidean" and "manhattan" return distances, where a low value means similar; every other method returns a similarity.
samplesOptional subset of samples: a character vector of names, a numeric or logical index, a tidyselect call, or a predicate over data's own sample table (e.g. age > 65). Omit it to use every sample. Only the selected samples enter the computation; the rest are NA rows and columns of the returned matrix.
useCharacter scalar, default "pairwise.complete.obs". Missing-value handling passed to stats::cor(); ignored by the distance and snp.* methods.
...Unused. Reserved for future method-specific arguments.
modelsOptional subset of models: a character vector of names, a numeric or logical index, a tidyselect call, or a predicate over data's own model table (e.g. gwas$trait == "AD"). Omit it to use every model. The rest are NA rows and columns of the returned matrix.

Value

A numeric square matrix: n.samples x n.samples for compute$similarity$samples(), n.models x n.models for compute$similarity$models(). Rows and columns carry no names: row and column i are the study's sample (model) i, and are NA wherever samples/models excluded it. The matrix carries a PolyGeniusProvenance() readable with provenance(), whose $misc holds data.name, layer, method, whether a subset was applied, the computed and total axis counts, and use for the correlation methods.

Details

Neither call writes into data. Store the result yourself, on the matching pair plane -- data$sample.pairs$<name> or data$model.pairs$<name>. That is why a subsetted computation still comes back full-sized: those planes validate a matrix against the full axis length and reject a shrunken one.

coop is used for Pearson correlation and for cosine similarity when it is installed, base R otherwise.

Each sample is taken as its vector of scores across models, and the metric is applied between those vectors: correlation for scale-invariant profile shape, cosine for direction alone, Euclidean or Manhattan for absolute separation.

A samples/models expression that selects nothing aborts, as does one that resolves to fetched columns rather than to a subset.

snp.jaccard and snp.weighted.overlap read each model's own variant table through data$library and key a variant on its canonical chr:position:allele-pair identity, so a model that recorded the allele pair swapped still matches. snp.jaccard divides the number of shared keys by the number of keys in the union; snp.weighted.overlap sums abs(beta.i * beta.j) over the shared keys and divides by sqrt(sum(beta.i^2) * sum(beta.j^2)). Two rows of one model that land on the same canonical key are summed first, so a within-model duplicate raises that model's own normalizer. Both methods force the diagonal to 1, and set a pair of models built on different genome builds to NA rather than 0, since their variant keys are not comparable. Both also abort before computing anything when the estimated cost is over the package ceiling: compare fewer models, or narrow their variant sets.

score.pearson and score.spearman correlate the models' score columns in data$scores[[layer]] across samples, so they need layer and the snp.* methods do not.

Functions

  • compute(similarity.samples): Sample-sample similarity from score profiles
  • compute(similarity.models): Model-model similarity from SNP overlap or score vectors

Examples

# Correlation between samples' score profiles
data$sample.pairs$cor <- compute$similarity$samples(data, X, method = "pearson")

# A distance instead -- low values mean similar
data$sample.pairs$dist <- compute$similarity$samples(data, X, method = "euclidean")

# Subset by name, or by a predicate over the sample table
data$sample.pairs$pair <- compute$similarity$samples(data, X, samples = c("subj_1", "subj_2"))
data$sample.pairs$elderly <- compute$similarity$samples(data, X, samples = age > 65)

```



```r

# SNP overlap -- no score layer needed
data$model.pairs$jaccard <- compute$similarity$models(data, method = "snp.jaccard")
data$model.pairs$weighted <- compute$similarity$models(data, method = "snp.weighted.overlap")

# Correlation of the models' score vectors
data$model.pairs$score.cor <- compute$similarity$models(data, X, method = "score.pearson")

# Subset by name, or by a predicate over the model table
data$model.pairs$best <- compute$similarity$models(data, X, models = c("LDpred2", "CT_5e8"))
data$model.pairs$AD <- compute$similarity$models(data, X, models = gwas$trait == "AD")

See Also

Aliases: compute-similarity, compute.similarity.samples, compute.similarity.models, compute$similarity, compute$similarity$samples, compute$similarity$models