PolyGenius
Contents

evaluate$redundancy

Redundancy between PRS models

evaluate$redundancy() asks how much two models overlap: in the variants they use, or in the scores they give the same samples.

Usage

evaluate$redundancy(
  data,
  scores.layer = X,
  method = NULL
)

Arguments

ArgumentDescription
dataA [PolyGeniusStudy](/reference/polygeniusstudy/). Holds the score layer and the model library. It is not modified.
scores.layerUnquoted name of an existing score layer, default X. Read only by score.pearson. A missing layer aborts only when score.pearson is requested. The layer is used as supplied: nothing is scored or standardised here.
methodCharacter vector drawn from "score.pearson", "snp.jaccard" and "snp.weighted.overlap", or NULL (default) for all three. Any other value aborts.

Value

A PolyGeniusEvaluation following the schema-evaluation schema. $results has the columns listed in evaluate, one row per method and unordered pair of models. model is the earlier of the pair and model.ref the later. With score.pearson, the order and labels are the layer's columns, as the other components label them. Without it, they are the study's model order. metric is the method and estimate the similarity. outcome and outcome.type are NA and stratum is "all". No row carries an interval or a p-value, and n and n.events are NA. There are no artifacts.

Details

Each method is one call to compute$similarity$models(), which defines it exactly:

  • score.pearson is the Pearson correlation of the two models' scores in scores.layer, across the samples where both scores are present.
  • snp.jaccard is the number of shared variants divided by the number in the union of the two variant sets.
  • snp.weighted.overlap is the overlap of the two sets weighted by the absolute product of the effect sizes, normalised to 1 for identical models.

The snp.* methods read each model's variants from the study's model library. A pair of models on different genome builds gets an NA estimate for these methods, because their variant keys are not comparable. They also abort before computing when the estimated cost exceeds the ceiling of compute$similarity$models(). A large library whose models share most variants reaches it. Request method = "score.pearson" then.

There is no outcome and no split.by: every row is on the whole study.

Examples

red <- evaluate$redundancy(data)

# Variant overlap only
red <- evaluate$redundancy(data, method = c("snp.jaccard", "snp.weighted.overlap"))

See Also

evaluate$profile(), which runs this with two or more models; visualize$evaluate$redundancy() plots one method as a model-by-model heatmap.

Other evaluate-components: evaluate, evaluate.association(), evaluate.compare(), evaluate.incremental(), evaluate.performance(), evaluate.profile(), evaluate.select(), evaluate.stratification()

Aliases: evaluate.redundancy, evaluate$redundancy