PolyGenius
Contents

compute$standardize

Standardize a score layer

Builds a new score layer from an existing one, centered and scaled per model, and tags every model with the (center, scale, scale.source) triple used. Returns the standardized matrix and writes nothing into data; the caller assigns it into a new data$scores layer.

Usage

compute$standardize(
  data,
  layer = X,
  scale.source = c("sample", "reference.panel", "external"),
  center = NULL,
  scale = NULL,
  by.genotype = FALSE
)

Arguments

ArgumentDescription
dataA PolyGeniusStudy. Study holding the layer to standardize.
layerUnquoted score-layer name, default X; a character scalar (layer = "X") is accepted too. Names an existing layer in data$scores, and aborts when that layer is absent or has no column names.
scale.sourceOne of "sample" (default), "reference.panel", "external". Where center/scale come from. "sample" — Computed here, per model, as colMeans() and the per-column sample standard deviation of data$scores[[layer]]. Supplying center/scale aborts. "reference.panel", "external" — Nothing is computed; center and scale must both be supplied, and the source is recorded as given. PolyGenius does not derive reference-panel statistics itself. Not supported together with by.genotype = TRUE.
centerNumeric scalar, a numeric vector named by model, or NULL (default). Required, and only used, when scale.source is "reference.panel" or "external". A scalar is recycled across every model in layer; a named vector must cover at least the layer's own models, and a superset such as a whole panel-derived table is fine. An unnamed vector of length above one aborts, as does any scale that is not strictly positive for every model.
scaleNumeric scalar, a numeric vector named by model, or NULL (default). Required, and only used, when scale.source is "reference.panel" or "external". A scalar is recycled across every model in layer; a named vector must cover at least the layer's own models, and a superset such as a whole panel-derived table is fine. An unnamed vector of length above one aborts, as does any scale that is not strictly positive for every model.
by.genotypeLogical scalar, default FALSE. When TRUE, center/scale are computed separately within each genotype fileset (data$sample.names.by.genotype) rather than pooled, which is the right choice when filesets differ in array, QC or ancestry mix. Aborts when scale.source is not "sample", when a row cannot be matched to a genotype, or when any genotype contributes fewer than two samples, since its in-sample standard deviation is then undefined.

Value

A numeric matrix with the same dimensions and dimnames as data$scores[[layer]], carrying a transform diagnostics entry (diagnostics(x, "transform")): a data.frame, one row per model in colnames(data$scores[[layer]]) order, with columns model, center, scale, scale.source and by.genotype. A model that could not be scaled has an all-NA column and a scale of NA; its center is kept, and is itself missing when the model has no score at all. That entry stays attached to this layer only — assigning the layer into data$scores copies it nowhere else; see diagnostics().

With by.genotype = TRUE a single (center, scale) pair per model cannot represent the transform, so both are recorded as NA in the transform entry while scale.source and by.genotype still identify what was done. The full per-(model, genotype) breakdown — model, genotype, n, center, scale — is written to provenance(x)$misc$by.genotype instead.

Details

Recording the transform is what keeps a PRS coefficient fit on a per-SD layer distinguishable from one fit on a per-raw-unit layer downstream — otherwise both arrive as ordinary coefficients under the same key.

A model whose computed scale is 0 or missing cannot be standardized: it gave all samples the same score (SD 0), it has fewer than two non-missing scores, or it has an infinite score. Its column comes back NA, the other models are standardized as usual, and one warning names every such model and why. A model scoring every sample alike has usually scored none of its variants, so the warning points to diagnostics(data$scores$<layer>, "coverage"). Under by.genotype = TRUE only that genotype's rows come back NA, with one warning per genotype naming it. A supplied scale that is not positive is a bad argument, and aborts naming the models.

Examples

data$scores$X.scaled <- compute$standardize(data, layer = X)
data$scores$X.panel  <- compute$standardize(
  data, layer = X, scale.source = "reference.panel",
  center = panel.center, scale = panel.scale)
data$scores$X.local <- compute$standardize(data, layer = X, by.genotype = TRUE)

See Also

compute$scores(), which produces the layer this standardizes.

Aliases: compute-standardize, compute.standardize, compute$standardize