Contents
compute$standardize
Standardize a score layer
Builds a new score layer from an existing one, centered and scaled per model, and tags every
model with the (center, scale, scale.source) triple used. Returns the standardized matrix and
writes nothing into data; the caller assigns it into a new data$scores layer.
Usage
compute$standardize(
data,
layer = X,
scale.source = c("sample", "reference.panel", "external"),
center = NULL,
scale = NULL,
by.genotype = FALSE
)Arguments
| Argument | Description |
|---|---|
data | A PolyGeniusStudy. Study holding the layer to standardize. |
layer | Unquoted score-layer name, default X; a character scalar (layer = "X") is accepted too. Names an existing layer in data$scores, and aborts when that layer is absent or has no column names. |
scale.source | One of "sample" (default), "reference.panel", "external". Where center/scale come from. "sample" — Computed here, per model, as colMeans() and the per-column sample standard deviation of data$scores[[layer]]. Supplying center/scale aborts. "reference.panel", "external" — Nothing is computed; center and scale must both be supplied, and the source is recorded as given. PolyGenius does not derive reference-panel statistics itself. Not supported together with by.genotype = TRUE. |
center | Numeric scalar, a numeric vector named by model, or NULL (default). Required, and only used, when scale.source is "reference.panel" or "external". A scalar is recycled across every model in layer; a named vector must cover at least the layer's own models, and a superset such as a whole panel-derived table is fine. An unnamed vector of length above one aborts, as does any scale that is not strictly positive for every model. |
scale | Numeric scalar, a numeric vector named by model, or NULL (default). Required, and only used, when scale.source is "reference.panel" or "external". A scalar is recycled across every model in layer; a named vector must cover at least the layer's own models, and a superset such as a whole panel-derived table is fine. An unnamed vector of length above one aborts, as does any scale that is not strictly positive for every model. |
by.genotype | Logical scalar, default FALSE. When TRUE, center/scale are computed separately within each genotype fileset (data$sample.names.by.genotype) rather than pooled, which is the right choice when filesets differ in array, QC or ancestry mix. Aborts when scale.source is not "sample", when a row cannot be matched to a genotype, or when any genotype contributes fewer than two samples, since its in-sample standard deviation is then undefined. |
Value
A numeric matrix with the same dimensions and dimnames as data$scores[[layer]], carrying a
transform diagnostics entry (diagnostics(x, "transform")): a data.frame, one row per
model in colnames(data$scores[[layer]]) order, with columns model, center, scale,
scale.source and by.genotype. A model that could not be scaled has an all-NA column and a
scale of NA; its center is kept, and is itself missing when the model has no score at all. That entry stays attached to this layer only — assigning the
layer into data$scores copies it nowhere else; see diagnostics().
With by.genotype = TRUE a single (center, scale) pair per model cannot represent the
transform, so both are recorded as NA in the transform entry while scale.source and
by.genotype still identify what was done. The full per-(model, genotype) breakdown — model,
genotype, n, center, scale — is written to provenance(x)$misc$by.genotype instead.
Details
Recording the transform is what keeps a PRS coefficient fit on a per-SD layer distinguishable from one fit on a per-raw-unit layer downstream — otherwise both arrive as ordinary coefficients under the same key.
A model whose computed scale is 0 or missing cannot be standardized: it gave all samples the
same score (SD 0), it has fewer than two non-missing scores, or it has an infinite score. Its
column comes back NA, the
other models are standardized as usual, and one warning names every such model and why. A
model scoring every sample alike has usually scored none of its variants, so the warning points
to diagnostics(data$scores$<layer>, "coverage"). Under by.genotype = TRUE only that
genotype's rows come back NA, with one warning per genotype naming it. A supplied scale
that is not positive is a bad argument, and aborts naming the models.
Examples
data$scores$X.scaled <- compute$standardize(data, layer = X)
data$scores$X.panel <- compute$standardize(
data, layer = X, scale.source = "reference.panel",
center = panel.center, scale = panel.scale)
data$scores$X.local <- compute$standardize(data, layer = X, by.genotype = TRUE)See Also
compute$scores(), which produces the layer this standardizes.