PolyGenius
Contents

visualize$data$scores$distribution

visualize PRS score distributions

Draws one data$scores layer as a per-model distribution -- violin, density, histogram or boxplot -- optionally grouped and faceted by sample phenotypes.

Usage

visualize.scores.distribution(
  data,
  models = NULL,
  scores.layer = X,
  samples = NULL,
  split.by = NULL,
  group.by = NULL,
  label = NULL,
  type = c("violin", "density", "histogram", "boxplot"),
  show.outliers = TRUE,
  show.points = FALSE,
  raster = FALSE,
  color.by = c("model", "group"),
  pattern.args = list(),
  geom.args = list(),
  raster.args = list(),
  facet.args = list(),
  palette = NULL,
  theme = c("polygenius", "none"),
  ...
)

Arguments

ArgumentDescription
dataA PolyGeniusStudy carrying at least one scores layer. A summary-mode study (no genotype backend attached, or an empty samples table) is refused, because this plot reads sample-level scores directly.
modelsUnquoted model selector, or NULL (default): positions (c(1, 3, 5)), bare model names (c(AD_PRS, PD_PRS)), or a tidyselect expression (starts_with("AD")). NULL, or omitting the argument, takes every model when there are 5 or fewer and aborts above that.
scores.layerUnquoted score-layer name, default X. Layer of data$scores to draw; data$scores.keys lists the available layers.
samplesUnquoted sample-side predicate or index vector, or NULL (default) for every sample. Must resolve to a single selector.
split.byUnquoted sample-side expression, or NULL (default) for a single panel. Every level becomes one facet_wrap() panel.
group.byUnquoted sample-side expression, or NULL (default). Every level becomes a separate distribution within each panel; NULL draws one distribution per model.
labelUnquoted model-side expression, or NULL (default) to use the model names. Resolved with data$fetch(.context = "models") and used on the axis and in the legend.
typeOne of "violin" (default), "density", "histogram", "boxplot". Distribution geometry.
show.outliersLogical scalar, default TRUE. Draw boxplot outlier points. Ignored for the other three types.
show.pointsLogical scalar, default FALSE. Overlay individual samples -- jittered points for violin and boxplot, a bottom rug for density. Ignored for histogram.
rasterLogical scalar, default FALSE. Rasterizes the overlaid point layer with ggrastr::geom_point_rast(), so it applies only with show.points = TRUE on a violin or boxplot. Without ggrastr it warns, falls back to geom_point() and drops raster.args.
color.byOne of "model" (default), "group". Which dimension carries color when several models and several group.by levels are both present; the other takes the pattern texture. Ignored when only one dimension varies.
pattern.argsNamed list, default list(). Extra arguments for the ggpattern geoms (e.g. pattern_spacing), merged after the pattern_density = 0.2 default. Used only on the pattern path.
geom.argsNamed list, default list(). Extra arguments for the main geom, e.g. alpha, bins (histogram, 30 otherwise) or bw (density).
raster.argsNamed list, default list(). Extra arguments forwarded to ggrastr::geom_point_rast(), over a default raster.dpi = 300.
facet.argsNamed list, default list(). Extra arguments for facet_wrap(), e.g. ncol, nrow, scales. Used only with split.by.
paletteCategorical color palette: a role or hue name, a vector of two or more colors, a ramp function, or NULL (default) for the package categorical palette.
themeOne of "polygenius" (default), "none". Plot theme. "none" gives a bare theme_minimal() to style yourself; palette colors are applied either way.
...Further arguments passed to data$fetch().

Value

A ggplot object.

Details

Scores are read from data$scores[[scores.layer]] by model position, so two models that share a name stay two models; every sample-side selector is resolved through data$fetch(), and nothing is recomputed. A study with no scores layer at all aborts asking for compute$scores(), and a scores.layer that is not a key of data$scores aborts listing the available keys. A samples, split.by or group.by expression the fetch reports as unresolved or ambiguous aborts naming that argument, rather than silently drawing an unfiltered figure. NA levels of split.by and group.by are kept as their own "NA" level. A name or label shared by more than one drawn model carries a [#<position>] suffix, the model's position in data; it distinguishes models within this figure and is not a persistent identifier.

Color and pattern split the two dimensions only when both vary: with several models and several group.by levels, color.by takes color and the other takes a ggpattern texture. Without ggpattern installed the plot warns once and separates the second dimension by position dodge (violin, boxplot) or by linetype (density, histogram).

Examples

visualize$data$scores$distribution(study, models = 1, type = "density")

# Models coloured, sex drawn as a pattern texture, one panel per diagnosis.
visualize$data$scores$distribution(
  study,
  models    = c(1, 2),
  group.by  = sex,
  split.by  = diagnosis,
  color.by  = "model"
)

See Also

Aliases: visualize.scores.distribution, visualize$data$scores$distribution, visualize_scores_distribution