Plots
What each visualize family draws, what comes back, and how to get a figure out
By now you have drawn a few plots without much explanation. This chapter is the systematic version.
The families
visualize$data$* draws things about your cohort: score distributions, sample or model similarity, embeddings.
visualize$models$* draws things about a model library: sizes, variant reuse, uniqueness, top variants.
visualize$associations$* draws association results: forest plots, heatmaps, survival curves, QQ plots, and a landscape.
The landscape is the figure for a scan over many scores: one point per score, grouped along the x-axis by any grouping you give it, at the minus-log p-value of the column you name. group.by = group.of[predictor] reads a lookup vector of your own, and statistic = "signed.neglog10p" splits risk from protective associations above and below zero.
It reads the p-value column as it stands; it never adjusts one. The threshold line is only as meaningful as that column: against the default adj.pval, a threshold of 0.05 is an FDR cut, against raw pval it is nominal. A result with nothing to plot draws an empty frame rather than stopping.
visualize$evaluate$* draws evaluation results: one plot per component — performance, incremental value, association, stratification, comparison, redundancy — plus profile(), the composite view of an evaluate$profile() result.
visualize$genome$* composes signals on a shared genome axis — chapter 12.2.
visualize$palette$* is the colour system — chapter 12.1.
visualize$execution$* renders run dashboards, which chapter 10 uses.
Calling plot() on a result gives you that result's default figure, which is often what you want.
The rule that shapes all of it
A plot renders what the analysis already emitted. It does not compute the number it shows.
In practice: if a plot says a quantity is not available, no argument will conjure it — you re-run the analysis so that it emits the artifact. The fix is upstream, always.
The payoff is that a figure and the table it came from cannot disagree.
One place currently crosses that line, and you should know it: the QQ plot recomputes the genomic inflation factor when the analysis did not record one, using the filtered p-values rather than all of them. It is being closed by having the analysis emit what the plot needs. The evaluate performance plots read their AUC and R-squared from the result row rather than re-fitting, so they do not.
Which plot needs which analysis
Artifacts arrive as a bundle per analysis. That is the rule that makes this predictable: a
stratification result cannot draw a confusion-matrix ROC curve, because performance was never run.
An evaluate$profile() result can draw everything its own components produced.
| Plot family | Needs |
|---|---|
| performance curves, ROC, precision-recall | a confusion artifact (binary) or a binned score-outcome grid (continuous) |
| score distribution, predicted-vs-observed, residuals | a per-sample score artifact |
| stratification bands and contrasts | no artifact — read straight from the result rows |
| redundancy heatmaps | a similarity matrix (paired needs two) |
| survival curves | predicted or observed curves |
A missing artifact is a clear error naming what is absent, not a blank panel. The two usual causes are that a filter removed everything, or that the artifact belongs to an analysis you did not run.
After filtering a result, artifacts are pruned along with the rows — so a filtered object's plots match its table.
Empty results error
An empty or all-missing result does not render an empty frame; it stops with a message. That is worth knowing so you recognise it as a filtering problem rather than a bug. The genome overview is the one exception: it drops what it cannot build, warns, and draws the rest.
What comes back
Five shapes, and the difference is visible to you:
A ggplot — add to it with + as usual.
A patchwork of several panels — forest plots, survival plots, genome stacks. + lands on the last panel only; use & to reach them all.
A ComplexHeatmap — most heatmaps. No ggplot theming applies, and it is drawn rather than printed. visualize$evaluate$redundancy() is not one of these — it draws a geom_tile ggplot.
A genome track — a specification you can save and re-render, not a plot until you print it.
A named list of plots — the execution performance family.
One special case: some functions return either a single heatmap or a list of them depending on whether you faceted.
The evaluate profile plot, specifically
visualize$evaluate$profile() is the composite view of an evaluate$profile() result: one row per model, one column per metric, faceted by outcome. A model's point is filled when evaluate$select()'s tied is TRUE for it and hollow otherwise — see chapter 11 for how to read that. group.by turns the rows into panels of models grouped your way, each sorted by score within its group.
Theming, briefly
Most single-panel plots take a theme switch with two settings: the package look, or a plain one. Colours are applied independently, so the plain setting is still on-brand.
Some plots do not take it at all — forest and survival plots are finished multi-panel layouts, and heatmaps are not ggplots. Chapter 12.1 covers the rest.
Getting a figure out
This depends on which shape you have, which is why it comes last:
- ggplot or patchwork —
ggsave(). - Heatmap — open a device,
draw()it, close the device. - Genome track —
plot()it to get a ggplot, then save that. - A list of plots — pick one first.
For a publication-ready result: set the base size to your journal's font size, use the palette accessors so a global palette change reaches your own plots too, and know that heatmap typography is not currently under the package's control — so a heatmap will not match an 8-point ggplot panel in a composed figure.
On coverage
Test coverage is uneven rather than thin overall. The genome family accounts for most of it; survival plots, evaluate plots and every heatmap have none. Check those by eye before publishing, and treat "it runs" as different from "it is right".