Comparing and Choosing Models
Predictive performance, incremental value, stratification, and how a shortlist is built
You have a set of candidates from chapter 09, scored into a layer with compute$scores(). This chapter is how you ask which one to carry forward: how well each predicts, what it adds beyond covariates you already collect, how outcome risk looks across score bands, and how the candidates compare with and overlap each other.
associate and evaluate both fit models on the same kind of data, but they answer different
questions. associate$regression() gives you one trustworthy adjusted effect. evaluate is about
predictive value and ranking a set of candidates against each other — see
Choosing an Association Analysis for the fuller contrast.
The score layer is used as given
Every evaluate$*() call reads one score layer (scores.layer, default X) and takes it exactly
as it stands: nothing is scored or standardised on your behalf, and a layer that is not already on
data aborts, naming the layers that are. If you want per-SD effects, standardise the layer
yourself first with compute$standardize() before calling evaluate.
Outcomes: binary, continuous, and the liability scale
An outcome is binary or continuous, inferred from its values unless you set the type explicitly
with outcome(). A surv() outcome aborts — evaluate has no survival component; a time-to-event
question belongs to survival associations.
# Type inferred: a two-level factor is binary, a numeric column with more than
# two distinct values is continuous.
evaluate$performance(data, outcomes = c(dementia = diagnosis, bmi = bmi))A binary outcome that needs an explicit event level, a transform, or a population prevalence takes
the same outcome() descriptor associate uses:
outcome(diagnosis, event = "case", prevalence = 0.05)prevalence is what unlocks the liability-scale R-squared metrics, r2.liability and
delta.r2.liability (Lee et al. 2012, eq. 15): the variance a score explains on the underlying
liability scale rather than the observed 0/1 scale, which is what you want when your sample's case
fraction does not match the population's — the usual situation in a case-enriched cohort. One
prevalence applies to every stratum, and setting it on an outcome that resolves continuous aborts.
What each component answers
Each is its own call, so you only pay for what you ask, and each returns a PolyGeniusEvaluation
with the same shape: $results, $artifacts, $diagnostics, $metadata, $provenance.
evaluate$performance() — how well the score alone predicts: AUC (and liability R-squared, with
a prevalence) for a binary outcome, R-squared for a continuous one.
evaluate$incremental() — what the score adds to a model already fitted on your covariates: a
likelihood-ratio test, plus the AUC or R-squared delta. Requires covariates.
evaluate$association() — an odds ratio or beta per model, adjusted for whatever covariates
you pass, by delegating straight to
associate$regression(). It is the same estimate associate$regression()
would give you for the same call, placed beside evaluate's other metrics; reach for
associate$regression() directly for interactions, another regression family, or the fitted models
themselves.
evaluate$stratification() — cuts each model's score into percentile bands (default
c(0.2, 0.8), giving three bands) and reports the outcome in every band, plus a covariate-adjusted
contrast of each band against a reference band.
evaluate$compare() — each model's metric minus one reference model's, as a paired contrast
(DeLong for AUC, bootstrap otherwise).
evaluate$redundancy() — how alike two models are: score correlation, or variant-set overlap.
Takes no outcome.
evaluate$profile() — the components above that apply, in one call and one object.
One call: evaluate$profile()
For most candidate-shortlisting work, profile() is the entry point:
ev <- evaluate$profile(
data,
outcomes = outcome(diagnosis, event = "case", prevalence = 0.05),
covariates = c(age, sex, PC1, PC2)
)It always runs performance, association and stratification. It adds incremental only when you
supply covariates — leaving them out silently drops one of the three components select() ranks
on. With two or more models in the layer it also runs redundancy (score correlation only) and
compare, ranking the candidates with select() first and contrasting every model against the
per-outcome top unless you name reference.model.
$results binds every component's rows, told apart by analysis. A model's n is not the same
number in every row: performance and compare take no covariates, so their complete-case count
can differ from association, incremental and stratification's.
Building a shortlist: evaluate$select()
evaluate$select() reads an evaluation and ranks its models per outcome:
ranked <- evaluate$select(ev)
ranked[order(outcome, rank)]It scores only the performance and incremental metrics with a "higher is better" direction. By
default that means AUC plus its incremental delta, and the incremental liability R-squared delta
when a prevalence is set (continuous outcomes take R-squared and its delta instead); each is
min-max scaled across the candidate models and averaged into one number per model. metrics takes one name, an
unweighted vector, or a named weight vector (c(auc = 2, delta.auc = 1)) to change the composite.
The returned table has outcome, model, score, rank and tied. tied is how you read
whether two models are distinguishable: it is TRUE for the top model, and for every other model it
reads a compare row between it and the top — TRUE when that contrast's interval crosses zero,
NA when no such contrast exists. evaluate$profile() runs compare against select's own top, so
tied is populated for a profile result out of the box.
The score is relative to the candidate set — add or drop one model and every other model's score moves — so it is a ranking, not a number to report on its own.
The tune/test pattern
Selecting and reporting on the same samples overstates the model you picked: it won this
comparison partly by chance, and nothing about select() or compare() corrects for that. Split
your samples before you rank, keeping relatives together in one split:
tune <- subset(data, samples = split == "tune")
test <- subset(data, samples = split == "test")
ranked <- evaluate$select(
evaluate$profile(tune, outcomes = c(case = diagnosis), covariates = c(age, sex, PC1, PC2))
)
best <- ranked[rank == 1L, model]
evaluate$profile(
subset(test, models = best),
outcomes = c(case = diagnosis),
covariates = c(age, sex, PC1, PC2)
)split.by is not a substitute for this. Its levels are not held out — every level still contributes
to the "all" rows evaluate$select() reads — so stratifying by, say, sex does not give you an
honest post-selection estimate the way a disjoint tune/test split does.
Strata with split.by
split.by adds one set of rows per level of a factor, character or logical column, beside the
overall stratum = "all" rows. A sample with a missing split.by value enters "all" only, and a
covariate that is also named in split.by is dropped from the within-level fits, where it would be
constant:
evaluate$association(data, outcomes = diagnosis, covariates = c(age, PC1, PC2), split.by = sex)evaluate$stratification() cuts its score bands once, on the "all" sample, and reuses the same
breaks in every level — a level's bands are not its own percentiles.
Multiple testing
Each component adjusts its own p-values once, with p.adjust.method (default "BH"), within
analysis, outcome, metric and stratum. evaluate$profile() keeps each component's families as they
were and never re-adjusts across them; neither does merge()/c() of two evaluations.
evaluate$performance() and evaluate$redundancy() carry no p-value at all.
Worth stating plainly: with the same covariates, evaluate$association()'s Wald test and
evaluate$incremental()'s likelihood-ratio test address the same null hypothesis — whether the
score has a nonzero effect once the covariates are accounted for. They are separate adjustment
families, not two independent pieces of evidence, and treating agreement between them as
confirmation double-counts the same test.
Interpretation caveats
Every metric here is in-sample: nothing in evaluate cross-validates or holds out data for you.
The tune/test split above is how you turn a selection into an honest estimate; without it, treat a
number as descriptive of this sample rather than as a generalisation claim.
evaluate$stratification()'s bands are percentiles of this sample, not of a reference population
— the same label covers a different score range in another cohort. Its odds ratio is conditional on
the covariates and, because odds ratios do not collapse, can differ from the unadjusted contrast
even without confounding.
evaluate$compare()'s default reference is the best-scoring model in the same data it is contrasted
against, which biases the contrasts away from zero. An honest comparison names a reference chosen on
independent data — a tuning split ranked with evaluate$select(), as above.
Evaluations do not meta-analyse across cohorts. merge()/c() on two PolyGeniusEvaluation
objects concatenate rows — useful for stitching a profile's own components together, or labelling
several cohorts' results with .id for a side-by-side table — but nothing pools an estimate across
them the way associate$meta() pools association results.
The plots
Each component has a matching plot: visualize$evaluate$performance(),
incremental(), association(), stratification(), compare() and redundancy() each draw that
component's own result rows — a ROC or precision-recall curve for a binary performance() result, a
binned score-outcome view for a continuous one, forests for incremental/association/compare,
band-by-band contrasts for stratification, a model-by-model tile grid for redundancy.
visualize$evaluate$profile() is the composite view for a profile() result: one row per model,
one column per metric, filled points marking the models tied calls indistinguishable from the top.
plot() on any single-analysis evaluation dispatches to that analysis's own plot; on a
profile()-shaped object it dispatches to visualize$evaluate$profile() directly.
What to do next
Choosing an Association Analysis for when to reach for associate
instead. Construction Algorithms for where the candidates in your score
layer come from. Plots for how to save and compose the figures above.