Contents
evaluate
Evaluation analyses for PolyGenius
evaluate bundles the predictive assessment of the models in a score layer.
Each component takes a PolyGeniusStudy and returns a
PolyGeniusEvaluation with $results, $artifacts, $diagnostics,
$metadata and $provenance. evaluate$select() is the exception: it reads
such an object and returns a data.table ranking.
Usage
evaluateValue
An environment of class evaluate, holding the components listed
under Format.
Details
Reach for evaluate when the question is predictive: how well each score
discriminates or explains the outcome, what it adds to the covariates, and
which model to carry forward. Reach for associate when the question is an
effect estimate: its size, how it changes under adjustment, or a survival or
interaction model. evaluate$association() places the main-term effects from
associate$regression() beside the metrics.
| Call | Families fitted | Result class | Schema(s) |
|---|---|---|---|
| evaluate$performance() | direct lm/glm fit |
PolyGeniusEvaluation | schema-evaluation |
| evaluate$incremental() | direct lm/glm fit |
PolyGeniusEvaluation | schema-evaluation |
| evaluate$association() | LinearRegression or LogisticRegression, through associate$regression() |
PolyGeniusEvaluation | schema-evaluation |
| evaluate$stratification() | direct lm/glm fit |
PolyGeniusEvaluation | schema-evaluation |
| evaluate$compare() | none | PolyGeniusEvaluation | schema-evaluation |
| evaluate$redundancy() | none | PolyGeniusEvaluation | schema-evaluation |
| evaluate$profile() | those of the components it runs | PolyGeniusEvaluation | schema-evaluation |
| evaluate$select() | none, reads an evaluation and returns a data.table |
none | none |
The score layer must already exist on data, or the call aborts. It is used
as supplied: nothing is scored or standardised.
ev <- evaluate$profile(
data,
outcomes = outcome(case, prevalence = 0.05),
covariates = c(age, sex, PC1)
)
# Rank the models, best first, per outcome.
ranking <- evaluate$select(ev)
Outcomes
An outcome is binary or continuous. The type is inferred from the values
unless an outcome() descriptor sets it. A surv() outcome aborts.
outcome(prevalence = ) gives a binary outcome's population prevalence and
enables the liability-scale R-squared metrics of Lee et al. (2012),
r2.liability and delta.r2.liability. A prevalence must be a single
number in (0, 1). It aborts on a continuous outcome. One prevalence applies
to every stratum. evaluate$redundancy() takes no outcome.
Complete cases
Complete cases are taken per model: the outcome, the covariates and that
model's score. evaluate$compare() takes them per pair of models. An all-NA
score column gives NA rows for that model only. Every row with an outcome
reports its own n and n.events. evaluate$performance() and evaluate$compare() take no
covariates, so one model's n can differ between the components of a
profile.
Strata
split.by adds rows for each of its levels beside the overall
stratum = "all" rows. A sample with a missing split.by value enters
"all" only. Several split.by columns give one stratum per observed
combination of their levels, labelled with the levels joined by ".". A
stratum labelled "all" aborts, as does a split.by column that is not
factor, character or logical. A covariate that is also a split.by column
is dropped from the fits within levels, where it is constant.
Results columns
$results has these columns, in this order:
analysis-- the component that produced the row, e.g."performance".outcome-- the outcome label. NA on redundancy rows.outcome.type--"binary"or"continuous". NA on redundancy rows.model-- the model label, a column name of the score layer.model.ref-- the reference model on compare rows, the pair's second model on redundancy rows. NA elsewhere.stratum--"all"or asplit.bylevel. Never NA.tier-- the score percentile band on stratification rows, e.g."0-20%". NA elsewhere.tier.ref-- the reference band on stratification contrast rows. NA elsewhere.metric-- the metric name.analysisandmetrictogether identify a metric.estimate-- the point estimate.lower-- the lower interval bound atconf.level. NA when no interval is computed.upper-- the upper interval bound atconf.level. NA when no interval is computed.pval-- the raw p-value. NA except on the rows listed under Multiple testing.adj.pval--pvaladjusted withp.adjust.method. NA wherepvalis.adj.family.id-- the adjustment family, a key into$indices$multiplicity. NA wherepvalis.n-- the samples behind the estimate. On a band row, the band's samples. NA on redundancy rows.n.events-- the cases among them. NA for a continuous outcome and on redundancy rows.
schema-evaluation generates the column, metric and artifact tables.
Multiple testing
Only these rows carry a p-value:
- incremental
likelihood.ratio, the likelihood-ratio test of adding the score to the covariates; - association
odds.ratioandbeta, the Wald test fromassociate$regression(); - stratification
odds.ratioandmean.difference, the Wald test of each band against the reference band; - compare
delta.auc, the paired DeLong test.
Every other row has NA pval, adj.pval and adj.family.id. Each component
adjusts its own p-values once, with p.adjust.method, within analysis x
outcome x metric x stratum. A stratification family therefore spans every
band and model of one outcome and stratum. evaluate$profile() keeps each
component's families and does not re-adjust. Neither does merge() or
c() of two evaluations. With the same covariates, the association Wald
test and the incremental likelihood-ratio test address the same null
hypothesis. They are separate families, not independent evidence. The
families are recorded in $indices$multiplicity.
Intervals
Bootstrap intervals are percentile intervals over bootstrap replicates.
For a binary outcome, cases and controls are resampled separately. Every
replicate refits every model it evaluates. A failed replicate is dropped and
counted in $diagnostics. Resampling draws from the caller's random number
stream, so calling set.seed() first makes it reproducible. bootstrap = 0
leaves the bootstrap intervals NA.
Analytic intervals do not depend on bootstrap: DeLong for auc and for
compare's delta.auc, Wald for the association and stratification contrasts,
Wilson score for outcome.rate, and t for outcome.mean.
likelihood.ratio has no interval.
Diagnostics and metadata
$diagnostics$messages is one table with columns analysis, outcome,
model, stratum and message. It records failed fits, strata where a
binary outcome has a single class, degenerate bands, identical score
columns, and failed bootstrap replicates as "k of B replicates failed".
$metadata records the arguments the component takes. evaluate$compare()
adds reference.model, named by outcome, and reference.selected, TRUE
when the reference was data-selected. evaluate$stratification() adds
quantiles and reference.
Selecting models
evaluate$select() ranks the models of an evaluation on its performance and
incremental metrics. Selecting and reporting on the same samples overstates
the chosen model. Select on one split of the samples, then profile the chosen
models on a disjoint split. Keep relatives in the same split, so the test
split shares no relatives with the one that chose. split.by is not a
substitute: its levels are not held out, and evaluate$select() reads only
the "all" rows, which pool them.