PolyGenius
Contents

evaluate

Evaluation analyses for PolyGenius

evaluate bundles the predictive assessment of the models in a score layer. Each component takes a PolyGeniusStudy and returns a PolyGeniusEvaluation with $results, $artifacts, $diagnostics, $metadata and $provenance. evaluate$select() is the exception: it reads such an object and returns a data.table ranking.

Usage

evaluate

Value

An environment of class evaluate, holding the components listed under Format.

Details

Reach for evaluate when the question is predictive: how well each score discriminates or explains the outcome, what it adds to the covariates, and which model to carry forward. Reach for associate when the question is an effect estimate: its size, how it changes under adjustment, or a survival or interaction model. evaluate$association() places the main-term effects from associate$regression() beside the metrics.

Call Families fitted Result class Schema(s)
evaluate$performance() direct lm/glm fit PolyGeniusEvaluation schema-evaluation
evaluate$incremental() direct lm/glm fit PolyGeniusEvaluation schema-evaluation
evaluate$association() LinearRegression or LogisticRegression, through associate$regression() PolyGeniusEvaluation schema-evaluation
evaluate$stratification() direct lm/glm fit PolyGeniusEvaluation schema-evaluation
evaluate$compare() none PolyGeniusEvaluation schema-evaluation
evaluate$redundancy() none PolyGeniusEvaluation schema-evaluation
evaluate$profile() those of the components it runs PolyGeniusEvaluation schema-evaluation
evaluate$select() none, reads an evaluation and returns a data.table none none

The score layer must already exist on data, or the call aborts. It is used as supplied: nothing is scored or standardised.

ev <- evaluate$profile(
  data,
  outcomes   = outcome(case, prevalence = 0.05),
  covariates = c(age, sex, PC1)
)

# Rank the models, best first, per outcome.
ranking <- evaluate$select(ev)

Outcomes

An outcome is binary or continuous. The type is inferred from the values unless an outcome() descriptor sets it. A surv() outcome aborts. outcome(prevalence = ) gives a binary outcome's population prevalence and enables the liability-scale R-squared metrics of Lee et al. (2012), r2.liability and delta.r2.liability. A prevalence must be a single number in (0, 1). It aborts on a continuous outcome. One prevalence applies to every stratum. evaluate$redundancy() takes no outcome.

Complete cases

Complete cases are taken per model: the outcome, the covariates and that model's score. evaluate$compare() takes them per pair of models. An all-NA score column gives NA rows for that model only. Every row with an outcome reports its own n and n.events. evaluate$performance() and evaluate$compare() take no covariates, so one model's n can differ between the components of a profile.

Strata

split.by adds rows for each of its levels beside the overall stratum = "all" rows. A sample with a missing split.by value enters "all" only. Several split.by columns give one stratum per observed combination of their levels, labelled with the levels joined by ".". A stratum labelled "all" aborts, as does a split.by column that is not factor, character or logical. A covariate that is also a split.by column is dropped from the fits within levels, where it is constant.

Results columns

$results has these columns, in this order:

  • analysis -- the component that produced the row, e.g. "performance".
  • outcome -- the outcome label. NA on redundancy rows.
  • outcome.type -- "binary" or "continuous". NA on redundancy rows.
  • model -- the model label, a column name of the score layer.
  • model.ref -- the reference model on compare rows, the pair's second model on redundancy rows. NA elsewhere.
  • stratum -- "all" or a split.by level. Never NA.
  • tier -- the score percentile band on stratification rows, e.g. "0-20%". NA elsewhere.
  • tier.ref -- the reference band on stratification contrast rows. NA elsewhere.
  • metric -- the metric name. analysis and metric together identify a metric.
  • estimate -- the point estimate.
  • lower -- the lower interval bound at conf.level. NA when no interval is computed.
  • upper -- the upper interval bound at conf.level. NA when no interval is computed.
  • pval -- the raw p-value. NA except on the rows listed under Multiple testing.
  • adj.pval -- pval adjusted with p.adjust.method. NA where pval is.
  • adj.family.id -- the adjustment family, a key into $indices$multiplicity. NA where pval is.
  • n -- the samples behind the estimate. On a band row, the band's samples. NA on redundancy rows.
  • n.events -- the cases among them. NA for a continuous outcome and on redundancy rows.

schema-evaluation generates the column, metric and artifact tables.

Multiple testing

Only these rows carry a p-value:

  • incremental likelihood.ratio, the likelihood-ratio test of adding the score to the covariates;
  • association odds.ratio and beta, the Wald test from associate$regression();
  • stratification odds.ratio and mean.difference, the Wald test of each band against the reference band;
  • compare delta.auc, the paired DeLong test.

Every other row has NA pval, adj.pval and adj.family.id. Each component adjusts its own p-values once, with p.adjust.method, within analysis x outcome x metric x stratum. A stratification family therefore spans every band and model of one outcome and stratum. evaluate$profile() keeps each component's families and does not re-adjust. Neither does merge() or c() of two evaluations. With the same covariates, the association Wald test and the incremental likelihood-ratio test address the same null hypothesis. They are separate families, not independent evidence. The families are recorded in $indices$multiplicity.

Intervals

Bootstrap intervals are percentile intervals over bootstrap replicates. For a binary outcome, cases and controls are resampled separately. Every replicate refits every model it evaluates. A failed replicate is dropped and counted in $diagnostics. Resampling draws from the caller's random number stream, so calling set.seed() first makes it reproducible. bootstrap = 0 leaves the bootstrap intervals NA.

Analytic intervals do not depend on bootstrap: DeLong for auc and for compare's delta.auc, Wald for the association and stratification contrasts, Wilson score for outcome.rate, and t for outcome.mean. likelihood.ratio has no interval.

Diagnostics and metadata

$diagnostics$messages is one table with columns analysis, outcome, model, stratum and message. It records failed fits, strata where a binary outcome has a single class, degenerate bands, identical score columns, and failed bootstrap replicates as "k of B replicates failed".

$metadata records the arguments the component takes. evaluate$compare() adds reference.model, named by outcome, and reference.selected, TRUE when the reference was data-selected. evaluate$stratification() adds quantiles and reference.

Selecting models

evaluate$select() ranks the models of an evaluation on its performance and incremental metrics. Selecting and reporting on the same samples overstates the chosen model. Select on one split of the samples, then profile the chosen models on a disjoint split. Keep relatives in the same split, so the test split shares no relatives with the one that chose. split.by is not a substitute: its levels are not held out, and evaluate$select() reads only the "all" rows, which pool them.

See Also