Contents
evaluate$performance
Predictive performance of each score
evaluate$performance() asks how well each score, on its own, predicts
each outcome: how it separates cases from controls, or how much outcome
variance it explains.
Usage
evaluate$performance(
data,
outcomes,
scores.layer = X,
split.by = NULL,
bootstrap = 1000,
conf.level = 0.95
)Arguments
| Argument | Description |
|---|---|
data | A [PolyGeniusStudy](/reference/polygeniusstudy/). Holds the score layer and the sample columns the other arguments name. It is not modified. |
outcomes | Unquoted outcome expression resolved from data: a bare column, an expression such as status == "case", an [outcome()](/reference/outcome/) descriptor, or a c() or list() of these whose names label the outcomes. The type is binary or continuous, inferred from the values unless the descriptor sets it. A [surv()](/reference/surv/) outcome aborts. |
scores.layer | Unquoted name of an existing score layer, default X. The layer is used as supplied: nothing is scored or standardised here. |
split.by | Unquoted expression naming factor, character or logical sample columns, or NULL (default). Adds rows for each level beside the overall stratum = "all" rows. A level named "all" aborts. |
bootstrap | Non-negative integer, default 1000. Replicates behind the percentile intervals of r2.liability and r.squared. 0 returns estimates without bootstrap intervals. Resampling draws from the caller's random number stream. The DeLong interval of auc does not use it. |
conf.level | Numeric scalar in (0, 1), default 0.95. |
Value
A PolyGeniusEvaluation following the schema-evaluation schema. $results has the columns listed in
evaluate, one row per outcome, model, stratum and metric, with
analysis = "performance". Binary outcomes give auc, and r2.liability
when a prevalence is set. Continuous outcomes give r.squared. pval,
adj.pval and adj.family.id are NA on every row. Artifacts, read with
artifacts():
confusion(binary): one row per outcome, model, stratum and threshold, with columnsoutcome, model, stratum, threshold, tp, fp, tn, fn. The thresholds are the distinct quantiles of that outcome, model and stratum's scores at 100 evenly spaced probabilities from 0 to 1, so there are at most 100. A sample counts as positive when its score is at or above the threshold. A stratum with a single class, or with no complete cases, has no rows.score.outcome.bins(continuous): a count grid of 50 equal-width score bins by 50 equal-width outcome bins, spanning each axis's range within that outcome, model and stratum, one row per non-empty grid cell, with columnsoutcome, model, stratum, score.lo, score.hi, outcome.lo, outcome.hi, n.
Neither artifact is federation-safe: both can disclose individual-level information in small strata.
$diagnostics$messages is always present, recording a stratum with a
single class or no complete cases, a non-finite continuous r.squared, a
non-finite DeLong estimate or bound (set to NA), each warning from pROC,
and the number of failed bootstrap replicates. pROC's warnings go there
instead of to the console. There is no $indices$multiplicity.
$metadata holds the outcome labels, scores.layer, the split.by
columns, bootstrap and conf.level.
Details
The score is the only predictor. Covariates are not used. The shared rules on outcomes, complete cases, strata and intervals are in evaluate.
A binary outcome gives auc, the area under the ROC curve. The direction is
fixed: a higher score predicts a case. A score negatively associated with
the outcome therefore gives an AUC below 0.5. Its interval is the DeLong
interval. It is analytic, so it is reported even with bootstrap = 0.
A binary outcome also gives r2.liability when its outcome() descriptor
sets prevalence. It is the variance the score explains on the liability
scale (Lee et al. 2012, eq. 15). R^2_o is
summary(lm(y ~ score))$r.squared of the 0/1 outcome: the unadjusted
R-squared, not a pseudo-R-squared. P is the case fraction in the
complete-case sample of that model and stratum. K is the prevalence,
the same in every stratum. The interval is bootstrap.
A continuous outcome gives r.squared, the unadjusted R-squared of
lm(y ~ score). The interval is bootstrap.
No row carries a p-value. evaluate$association() tests the link between
score and outcome.
A stratum in which a binary outcome has only cases or only controls gives NA rows and a diagnostic.
Examples
perf <- evaluate$performance(study, outcomes = c(case = status == "case"))
artifacts(perf, "confusion")
# Liability-scale R-squared needs the population prevalence.
evaluate$performance(study,
outcomes = outcome(status, event = "case", prevalence = 0.05),
split.by = sex)
# A continuous outcome, estimates only.
evaluate$performance(study, outcomes = bmi, bootstrap = 0)See Also
evaluate$profile() runs this together with the other components.
evaluate$compare() contrasts these metrics between models.
visualize$evaluate$performance() plots the artifacts.
Other evaluate-components:
evaluate,
evaluate.association(),
evaluate.compare(),
evaluate.incremental(),
evaluate.profile(),
evaluate.redundancy(),
evaluate.select(),
evaluate.stratification()
References
Lee et al. (2012). A better coefficient of determination for genetic profile analysis. Genetic Epidemiology 36(3), 214-224. doi:10.1002/gepi.21614