PolyGenius
Contents

evaluate$incremental

Incremental value of a score beyond covariates

evaluate$incremental() asks what each score adds to a model on the covariates alone. It fits a baseline model and the baseline plus the score, and reports the difference.

Usage

evaluate$incremental(
  data,
  outcomes,
  covariates,
  scores.layer = X,
  split.by = NULL,
  bootstrap = 1000,
  conf.level = 0.95,
  p.adjust.method = "BH"
)

Arguments

ArgumentDescription
dataA [PolyGeniusStudy](/reference/polygeniusstudy/). Holds the score layer and the sample columns the other arguments name. It is not modified.
outcomesUnquoted outcome expression resolved from data: a bare column, an expression such as status == "case", an [outcome()](/reference/outcome/) descriptor, or a c() or list() of these whose names label the outcomes. The type is binary or continuous, inferred from the values unless the descriptor sets it. A [surv()](/reference/surv/) outcome aborts.
covariatesUnquoted covariate expression resolved from data, such as c(age, sex, PC1). Required, because it defines the baseline model: omitting it or passing NULL aborts. A character vector aborts.
scores.layerUnquoted name of an existing score layer, default X. The layer is used as supplied: nothing is scored or standardised here.
split.byUnquoted expression naming factor, character or logical sample columns, or NULL (default). Adds rows for each level beside the overall stratum = "all" rows. A level named "all" aborts.
bootstrapNon-negative integer, default 1000. Replicates behind the percentile intervals. 0 returns estimates without bootstrap intervals. Resampling draws from the caller's random number stream.
conf.levelNumeric scalar in (0, 1), default 0.95.
p.adjust.methodOne of stats::p.adjust.methods, default "BH". Applied within analysis, outcome, metric and stratum.

Value

A PolyGeniusEvaluation following the schema-evaluation schema. $results has the columns listed in evaluate, one row per outcome, model, stratum and metric, with analysis = "incremental". Binary outcomes give delta.auc, likelihood.ratio and, when prevalence is set, delta.r2.liability. Continuous outcomes give delta.r.squared and likelihood.ratio. Only likelihood.ratio carries pval and adj.pval, and its lower and upper are NA. $indices$multiplicity records the adjustment families. There are no artifacts.

$diagnostics$messages is always present, recording a stratum with no complete cases or a single class, a dropped covariate, an aliased score, each fit warning or error, a non-finite estimate, and the number of failed bootstrap replicates.

$metadata holds the outcome labels, the covariate columns, scores.layer, the split.by columns, bootstrap, conf.level and p.adjust.method.

Details

Per outcome, model and stratum, two nested models are fitted on the same complete-case rows: the baseline outcome ~ covariates and the full outcome ~ covariates + score. A binary outcome uses stats::glm() with the binomial family. A continuous outcome uses stats::lm(). The baseline is refitted for every model, on that model's complete cases. Within a split.by level, a covariate that is also a split.by column is dropped from both models. A factor or character covariate with fewer than two observed levels in a cell is dropped there, with a diagnostic. When no covariate is left, the baseline is the intercept alone, and delta.auc is the full model's AUC minus 0.5. The shared rules on outcomes, complete cases and strata are in evaluate.

The metrics:

  • delta.auc (binary): the AUC of the full model's fitted values minus the AUC of the baseline's.
  • delta.r2.liability (binary): the liability-scale R^2 of the full model minus that of the baseline. Emitted only when the outcome() descriptor sets prevalence. Each model's observed-scale R^2 is that of stats::lm() of the 0/1 outcome on the same terms, converted with Lee et al. (2012, Genet Epidemiol 36:214), eq. 15. K is the prevalence and P is the case fraction of the fitted rows. Applying eq. 15 to each of two nested models is an extension of Lee et al.
  • delta.r.squared (continuous): the R^2 of the full model minus that of the baseline.
  • likelihood.ratio (both types): twice the log-likelihood difference, 2 * (logLik(full) - logLik(baseline)), from stats::logLik(). Its p-value is the chi-squared test on the difference in degrees of freedom. For a continuous outcome this is the asymptotic test, not the nested F test. A score aliased with the covariates, so that the full model adds no degree of freedom, gives NA rows and a diagnostic.

likelihood.ratio is the only row with a p-value. A test of the AUC difference between nested models is invalid (Pepe et al. 2013, Stat Med 32:1467), so the likelihood-ratio test is the test of added value.

The three delta metrics carry percentile bootstrap intervals, and every replicate refits both models. likelihood.ratio has no interval. delta.auc is in-sample (apparent): it is measured on the rows the models were fitted on. Its interval's coverage is unreliable when the true difference is near 0. delta.r.squared and delta.r2.liability are non-negative by construction. Their intervals describe precision and are not tests.

Examples

inc <- evaluate$incremental(
  study,
  outcomes   = c(ad = status == "case"),
  covariates = c(age, sex, PC1, PC2)
)
inc$results[metric == "likelihood.ratio"]

# Liability-scale increment, within each level of a grouping column.
evaluate$incremental(
  study,
  outcomes   = outcome(status, type = "binary", event = "case",
                       prevalence = 0.05),
  covariates = c(age, PC1, PC2),
  split.by   = sex
)

See Also

evaluate$profile() runs this with the other components when covariates is given. visualize$evaluate$incremental() plots the delta rows.

Other evaluate-components: evaluate, evaluate.association(), evaluate.compare(), evaluate.performance(), evaluate.profile(), evaluate.redundancy(), evaluate.select(), evaluate.stratification()

References

Lee et al. (2012). A better coefficient of determination for genetic profile analysis. Genetic Epidemiology 36(3), 214-224. doi:10.1002/gepi.21614

Pepe et al. (2013). Testing for improvement in prediction model performance. Statistics in Medicine 32(9), 1467-1482. doi:10.1002/sim.5727

Aliases: evaluate.incremental, evaluate$incremental