PolyGenius
Contents

KaplanMeierRegression

Kaplan-Meier log-rank family for associate$regression()

Fits grouped Kaplan-Meier curves with survival::survfit() and tests them with survival::survdiff(). associate$regression instantiates this family only for an explicit model = "km"; "auto" never infers it.

Details

The predictor must be a factor, character or logical column and becomes the grouping variable; a numeric predictor aborts, so groups must be created before the call. An interaction column, itself categorical, is crossed with the predictor via base::interaction() so each group level is one predictor x interaction combination. Covariates and case weights play no part in the fit. Surv(time, event) ~ group is then fitted on complete cases; a surv(entry = ) column is not read by this family.

The result table carries exactly one row per fit, following the "km" association schema (schemaKm() in associate-schemas.R): the omnibus log-rank test (test = "logrank", term.type = "omnibus", effect.scale = "median.time"), whose statistic is the chi-squared value and pval its upper tail on max(1, groups - 1) degrees of freedom. There is no coefficient, so estimate, se, lower and upper are all NA_real_. The schema declares no meta-analysis keys, so km rows are not poolable by associate$meta.

Group-level detail lives in the artifacts: observed.curves (per-group curve steps with confidence limits and at-risk, event and censoring counts), risk.table (the same rows without the estimates) and group.summary (per group n, n.events and median survival with its limits).

Super class

PolyGenius::RegressionFamily -> KaplanMeierRegression

Methods

Public methods

  • KaplanMeierRegression$new()
  • KaplanMeierRegression$validate.frame()
  • KaplanMeierRegression$prepare.frame()
  • KaplanMeierRegression$describe.model.terms()
  • KaplanMeierRegression$build.formula()
  • KaplanMeierRegression$fit.model()
  • KaplanMeierRegression$build.summary.rows()
  • KaplanMeierRegression$build.artifacts()
  • KaplanMeierRegression$clone()

Method new()

Register this family under the key "km".

Usage

KaplanMeierRegression$new()

Returns

A new KaplanMeierRegression.

Method validate.frame()

Check the time and predictor roles resolved onto the frame before any model is built.

Usage

KaplanMeierRegression$validate.frame(frame, cell, fit.specs, conf.level)

Arguments

frame — A data.table resolved from a PolyGeniusStudy, whose roles() tags identify the time, outcome, predictor and interaction columns.

cell — Named list describing one fit cell: fit.id, analysis.id, family, outcome, predictor, interaction, stratum. Unused.

fit.specs — Named list of extra fitting options. Unused.

conf.level — Numeric scalar in (0, 1). Unused.

Returns

TRUE, invisibly. Aborts unless exactly one column carries the time role and the predictor is a factor, character or logical column.

Method prepare.frame()

Shape the data for survival::survfit() and survival::survdiff(): time, event (0/1) and group, the predictor as a factor (releveled when reference.level was given), crossed with the interaction column when one is present. Rows missing any of them are dropped.

Usage

KaplanMeierRegression$prepare.frame(frame, cell, fit.specs, conf.level)

Arguments

frame — A data.table resolved from a PolyGeniusStudy, whose roles() tags identify the time, outcome, predictor and interaction columns.

cell — Named list describing one fit cell. Unused; columns come from the frame's roles.

fit.specs — Named list of extra fitting options. reference.level is read here.

conf.level — Numeric scalar in (0, 1). Stored on the returned list and used later as survfit(conf.int = ).

Returns

Named list: data (complete-case data.frame with time, event and group), time.name, event.name, predictor.name, predictor.label, interaction.name, conf.level and row.id, the retained row indices into frame. Aborts when the interaction column is not factor, character or logical.

Method describe.model.terms()

Describe the single term whose estimability is checked. Kaplan-Meier compares survival across the grouping predictor only, with no covariates, so the fit hinges on the group having at least two observed levels.

Usage

KaplanMeierRegression$describe.model.terms(fit.data)

Arguments

fit.data — Named list returned by prepare.frame().

Returns

Single-element list holding one list(column, role, label) descriptor with role = "group".

Method build.formula()

Build Surv(time, event) ~ group over the combined group column from prepare.frame().

Usage

KaplanMeierRegression$build.formula(frame)

Arguments

frame — Named list returned by prepare.frame().

Returns

A formula over the internal column names, carrying the original-column-name form as the "display.formula" attribute.

Method fit.model()

Fit both survival::survfit(), for the curves, and survival::survdiff(), for the log-rank statistic, from the same formula.

Usage

KaplanMeierRegression$fit.model(formula, frame, ...)

Arguments

formula — A formula from build.formula().

frame — Named list returned by prepare.frame(); $conf.level becomes survfit(conf.int = ).

... — Accepted and ignored; neither underlying call receives them.

Returns

Named list with elements survfit (a survfit object) and survdiff (a survdiff object).

Method build.summary.rows()

Build the single omnibus log-rank result row -- chi-squared statistic, degrees of freedom and p-value. Per-group estimates are in the artifacts instead.

Usage

KaplanMeierRegression$build.summary.rows(
  fit,
  fit.data,
  cell,
  fit.specs,
  conf.level,
  formula
)

Arguments

fit — Named list with survfit and survdiff, from fit.model().

fit.data — Named list returned by prepare.frame().

cell — Named list describing one fit cell: fit.id, analysis.id, family, outcome, predictor, stratum.

fit.specs — Named list of extra fitting options. Unused.

conf.level — Numeric scalar in (0, 1). Unused; the row carries no interval.

formula — Character scalar. Display formula recorded on the row.

Returns

A single-row data.table of "km"-schema columns with term.type = "omnibus" and NA_real_ for estimate, se, lower and upper.

Method build.artifacts()

Build the per-group curve steps, risk table and group-level summary statistics.

Usage

KaplanMeierRegression$build.artifacts(
  fit,
  fit.data,
  cell,
  fit.specs,
  conf.level,
  formula
)

Arguments

fit — Named list with survfit and survdiff, from fit.model().

fit.data — Named list returned by prepare.frame().

cell — Named list describing one fit cell.

fit.specs — Named list of extra fitting options. Unused.

conf.level — Numeric scalar in (0, 1). Unused; the curve limits come from the conf.int already applied in fit.model().

formula — Character scalar. Unused; present for interface parity.

Returns

Named list of data.frames: observed.curves (one row per group x time step, with estimate, lower, upper, n.risk, n.event, n.censor), risk.table (the same rows without the estimates) and group.summary (per group n, n.events, median.time and its limits).

Method clone()

The objects of this class are cloneable with this method.

Usage

KaplanMeierRegression$clone(deep = FALSE)

Arguments

deep — Whether to make a deep clone.

Examples

# Log-rank comparison across PRS tertiles -- "km" must be asked for
associate$regression(data,
  outcomes = surv(time = age_obs, event = dementia),
  predictors = PRS_AD_tertile,
  model = "km")

# Same test with an explicit reference group, run within each sex
associate$regression(data,
  outcomes = surv(time = age_obs, event = dementia),
  predictors = PRS_AD_tertile,
  model = "km", reference.level = "low", split.by = sex)

See Also

associate$regression, CoxRegression for the adjusted hazard-ratio fit, CompetingRiskRegression when a competing event is declared.

Other regression-families: CompetingRiskRegression, CoxRegression, LinearRegression, LogisticRegression, RegressionFamily