Contents
KaplanMeierRegression
Kaplan-Meier log-rank family for associate$regression()
Fits grouped Kaplan-Meier curves with survival::survfit() and tests them
with survival::survdiff(). associate$regression
instantiates this family only for an explicit model = "km"; "auto" never
infers it.
Details
The predictor must be a factor, character or logical column and becomes the
grouping variable; a numeric predictor aborts, so groups must be created
before the call. An interaction column, itself categorical, is crossed with
the predictor via base::interaction() so each group level is one predictor
x interaction combination. Covariates and case weights play no part in the
fit. Surv(time, event) ~ group is then fitted on complete cases; a
surv(entry = ) column is not read by this family.
The result table carries exactly one row per fit, following the "km"
association schema (schemaKm() in associate-schemas.R): the omnibus
log-rank test (test = "logrank", term.type = "omnibus",
effect.scale = "median.time"), whose statistic is the chi-squared value
and pval its upper tail on max(1, groups - 1) degrees of freedom. There
is no coefficient, so estimate, se, lower and upper are all
NA_real_. The schema declares no meta-analysis keys, so km rows are not
poolable by associate$meta.
Group-level detail lives in the artifacts: observed.curves (per-group
curve steps with confidence limits and at-risk, event and censoring counts),
risk.table (the same rows without the estimates) and group.summary (per
group n, n.events and median survival with its limits).
Super class
PolyGenius::RegressionFamily -> KaplanMeierRegression
Methods
Public methods
KaplanMeierRegression$new()KaplanMeierRegression$validate.frame()KaplanMeierRegression$prepare.frame()KaplanMeierRegression$describe.model.terms()KaplanMeierRegression$build.formula()KaplanMeierRegression$fit.model()KaplanMeierRegression$build.summary.rows()KaplanMeierRegression$build.artifacts()KaplanMeierRegression$clone()
Method new()
Register this family under the key "km".
Usage
KaplanMeierRegression$new()
Returns
A new KaplanMeierRegression.
Method validate.frame()
Check the time and predictor roles resolved onto the frame before any model is built.
Usage
KaplanMeierRegression$validate.frame(frame, cell, fit.specs, conf.level)
Arguments
frame — A data.table resolved from a PolyGeniusStudy, whose
roles() tags identify the time, outcome, predictor and interaction
columns.
cell — Named list describing one fit cell: fit.id, analysis.id,
family, outcome, predictor, interaction, stratum. Unused.
fit.specs — Named list of extra fitting options. Unused.
conf.level — Numeric scalar in (0, 1). Unused.
Returns
TRUE, invisibly. Aborts unless exactly one column carries the
time role and the predictor is a factor, character or logical
column.
Method prepare.frame()
Shape the data for survival::survfit() and
survival::survdiff(): time, event (0/1) and group, the
predictor as a factor (releveled when reference.level was given),
crossed with the interaction column when one is present. Rows missing
any of them are dropped.
Usage
KaplanMeierRegression$prepare.frame(frame, cell, fit.specs, conf.level)
Arguments
frame — A data.table resolved from a PolyGeniusStudy, whose
roles() tags identify the time, outcome, predictor and interaction
columns.
cell — Named list describing one fit cell. Unused; columns come
from the frame's roles.
fit.specs — Named list of extra fitting options. reference.level
is read here.
conf.level — Numeric scalar in (0, 1). Stored on the returned list
and used later as survfit(conf.int = ).
Returns
Named list: data (complete-case data.frame with time,
event and group), time.name, event.name, predictor.name,
predictor.label, interaction.name, conf.level and row.id, the
retained row indices into frame. Aborts when the interaction column
is not factor, character or logical.
Method describe.model.terms()
Describe the single term whose estimability is checked. Kaplan-Meier compares survival across the grouping predictor only, with no covariates, so the fit hinges on the group having at least two observed levels.
Usage
KaplanMeierRegression$describe.model.terms(fit.data)
Arguments
fit.data — Named list returned by prepare.frame().
Returns
Single-element list holding one list(column, role, label)
descriptor with role = "group".
Method build.formula()
Build Surv(time, event) ~ group over the combined group
column from prepare.frame().
Usage
KaplanMeierRegression$build.formula(frame)
Arguments
frame — Named list returned by prepare.frame().
Returns
A formula over the internal column names, carrying the
original-column-name form as the "display.formula" attribute.
Method fit.model()
Fit both survival::survfit(), for the curves, and
survival::survdiff(), for the log-rank statistic, from the same
formula.
Usage
KaplanMeierRegression$fit.model(formula, frame, ...)
Arguments
formula — A formula from build.formula().
frame — Named list returned by prepare.frame(); $conf.level
becomes survfit(conf.int = ).
... — Accepted and ignored; neither underlying call receives them.
Returns
Named list with elements survfit (a survfit object) and
survdiff (a survdiff object).
Method build.summary.rows()
Build the single omnibus log-rank result row -- chi-squared statistic, degrees of freedom and p-value. Per-group estimates are in the artifacts instead.
Usage
KaplanMeierRegression$build.summary.rows(
fit,
fit.data,
cell,
fit.specs,
conf.level,
formula
)
Arguments
fit — Named list with survfit and survdiff, from fit.model().
fit.data — Named list returned by prepare.frame().
cell — Named list describing one fit cell: fit.id, analysis.id,
family, outcome, predictor, stratum.
fit.specs — Named list of extra fitting options. Unused.
conf.level — Numeric scalar in (0, 1). Unused; the row carries no
interval.
formula — Character scalar. Display formula recorded on the row.
Returns
A single-row data.table of "km"-schema columns with
term.type = "omnibus" and NA_real_ for estimate, se, lower
and upper.
Method build.artifacts()
Build the per-group curve steps, risk table and group-level summary statistics.
Usage
KaplanMeierRegression$build.artifacts(
fit,
fit.data,
cell,
fit.specs,
conf.level,
formula
)
Arguments
fit — Named list with survfit and survdiff, from fit.model().
fit.data — Named list returned by prepare.frame().
cell — Named list describing one fit cell.
fit.specs — Named list of extra fitting options. Unused.
conf.level — Numeric scalar in (0, 1). Unused; the curve limits
come from the conf.int already applied in fit.model().
formula — Character scalar. Unused; present for interface parity.
Returns
Named list of data.frames: observed.curves (one row per group
x time step, with estimate, lower, upper, n.risk, n.event,
n.censor), risk.table (the same rows without the estimates) and
group.summary (per group n, n.events, median.time and its
limits).
Method clone()
The objects of this class are cloneable with this method.
Usage
KaplanMeierRegression$clone(deep = FALSE)
Arguments
deep — Whether to make a deep clone.
Examples
# Log-rank comparison across PRS tertiles -- "km" must be asked for
associate$regression(data,
outcomes = surv(time = age_obs, event = dementia),
predictors = PRS_AD_tertile,
model = "km")
# Same test with an explicit reference group, run within each sex
associate$regression(data,
outcomes = surv(time = age_obs, event = dementia),
predictors = PRS_AD_tertile,
model = "km", reference.level = "low", split.by = sex)See Also
associate$regression, CoxRegression for the adjusted hazard-ratio fit, CompetingRiskRegression when a competing event is declared.
Other regression-families:
CompetingRiskRegression,
CoxRegression,
LinearRegression,
LogisticRegression,
RegressionFamily