PolyGenius
Contents

associate$regression

Regression-based associations for a PolyGeniusStudy

associate$regression() fits one association model per outcome x predictor x interaction x stratum cell resolved from a PolyGeniusStudy, picking a classical or survival family per outcome. It returns a PolyGeniusAssociation carrying long-format coefficient rows, plot-ready artifacts and per-fit diagnostics.

Usage

associate.regression(
  data,
  outcomes,
  predictors = everything(),
  scores.layer = X,
  split.by = NULL,
  interactions = NULL,
  covariates = NULL,
  model = c("auto", "lm", "glm", "cox", "crr", "km"),
  reference.level = NULL,
  weights = NULL,
  conf.level = 0.95,
  p.adjust.method = "BH",
  output = c("summary", "models"),
  artifacts = c("full", "minimal", "none"),
  grid.resolution = 50L,
  ...
)

Arguments

ArgumentDescription
dataA PolyGeniusStudy.
outcomesUnquoted outcome expression, resolved from data. Accepts a bare column, an [outcome()](/reference/outcome/) wrapper, a [surv()](/reference/surv/) descriptor, or a c() of these. Mixing surv() and plain outcomes in one call aborts.
predictorsUnquoted predictor expression, default everything(). everything() -- alone or as an element of a c() -- expands to every model column of scores.layer.
scores.layerUnquoted score-layer name, default X. Layer read when predictors resolves to PRS model names.
split.byUnquoted stratification expression, or NULL (default). One fit is run per observed level of the interaction of the resolved columns; each column must be factor, character or logical.
interactionsUnquoted interaction expression, or NULL (default). Each resolved column gets its own fit carrying a predictor x term product.
covariatesUnquoted adjustment expression, or NULL (default). Added to every fit.
modelOne of "auto" (default), "lm", "glm", "cox", "crr", "km". "auto" infers the family per outcome; any other value forces it.
reference.levelNon-empty character scalar, or NULL (default). Reference level for a categorical predictor; an error is raised if the predictor is not categorical or the level is unobserved.
weightsUnquoted expression resolving to one column, a numeric vector of length data$n.samples, or NULL (default). Values must be non-negative.
conf.levelNumeric scalar in (0, 1), default 0.95. Confidence level for lower/upper.
p.adjust.methodCharacter scalar, default "BH". Method passed to stats::p.adjust().
outputOne of "summary" (default), "models". "models" returns a fit index rather than a result object -- see Value.
artifactsOne of "full" (default), "minimal", "none". "full" keeps every artifact the family builds, "minimal" keeps only group.summary (so nothing survives for lm/glm), and "none" skips artifact construction entirely. Artifacts are recomputable from the data and the fit, so dropping them loses no inferential content.
grid.resolutionInteger scalar >= 2, default 50L. Number of points in the prediction.grid artifact built by lm and glm.
...Further arguments forwarded to the family's fitting function -- stats::lm(), stats::glm() or survival::coxph(). "km" accepts them and ignores them.

Value

With output = "summary" (default), a PolyGeniusAssociation: $results — One row per coefficient term, following the schema named in each row's schema column (schema-lm, schema-glm, schema-cox, schema-crr or schema-km): identity (.fit, fit.id, family, outcome, predictor, interaction, covariates, term, term.type, stratum, formula), inference (estimate, se, lower, upper, statistic, pval, adj.pval, adj.family.id), and n, effect.scale, test. Survival schemas add n.events, entry and time.origin; glm adds n.cases and n.controls. A km row is an omnibus log-rank test and carries no coefficient.

$artifacts — prediction.grid for lm/glm; observed.curves, predicted.curves, risk.table and group.summary for cox/crr; observed.curves, risk.table and group.summary for km. Compacted to a .fit key -- use artifacts() to rejoin the identity columns. Empty when artifacts = "none".

$diagnostics — fit.messages, one row per warning, error or not-fitted cell, kept wide with its full fit context.

$fits — One row per fit: identity plus covariates, formula, n, n.cases, n.controls, n.events, entry, time.origin. Failed fits are included.

$metadata — outcome.type (or "mixed" when families disagree).

$provenance — the captured call; see $metadata and $provenance in PolyGeniusResult.

$indices$multiplicity — One row per declared adjustment family, keyed from each result row's adj.family.id. With output = "models", a plain data.frame indexing the fits that ran, with columns fit.id, family, outcome and predictor. It does not carry the fitted model objects.

Details

Outcome descriptors

outcomes accepts a bare column, an outcome() wrapper, a surv() descriptor, or a c() of these. Survival descriptors are read from the unevaluated call, so only time, event, competing (or competing.event), entry and origin are honoured here; surv()'s event.code, competing.code, type and label are ignored on this path.

# Bare column -- family inferred from the column's type
associate$regression(data, outcomes = dementia)

# Several plain outcomes in one call
associate$regression(data, outcomes = c(bmi, dementia))

# Survival outcome (Cox), then with a competing risk (Fine-Gray)
associate$regression(data, outcomes = surv(time = age, event = dementia))
associate$regression(data,
  outcomes = surv(time = age, event = dementia, competing = death))

# Several survival outcomes sharing one time variable
associate$regression(data,
  outcomes = c(
    dementia = surv(time = age, event = dementia),
    mci      = surv(time = age, event = mci)
  ))

# Left-truncated (delayed-entry) Cox
associate$regression(data,
  outcomes = surv(time = age_exit, event = dementia, entry = age_entry,
                  origin = "age since birth"))

Family selection

model = "auto" resolves per outcome: a surv() descriptor with a competing event gives "crr" and without one gives "cox"; a factor or logical column gives "glm", as does a numeric column whose non-missing values are all in {0, 1} and a character column with at most two distinct values; anything else gives "lm". "km" is never inferred and must be requested explicitly. A multi-level factor outcome resolves to "glm" and then fails in the fit.

Family Outcome Effect scale
"lm" continuous identity
"glm" binary log-odds
"cox" right-censored survival log-hazard
"crr" competing-risk survival log-subdistribution-hazard
"km" grouped survival, omnibus log-rank test only none (no coefficient)

Left truncation

surv(entry = ) declares a delayed-entry time: "cox" and "crr" then fit Surv(entry, time, event), restricting each event's risk set to subjects with entry <= time. surv(origin = ) is a free-text label for what time = 0 means; it is never checked against the data, but it is recorded on every result row and in $fits as time.origin (alongside the resolved entry column name) so associate$meta can warn when pooled studies disagree.

Complete cases

Each fit drops rows missing any variable it uses -- outcome, time, event, competing, entry, predictor, interaction or covariate -- and reports that count as n. Dropping is per fit, so strata and families can retain different sample sizes. A subject lost to follow-up must be coded as censored (last contact in time, event = 0), not as NA: censored rows define the at-risk sets and, for "crr", the censoring distribution behind the Fine-Gray weights, whereas an NA time silently removes the subject. For "crr", complete cases are selected before survival::finegray() builds the weighted risk set, so those weights are estimated on exactly the analysis sample.

Multiplicity

p.adjust.method is applied within each outcome x family x stratum group, not across the whole table, so screening many predictors for one outcome is the controlled family. km rows carry no adj.pval. The declared families are recorded in $indices$multiplicity and keyed from adj.family.id.

Failed fits

A cell that errors is dropped, recorded in $diagnostics$fit.messages with a reason.code, and summarized in one warning; the surviving cells are still returned. The call aborts only when every cell failed, or -- before any fit runs -- when a covariate or interaction term has a single observed level across the whole frame.

Examples

# Binary endpoint association
associate$regression(data, outcomes = dementia, predictors = PRS_AD,
                     covariates = c(age, sex, PC1, PC2))

# Cox regression over every model in the default score layer
associate$regression(data,
  outcomes = surv(time = age_obs, event = dementia),
  predictors = everything(),
  covariates = c(sex, PC1, PC2))

# Competing-risk regression, artifacts dropped for a large screen
associate$regression(data,
  outcomes = surv(time = age_obs, event = dementia, competing = death),
  predictors = PRS_AD,
  artifacts = "none")

See Also

Aliases: associate-regression, associate.regression, associate$regression