Contents
associate$regression
Regression-based associations for a PolyGeniusStudy
associate$regression() fits one association model per outcome x predictor x
interaction x stratum cell resolved from a PolyGeniusStudy, picking a
classical or survival family per outcome. It returns a
PolyGeniusAssociation carrying long-format coefficient rows, plot-ready
artifacts and per-fit diagnostics.
Usage
associate.regression(
data,
outcomes,
predictors = everything(),
scores.layer = X,
split.by = NULL,
interactions = NULL,
covariates = NULL,
model = c("auto", "lm", "glm", "cox", "crr", "km"),
reference.level = NULL,
weights = NULL,
conf.level = 0.95,
p.adjust.method = "BH",
output = c("summary", "models"),
artifacts = c("full", "minimal", "none"),
grid.resolution = 50L,
...
)Arguments
| Argument | Description |
|---|---|
data | A PolyGeniusStudy. |
outcomes | Unquoted outcome expression, resolved from data. Accepts a bare column, an [outcome()](/reference/outcome/) wrapper, a [surv()](/reference/surv/) descriptor, or a c() of these. Mixing surv() and plain outcomes in one call aborts. |
predictors | Unquoted predictor expression, default everything(). everything() -- alone or as an element of a c() -- expands to every model column of scores.layer. |
scores.layer | Unquoted score-layer name, default X. Layer read when predictors resolves to PRS model names. |
split.by | Unquoted stratification expression, or NULL (default). One fit is run per observed level of the interaction of the resolved columns; each column must be factor, character or logical. |
interactions | Unquoted interaction expression, or NULL (default). Each resolved column gets its own fit carrying a predictor x term product. |
covariates | Unquoted adjustment expression, or NULL (default). Added to every fit. |
model | One of "auto" (default), "lm", "glm", "cox", "crr", "km". "auto" infers the family per outcome; any other value forces it. |
reference.level | Non-empty character scalar, or NULL (default). Reference level for a categorical predictor; an error is raised if the predictor is not categorical or the level is unobserved. |
weights | Unquoted expression resolving to one column, a numeric vector of length data$n.samples, or NULL (default). Values must be non-negative. |
conf.level | Numeric scalar in (0, 1), default 0.95. Confidence level for lower/upper. |
p.adjust.method | Character scalar, default "BH". Method passed to stats::p.adjust(). |
output | One of "summary" (default), "models". "models" returns a fit index rather than a result object -- see Value. |
artifacts | One of "full" (default), "minimal", "none". "full" keeps every artifact the family builds, "minimal" keeps only group.summary (so nothing survives for lm/glm), and "none" skips artifact construction entirely. Artifacts are recomputable from the data and the fit, so dropping them loses no inferential content. |
grid.resolution | Integer scalar >= 2, default 50L. Number of points in the prediction.grid artifact built by lm and glm. |
... | Further arguments forwarded to the family's fitting function -- stats::lm(), stats::glm() or survival::coxph(). "km" accepts them and ignores them. |
Value
With output = "summary" (default), a PolyGeniusAssociation:
$results — One row per coefficient term, following the schema named
in each row's schema column (schema-lm, schema-glm, schema-cox, schema-crr or
schema-km): identity (.fit, fit.id, family, outcome, predictor,
interaction, covariates, term, term.type, stratum, formula),
inference (estimate, se, lower, upper, statistic, pval,
adj.pval, adj.family.id), and n, effect.scale, test. Survival
schemas add n.events, entry and time.origin; glm adds
n.cases and n.controls.
A km row is an omnibus log-rank test and carries no coefficient.
$artifacts — prediction.grid for lm/glm; observed.curves,
predicted.curves, risk.table and group.summary for cox/crr;
observed.curves, risk.table and group.summary for km. Compacted to
a .fit key -- use artifacts() to rejoin the identity columns. Empty
when artifacts = "none".
$diagnostics — fit.messages, one row per warning, error or
not-fitted cell, kept wide with its full fit context.
$fits — One row per fit: identity plus covariates, formula, n,
n.cases, n.controls, n.events, entry, time.origin. Failed fits
are included.
$metadata — outcome.type (or "mixed" when families disagree).
$provenance — the captured call; see $metadata and $provenance in
PolyGeniusResult.
$indices$multiplicity — One row per declared adjustment family,
keyed from each result row's adj.family.id.
With output = "models", a plain data.frame indexing the fits that ran, with
columns fit.id, family, outcome and predictor. It does not carry the
fitted model objects.
Details
Outcome descriptors
outcomes accepts a bare column, an outcome() wrapper, a surv()
descriptor, or a c() of these. Survival descriptors are read from the
unevaluated call, so only time, event, competing (or competing.event),
entry and origin are honoured here; surv()'s event.code,
competing.code, type and label are ignored on this path.
# Bare column -- family inferred from the column's type
associate$regression(data, outcomes = dementia)
# Several plain outcomes in one call
associate$regression(data, outcomes = c(bmi, dementia))
# Survival outcome (Cox), then with a competing risk (Fine-Gray)
associate$regression(data, outcomes = surv(time = age, event = dementia))
associate$regression(data,
outcomes = surv(time = age, event = dementia, competing = death))
# Several survival outcomes sharing one time variable
associate$regression(data,
outcomes = c(
dementia = surv(time = age, event = dementia),
mci = surv(time = age, event = mci)
))
# Left-truncated (delayed-entry) Cox
associate$regression(data,
outcomes = surv(time = age_exit, event = dementia, entry = age_entry,
origin = "age since birth"))
Family selection
model = "auto" resolves per outcome: a surv() descriptor with a competing
event gives "crr" and without one gives "cox"; a factor or logical column
gives "glm", as does a numeric column whose non-missing values are all in
{0, 1} and a character column with at most two distinct values; anything
else gives "lm". "km" is never inferred and must be requested explicitly.
A multi-level factor outcome resolves to "glm" and then fails in the fit.
| Family | Outcome | Effect scale |
|---|---|---|
"lm" |
continuous | identity |
"glm" |
binary | log-odds |
"cox" |
right-censored survival | log-hazard |
"crr" |
competing-risk survival | log-subdistribution-hazard |
"km" |
grouped survival, omnibus log-rank test only | none (no coefficient) |
Left truncation
surv(entry = ) declares a delayed-entry time: "cox" and "crr" then fit
Surv(entry, time, event), restricting each event's risk set to subjects with
entry <= time. surv(origin = ) is a free-text label for what time = 0
means; it is never checked against the data, but it is recorded on every
result row and in $fits as time.origin (alongside the resolved entry
column name) so associate$meta can warn when pooled studies
disagree.
Complete cases
Each fit drops rows missing any variable it uses -- outcome, time, event,
competing, entry, predictor, interaction or covariate -- and reports that
count as n. Dropping is per fit, so strata and families can retain different
sample sizes. A subject lost to follow-up must be coded as censored (last
contact in time, event = 0), not as NA: censored rows define the at-risk
sets and, for "crr", the censoring distribution behind the Fine-Gray
weights, whereas an NA time silently removes the subject. For "crr",
complete cases are selected before survival::finegray() builds the weighted
risk set, so those weights are estimated on exactly the analysis sample.
Multiplicity
p.adjust.method is applied within each outcome x family x stratum group, not
across the whole table, so screening many predictors for one outcome is the
controlled family. km rows carry no adj.pval. The declared families are
recorded in $indices$multiplicity and keyed from adj.family.id.
Failed fits
A cell that errors is dropped, recorded in $diagnostics$fit.messages with a
reason.code, and summarized in one warning; the surviving cells are still
returned. The call aborts only when every cell failed, or -- before any fit
runs -- when a covariate or interaction term has a single observed level
across the whole frame.
Examples
# Binary endpoint association
associate$regression(data, outcomes = dementia, predictors = PRS_AD,
covariates = c(age, sex, PC1, PC2))
# Cox regression over every model in the default score layer
associate$regression(data,
outcomes = surv(time = age_obs, event = dementia),
predictors = everything(),
covariates = c(sex, PC1, PC2))
# Competing-risk regression, artifacts dropped for a large screen
associate$regression(data,
outcomes = surv(time = age_obs, event = dementia, competing = death),
predictors = PRS_AD,
artifacts = "none")See Also
associate, associate$compare,
associate$meta, PolyGeniusAssociation, outcome(),
surv()
Other associate-analyses:
associate-compare,
associate-mediation,
associate-meta,
associate-mr,
associate-single-variant