Regression Associations
One call, many outcomes: continuous and binary models, subgroups, and comparisons
One call
associate$regression() fits a model for every combination of outcome, predictor, term and stratum you ask for, and returns one table of coefficients.
The result is a grid: one row per outcome × predictor × term × stratum. Everything else on this page is a variation on that grid.
Every row carries a term and a term.type, so the score's coefficient, the covariate coefficients and any interaction terms all sit in one table. Filter on term.type rather than assuming the row order.
The model family follows each outcome's column type unless you name one. Survival outcomes use the same call but have their own page — 07.2.
Multiplicity here: adjusted p-values are computed within outcome × family × stratum.
Continuous and binary outcomes
A continuous outcome gives you a linear model and an effect on the outcome's own scale.
A binary outcome gives you logistic regression and an effect in log-odds, with case and control counts on the row. Plots exponentiate for display; the stored estimate does not.
Set the reference level explicitly for a binary outcome. An unset factor takes R's level order, which silently flips the sign of every effect.
An omnibus test for a multi-level factor is not in the results table — that is a nested-model comparison, below.
The optional prediction-grid artifact gives you fitted values over a covariate grid, which is what the plotting functions draw.
Subgroups
split.by fits a separate model per stratum and labels each row with the level. Sample sizes and covariate availability differ per stratum, because each is an independent fit.
interactions does something different: one model with a product term, testing whether the slope differs.
These two look almost identical in the results table and mean different things. Say which you used.
Two consequences of split.by worth stating: adding it changes every adjusted p-value in the call, because the stratum is part of the grouping; and a difference between two strata's estimates is not a test of that difference.
Comparisons
associate$compare() answers questions the coefficient table cannot: does adding this score improve the model, and does the effect differ between groups.
type = "incremental" compares a nested pair of models. For a linear model you get a change in R² with an F test. For the others you get a likelihood-ratio test — and no effect estimate, only the statistic, its p-value and degrees of freedom.
type = "contrast" gives pairwise contrasts between group levels from a single interaction model. The number of rows grows quadratically in the number of levels, and they all share one adjustment family.
type = "heterogeneity" is the formal test of whether an effect differs across groups. It lives here rather than in regression().
Two things that distinguish comparisons from regressions: they use a different result schema and carry no per-fit metadata, and they cannot be pooled across cohorts. If your design depends on meta-analysis, that matters before you start.
Comparisons adjust within outcome × family, without the stratum — unlike regression(). Two tables side by side can therefore carry differently-pooled adjusted p-values.
Testing many models at once
Naming many predictors gives you one row per model, which is how a panel of scores gets tested against an outcome.
The grid you asked for is the family. That is worth pausing on when the grid is large and the models overlap: many candidate scores built from the same GWAS are close to a single test, while scores from unrelated traits are closer to independent ones, and an adjusted p-value cannot tell the difference.
A practical habit: compute score-to-score correlation first (chapter 11) and read it as your effective number of tests, then report the raw p-value and the family size alongside the adjusted one.
Sweeping many thresholds and keeping the best is model selection on the same sample, and nothing here corrects for it.
Artifacts and diagnostics
Artifacts come at three levels. "none" skips building them and is the fast option. "minimal" builds everything and then discards most of it, so it is not the speed answer people expect. "full" keeps them.
Failed fits are visible: the per-fit table includes them, so a results table with fewer rows than your grid is a diagnostic rather than an error. diagnostics is the only place a dropped stratum or a convergence failure appears.
output = "models" hands back the fitted objects, which is the escape hatch for anything PolyGenius does not emit.