Choosing an Association Analysis
What associate does, what comes back, and which analysis answers which question
Two domains, two questions
associate estimates effects: how strongly a score relates to an outcome, adjusted for the covariates you name.
evaluate compares candidate models: which of several scores predicts best. It lives in chapter 11.
The distinction has a concrete edge worth knowing early. evaluate$association() reports the same odds ratio or beta associate$regression() would, for the same covariates — it delegates to associate$regression() directly. What differs is role, not the number: evaluate$association() sits beside evaluate's other predictive metrics for comparing candidates, while associate$regression() is what you reach for directly when you need interactions, another regression family, or the fitted models themselves.
Neither computes a missing score layer for you: both abort, naming the layer, if scores.layer is not already on the data.
What comes back
Every association returns the same four parts. results is one row per statistical claim. artifacts are plot-ready derived tables. diagnostics is what went wrong, was dropped, or failed to converge. fits is per-fit metadata.
Every effect row carries an estimate and an effect.scale, and the estimate is always on the model's coefficient scale — log-odds, log-hazard — never an odds or hazard ratio. Plots exponentiate for display; the stored number does not.
Because the coefficient families share one table shape, a linear model's result still has time, event and n.events columns sitting empty. That is expected, not a bug.
Working with a result
The usual dplyr verbs work and return the same class, so you can filter to one outcome or arrange by effect size. summarise() and pull() step outside the class, as you would expect.
Row filters prune the artifacts and diagnostics along with the rows, so a filtered object's plots match its table.
One thing filtering does not do: re-compute adjusted p-values. Those were calculated over the family that existed when the analysis ran, and they stay that way.
Multiple testing, in principle
A family of tests is what was tested for one outcome within one stratum. Adjusted p-values are computed per call — never across calls, and never across objects.
Effect p-values and heterogeneity p-values are separate families and are never pooled together.
The grouping is a design decision rather than a knob, and it differs between analyses. Each of the pages below states its own, at the point where it matters.
Two consequences people get wrong: running a regression, then a comparison, then an evaluation gives you three independent adjustments and no overall control; and a within-cohort adjusted p-value is not comparable to a pooled one, because pooling re-adjusts in its own family.
Which analysis do you want?
| Your question | Page | Pools across cohorts? |
|---|---|---|
| What is the adjusted effect of this score on this outcome? | 07.1 Regression | Yes |
| Does it differ between subgroups? | 07.1 Regression | Yes |
| Does one model add anything over another? | 07.1 Regression | No |
| How does the effect look over time, or with a competing event? | 07.2 Survival | Cox and Fine–Gray yes; Kaplan–Meier no |
| How much of the effect runs through an intermediate? | 07.3 Mediation | Yes |
| Which individual variants associate with my phenotype? | 07.4 Single-variant | Yes |
| What is the combined effect across my cohorts? | 07.5 Meta-analysis | — |
| Which of my candidate scores is best? | 11 Evaluate | No |
| Mendelian randomization | Not implemented — the call errors | — |
The "pools" column is there for a reason: it is easy to build a whole multi-cohort design on an analysis whose results cannot be combined.
Two boundaries
associate$* never silently runs compute$* — the scores have to exist already.
Neither domain writes into your cohort object. Both return a result that you assign.