Meta-Analysis
Combining association results across cohorts without moving individual-level data
What it pools
associate$meta() takes association results from several cohorts and combines them cell by cell, on the coefficient scale.
Whether an analysis can be pooled is declared by the analysis itself rather than by a list here: the four coefficient families and mediation and single-variant scans can; Kaplan-Meier and model comparisons cannot. Chapter 07 has the table.
A cell is one comparable quantity — the same predictor, outcome, family and stratum across cohorts. You can subdivide further if you need to.
Fixed and random effects
Fixed effects weight each cohort by the inverse of its squared standard error, which gives the most precise combined estimate under the assumption that all cohorts share one true effect.
Random effects adds a between-cohort variance term, estimated by DerSimonian-Laird, and falls back to the fixed-effect answer when that variance estimates to zero.
Only DerSimonian-Laird is available — no REML, no Paule-Mandel, no Hartung-Knapp interval adjustment. It is biased low with few cohorts, and the heterogeneity percentage is barely interpretable below about five, so say which you used and how many cohorts you had.
A worked example is worth more than the formulae here: two cohorts with the same standard error and effects of 0.20 and 0.30 give a combined 0.25, a standard error of 0.0707, and a heterogeneity statistic of 0.5 on one degree of freedom — so no between-cohort variance, and random effects collapses to fixed.
Rules that are not negotiable
Pool on the coefficient scale, never a ratio. Log-odds, log-hazard, betas. Back-transformation is for display. This is why a proportion cannot be pooled at all.
Effect p-values and heterogeneity p-values are separate families, each adjusted in its own pool. This is the only page where you see two adjusted-p columns side by side, so it is where that principle earns its keep.
The family and the effect scale are part of what defines a cell, so a cause-specific hazard and a subdistribution hazard for the same predictor never combine.
Kaplan-Meier cannot be pooled. That is a statistical refusal, not a gap.
Reading a pooled result
The output uses one shape regardless of what went in, and records what it came from. Sample counts are summed across cohorts; frequencies are averaged with sample-size weighting.
There is one result row per pooled cell and nothing else. Each cohort's own rows stay in the object you passed in, so keep those objects if you want to show them next to the pooled estimate.
How much each cohort contributed to a cell is in the study.weights artifact: one row per cell and cohort, with the inverse-variance weight the estimator used and its percentage share of the cell.
Partial losses are quiet. A cohort whose estimate was not usable drops out of its cell, and only a cell that loses every cohort is recorded in the diagnostics. Compare the reported number of studies against how many you passed in.
For the same reason the study count can be smaller than the number of cohorts behind the sample size, because counts contribute even when an estimate does not.
The cohorts' own artifacts are not merged, deliberately: forest plots yes, per-cohort survival curves no, because there is no coherent average of a study-specific curve. Keep the per-cohort objects if you need them. The pooled result carries one artifact of its own, study.weights. It is built by default when pooling models, and not for variant-level pooling, where it would hold a row per variant and cohort; ask for it with artifacts = "full".
Variant-level pooling
Variant results are harmonised before pooling: alleles are put in a canonical order, and a cohort reporting the opposite orientation has its effect sign, frequency and interval flipped to match.
Palindromic variants, where strand cannot be resolved from the alleles alone, are dropped. A cross-cohort check also drops variants whose allele pair disagrees between cohorts. Both are reported.
Two-stage pooling
A pooled result can itself be pooled, which is how you combine several genotyping arrays within a cohort and then combine cohorts.
The manuscript's own design is the worked example: three genotype datasets within one cohort combined with fixed effects, then cohorts combined with random effects.
Two things the code cannot decide for you: whether a first-stage unit enters the second stage as one study, and what the study count then means. Two-stage equals one-stage only under fixed effects with consistent weights.
What PolyGenius cannot verify
Proportional hazards within each cohort, identical coding of a mediation across cohorts, strand resolution for palindromic variants, and whether your cohorts share samples.
A note on Mendelian randomization
associate$mr() is not implemented — it validates its arguments and then errors.
The note belongs on this page because MR is a pooling problem: two-sample MR is itself an inverse-variance meta-analysis over instruments, so combining MR estimates across cohorts would be a second stage on top of that.
If you need it today, build instrument-level effects with a single-variant scan and compute the estimate outside PolyGenius. Note that the strand handling above applies: palindromic variants are dropped, which for MR means losing instruments rather than silently recoding them.