PolyGenius
Contents

associate$meta

Meta-analysis of PolyGeniusAssociation results

associate$meta() pools the summary rows of two or more PolyGeniusAssociation objects that share one source schema, cell by cell, with a fixed- or random-effects estimator. Only $results is pooled; the inputs' artifacts are dropped, so the returned object supports inference, forest plots and heatmaps but not cohort-specific curve plots. By default on the model axis, each input's weight in each pooled cell is returned as the study.weights artifact.

Usage

associate$meta(
  ...,
  by = NULL,
  method = c("fixed", "random"),
  conf.level = 0.95,
  output = c("summary", "models"),
  artifacts = NULL
)

Arguments

ArgumentDescription
...One or more summary-mode [PolyGeniusAssociation](/reference/polygeniusassociation/) objects, or a single list of them. Argument names label each input as its study.id in study.weights; when any input is unnamed, all are labelled study1, study2, ... A single input is accepted and pools every cell over one study. A named argument this function does not take is read as an input too, so a misspelled argument fails the class check.
byCharacter vector of extra column names, or NULL (default). Added to the schema's own grouping keys; a name not present on the bound table is silently ignored.
methodOne of "fixed" (default), "random".
conf.levelNumeric scalar in (0, 1), default 0.95. Confidence level for lower/upper and for the prediction interval.
outputOne of "summary" (default), "models".
artifactsOne of "full", "minimal", "none", or NULL (default). "full" keeps study.weights, "minimal" keeps nothing, and "none" skips building it. NULL resolves to "full" for coefficient and mediation sources and to "none" for "single.variant". Ignored when output = "models". See Per-study weights.

Value

With output = "summary" (default), a PolyGeniusAssociation: $results — One row per pooled cell, schema = "meta" (schema-meta) and source.schema naming what it was pooled from, plus the source schema's grouping keys, group.id, estimate, se, lower, upper, statistic, pval, adj.pval, adj.family.id, n.studies, pool.method (mirrored as model), heterogeneity.q, heterogeneity.df, heterogeneity.pval, heterogeneity.adj.pval, heterogeneity.adj.family.id, tau2, i2, pi.lower, pi.upper, the source schema's summed count columns and weighted-mean columns. The column set depends only on the source schema and the by columns found: a key, count or mean column no input carries is NA. The inputs' own rows are not repeated here.

$artifacts — study.weights when artifacts resolves to "full", otherwise empty. The inputs' own artifacts are never pooled.

$diagnostics — dropped.cells when any cell was non-estimable, and harmonization (n.flipped, dropped.mismatch, dropped.ambiguous) for single-variant sources.

$fits — NULL -- meta rows are already narrow.

$metadata — analysis.type = "meta", input.schema, outcome.type (or "mixed"), and meta, the machine-readable re-entry record (source.schema, keys, counts, means) a second pass reads. With output = "models", a plain data.frame of the same pooled cells.

Details

Eligible sources

A schema is poolable when it declares meta-analysis keys: the four coefficient families ("lm", "glm", "cox", "crr"), "mediation", and "single.variant". "km" and "comparison" declare none and are rejected -- a log-rank omnibus row carries no coefficient-scale standard error. Every input must resolve to the same source schema, and an input that is itself a meta result re-enters on the schema it was pooled from, as one study. An input is a meta result when every row's schema is "meta"; an input that mixes meta rows with raw rows is rejected.

Cells and estimators

One cell is a unique combination of the schema's own keys plus any by columns. Coefficient cells key on family, outcome, predictor, term, term.type, stratum, effect scale and test; single-variant cells key on outcome, chromosome, position, the canonical allele pair, test, effect scale and stratum -- deliberately not on genotype, so one scan split across genotype datasets collapses into one row per variant.

"fixed" is inverse-variance weighting (w_i = 1/se_i^2). "random" is DerSimonian-Laird: \tau^2 = \max((Q - df)/C, 0) with df = k - 1 and C = \sum w_i - \sum w_i^2 / \sum w_i, re-weighting to w_i^* = 1/(se_i^2 + \tau^2). Both report Cochran's Q, its p-value and I^2 = \max((Q - df)/Q, 0); "fixed" reports \tau^2 = 0 by definition. Cell effect tests are two-sided Wald tests. No other \tau^2 estimator and no Hartung-Knapp-Sidik-Jonkman adjustment is implemented.

Effect p-values are BH-adjusted into adj.pval, one adjustment family per combination of outcome, family and stratum. A family holds the pooled cells of this call only, so adj.pval means the same thing on every row, and separate calls are separate families. The heterogeneity Q-tests are a separate family of hypotheses, BH-adjusted in their own pool on the same grouping into heterogeneity.adj.pval. Both families are recorded in $indices$multiplicity, keyed by adj.family.id and heterogeneity.adj.family.id. Each input's own estimates and adjusted p-values stay on that input.

Reading I-squared

Read n.studies next to I^2, always. I^2 is a ratio, not an absolute measure: at fixed \tau^2 it rises with per-study precision, so two pools with identical biological heterogeneity report different I^2 purely because their studies differ in size -- a real risk when per-cohort n spans orders of magnitude. Its confidence interval is wide at small k (at n.studies = 4 it can span nearly all of [0, 1]), and this function computes no such interval. The \max(\cdot, 0) floor puts a point mass at exactly 0, so i2 == 0 means Q \le df, not "no heterogeneity".

Prediction interval

pi.lower/pi.upper bound where a new, comparable study's true effect is expected to fall -- a different quantity from lower/upper, the confidence interval on the pooled mean. Computed as \hat\theta \pm t_{df}\sqrt{\tau^2 + se^2} with df = n.studies - 2 (Higgins, Thompson & Spiegelhalter 2009), using a t-quantile rather than qnorm(). NA for method = "fixed" (a fixed-effect model assumes one common true effect) and for n.studies < 3, never a silently emitted NaN.

Single-variant harmonization

Raw single-variant inputs are oriented to a cohort-independent allele pair before pooling: a1.canonical is the lexicographically larger allele and the one every signed effect is oriented to, a2.canonical the smaller, and beta, statistic, eaf, lower and upper are flipped on rows coded the other way. Strand-ambiguous palindromes (A/T, C/G) are dropped, as are rows whose allele pair disagrees with the modal pair at that (chr, position) -- the modal vote is weighted by n.studies, so a pre-pooled unit is not outvoted by one discordant cohort. An input that is already a meta result is trusted as canonical and skips this step. Counts land in $diagnostics$harmonization; the pool aborts if palindrome removal leaves no rows.

Non-estimable cells

An input row contributes to the pooled effect only with a finite estimate and a positive se. Non-estimable rows are dropped from their cell; a cell that loses every row is dropped from the pool and recorded in $diagnostics$dropped.cells with a reason.code of "missing.estimate", "missing.se" or "non.positive.se", and one warning summarizes the drop by reason. The call aborts only when no estimable cell remains. Summed counts (n, n.cases, ...) and weighted means (eaf) are taken over all rows of a surviving cell, so n.studies can be smaller than the number of studies backing its n. A count or mean that no row of a cell reports is NA, not 0.

Pooling "cox"/"crr" cells whose studies declare a different time.origin, or that disagree on whether an entry column was present, emits a warning and proceeds: such coefficients are only comparable on a shared time origin and truncation convention.

Per-study weights

$artifacts$study.weights records how much each input contributed to each pooled cell. It is a data.table with one row per cell and input: group.id — The pooled cell, as in $results$group.id.

study.id — The input, labelled as described under ....

weight — The inverse-variance weight the estimator gave the input: 1/se_i^2 for "fixed" and 1/(se_i^2 + \tau^2) for "random". An input with several rows in one cell, such as a scan split across genotype datasets, gets the sum of their weights.

weight.pct — weight as a percentage of the cell's total, so each cell's rows sum to 100. Only estimable rows carry a weight. An input with no estimable row in a cell has no row for that cell, and a dropped cell has no rows at all. An input that is itself a meta result counts as one study, weighted by its pooled se. The table carries no estimate or se; those are on the inputs. A row filter on the result keeps the weights of the cells that survive it. The artifact is built by default for coefficient and mediation sources, but not for "single.variant", where it holds one row per variant and input.

Examples

amsterdam <- associate$regression(study.a, outcomes = dementia,
                                  predictors = everything())
rotterdam <- associate$regression(study.b, outcomes = dementia,
                                  predictors = everything())

pooled <- associate$meta(amsterdam = amsterdam, rotterdam = rotterdam,
                         method = "random")
pooled$results[, c("predictor", "estimate", "se", "i2", "n.studies")]
pooled$artifacts$study.weights
pooled$diagnostics$dropped.cells

See Also

Aliases: associate-meta, associate.meta, associate$meta