Contents
associate$singleVariant
Single-variant association (GWAS)
associate$singleVariant() runs genome-wide or targeted single-variant tests
through PLINK2 --glm for one or more phenotypes resolved from a
PolyGeniusStudy. It returns a PolyGeniusAssociation with one row per
variant x phenotype x stratum, consumable by associate$meta
and the Manhattan / QQ plots.
Usage
associate.single.variant(
data,
phenotypes,
variants = NULL,
covariates = NULL,
split.by = NULL,
conf.level = 0.95,
reference.level = NULL,
maf = NULL,
mac = NULL,
vif = NULL,
p.adjust.method = "BH",
memory = NULL,
nthreads = NULL
)Arguments
| Argument | Description |
|---|---|
data | A PolyGeniusStudy with at least one attached GenotypeSource. |
phenotypes | Unquoted phenotype expression, resolved on the observation side from data. Numeric columns are treated as quantitative; logical, two-level factor or character, and {0, 1} numeric columns as binary. At least one column must resolve. |
variants | A data.frame with chr and position columns, or NULL (default). NULL takes every variant identity in data's attached model library and aborts if none is attached. The table needs no preparation: chr is normalized internally to the package's canonical spelling (chr1..chr22, chrX, chrY, chrMT), position is coerced to integer, and the result is made unique before it reaches PLINK. A supplied contig PolyGenius does not recognize is passed through as given; on the NULL path one that cannot be resolved is dropped and reported. |
covariates | Unquoted adjustment expression, or NULL (default). Resolved on the observation side and passed to PLINK as variance- standardized covariates. |
split.by | Unquoted stratification expression, or NULL (default). One scan is run per observed level of the interaction of the resolved columns; each column must be factor, character or logical. |
conf.level | Numeric scalar in (0, 1), default 0.95. Passed to PLINK's --ci and reported as lower/upper. |
reference.level | Named list mapping a binary phenotype name to its reference (control) level, or NULL (default). Names not matching a binary phenotype are ignored; unset binaries take PLINK's own deterministic default, and whichever level was used is echoed on every row and in $metadata$reference.levels. |
maf | Numeric scalar minor-allele-frequency floor and integer scalar minor-allele-count floor, or NULL (default). Passed as PLINK's --maf/--mac; NULL omits the flag. |
mac | Numeric scalar minor-allele-frequency floor and integer scalar minor-allele-count floor, or NULL (default). Passed as PLINK's --maf/--mac; NULL omits the flag. |
vif | Numeric scalar, or NULL (default). Covariate variance-inflation ceiling passed as PLINK's --vif. |
p.adjust.method | Character scalar, default "BH". Method passed to stats::p.adjust(), applied within each scan. |
memory | Positive numeric scalar (MB), or NULL (default). PLINK --memory cap; NULL leaves PLINK's own default. |
nthreads | Positive integer scalar, or NULL (default). PLINK --threads count; NULL leaves PLINK's own default. |
Value
A PolyGeniusAssociation whose $results follow the
schema-single-variant schema, one row per variant x phenotype x stratum. The
variant columns use the PolyGenius GWAS/model naming (chr, position,
ea, nea, beta, se, pval), so the table joins directly against
PGS variants and the generate GWAS format. Key $results columns:
outcome, stratum, genotype — Scan identity: phenotype, stratum
level, and genotype dataset the row belongs to.
chr, position, variant.id — Chromosome, base-pair position, and
variant identifier. chr is the package-canonical code (chr1..chr22,
chrX, chrY, chrMT; see pg.chr.canonical()), so it aligns with
model and genotype variants on any join or genome track.
ea, nea, eaf — Effect (tested, A1/ALT) allele, non-effect
(reference) allele, and effect-allele frequency.
beta, se, lower, upper — Effect size and its conf.level
interval, on the scale given by effect.scale: "log.odds" for binary
phenotypes (Firth-fallback logistic), "identity" for quantitative.
statistic, pval, adj.pval, adj.family.id — Test statistic, raw
p-value, the within-scan p.adjust.method adjusted p-value, and the key
into $indices$multiplicity.
n, n.cases, n.controls — Sample size, and case/control counts
for binary phenotypes.
family, outcome.type, effect.scale, reference.level — Fitted
family ("lm"/"glm"), phenotype type, effect scale, and the control
level for binary phenotypes.
lambda.gc — Per-scan genomic-inflation factor
(\lambda_{GC}), also surfaced in $diagnostics$genomic.inflation.
$diagnostics$genomic.inflation holds one row per scan (lambda.gc,
n.variants); $artifacts is empty and $fits is NULL, since the rows
are already narrow. $provenance records the call; $metadata carries
reference.levels and dropped.covariates. A scan that produced no rows returns an association
with the empty "single.variant" prototype and no diagnostics.
Details
Scans
One PLINK2 --glm run is issued per genotype dataset x stratum, restricted to
that stratum's sample IIDs; all phenotypes are tested in the same pass.
Quantitative phenotypes are fitted with linear regression and binary ones with
Firth-fallback logistic regression, decided per phenotype from its column
type. Covariates are written out with --covar-variance-standardize.
Variants sharing an ID are excluded (--rm-dup exclude-all), which removes every copy.
p.adjust.method is applied within each genotype x phenotype x stratum scan,
never across scans; the declared families are recorded in
$indices$multiplicity and keyed from adj.family.id.
Examples
gwas <- associate$singleVariant(
data,
phenotypes = dementia,
covariates = c(age, sex, PCA$PC1, PCA$PC2)
)
gwas$results
gwas$diagnostics$genomic.inflation