Reference
Every documented function, accessor and class, grouped by the namespace it belongs to and ordered the way an analysis runs.
13 topics
generate$
Model generation from GWAS sources and algorithms
- generategenerate is the entry point for building or importing polygenic score models.
- generate$modelsgenerate$models() cross-joins one or more source specifications with one or more algorithm specifications, then resolves the requested models through the execution engine.
- generate$algorithmsgenerate$algorithms declares how a model is built from summary statistics: ClumpingThresholding() (generate$algorithms$ClumpingThresholding), LDpred2() (generate$algorithms$LDpred2), lassosum2()...
- generate$algorithms$ClumpingThresholdingDeclares a clumping + p-value thresholding (C+T) algorithm for generate$models().
- generate$algorithms$lassosum2Declares a lassosum2 algorithm for generate$models().
- generate$algorithms$LDpred2Declares an LDpred2 algorithm for generate$models().
- generate$algorithms$PRScsDeclares a PRS-CS algorithm for generate$models().
- generate$from.pgs.cataloggenerate$from.pgs.catalog() imports one or more polygenic score models from the PGS Catalog over its REST API and resolves them through the PolyGenius execution engine.
- generate$from.pgs.filegenerate$from.pgs.file() reads PGS Catalog scoring files from disk into PolyGenius.
- generate$sourcesgenerate$sources declares where GWAS summary statistics come from: generate$sources$opengwas, generate$sources$gwascatalog and generate$sources$local.
- generate$sources$gwascatalogDeclares one or more GWAS Catalog studies as GWAS sources for generate$models().
- generate$sources$localDeclares in-memory or on-disk GWAS summary statistics as sources for generate$models().
- generate$sources$opengwasDeclares one or more IEU OpenGWAS studies as GWAS sources for generate$models().
15 topics
compute$
PRS score computation and related operations
- computeNamespaced entry point for every transformation that reads or writes a PolyGeniusStudy's sample, model or score axes.
- compute$scoresComputes PRS values for every model in a PolyGeniusStudy by driving PLINK2 --score over the study's genotypes.
- compute$embeddingcompute$embedding$samples() embeds samples and compute$embedding$models() embeds models, from any matrix-shaped study slot or from a pre-computed similarity matrix (PCA, MDS, UMAP, t-SNE, PHATE).
- compute$genomecompute$genome is not a general toolkit for primary genomic or variant analysis.
- compute$genome$attributioninternalLocalizes where on the genome each trait-PRS's outcome signal sits.
- compute$genome$concordanceinternalReduces a PGS library to a per-variant PolyGeniusGenomeSignal measuring whether the models that carry a variant agree on its harmonized effect direction, each model's vote weighted by the magnitude of the weight it...
- compute$genome$convergenceinternalShows where along the genome the model library's outcome signal concentrates.
- compute$genome$cumulativeWeightinternalSums the absolute harmonized effect weight |beta| a model library places at each variant -- a weighted alternative to the plain reuse count.
- compute$populationStructureComputes sample principal components from a study's genotypes with PLINK2.
- compute$relatednesscompute$relatedness finds related samples in a PolyGeniusStudy and chooses which to keep so that no two retained samples are related.
- compute$relatedness$kinshipRuns PLINK 2's KING-robust estimator within each genotype fileset of a PolyGeniusStudy and returns the related pairs as a sparse sample-by-sample matrix, ready to store in data$sample.pairs.
- compute$relatedness$pruneDecides which samples of a PolyGeniusStudy to keep so that no two kept samples are related above the requested degree, using the greedy algorithm of PLINK 2 --king-cutoff.
- compute$scores.planPredicts the peak RAM, batch count and scoring strategy compute$scores() would use for a workload, and reports what a given memory value resolves to.
- compute$similaritycompute$similarity$samples() compares samples by their PRS score profiles; compute$similarity$models() compares models by SNP overlap or by their score vectors.
- compute$standardizeBuilds a new score layer from an existing one, centered and scaled per model, and tags every model with the (center, scale, scale.source) triple used.
7 topics
associate$
Association analyses: regression, mediation, MR, meta
- associateassociate is an environment bundling the package's association entry points, each of which takes a PolyGeniusStudy and returns a PolyGeniusAssociation with $results, $artifacts, $diagnostics, $metadata...
- associate$regressionassociate$regression() fits one association model per outcome x predictor x interaction x stratum cell resolved from a PolyGeniusStudy, picking a classical or survival family per outcome.
- associate$compareassociate$compare() answers one of three comparison questions about predictors and outcomes resolved from a PolyGeniusStudy: does adding a predictor improve fit, is its effect homogeneous across groups, and how...
- associate$mediationassociate$mediation() decomposes the effect of one exposure on one outcome into an indirect (ACME), direct (ADE) and total effect through one mediator, all resolved from a PolyGeniusStudy.
- associate$metaassociate$meta() pools the summary rows of two or more PolyGeniusAssociation objects that share one source schema, cell by cell, with a fixed- or random-effects estimator.
- associate$mrassociate$mr() reserves the Mendelian-randomization surface.
- associate$singleVariantassociate$singleVariant() runs genome-wide or targeted single-variant tests through PLINK2 --glm for one or more phenotypes resolved from a PolyGeniusStudy.
9 topics
evaluate$
Discrimination, calibration, stratification, incremental
- evaluateevaluate bundles the predictive assessment of the models in a score layer.
- evaluate$profileevaluate$profile() runs the evaluate components in one call, so one object shows how well each PRS model predicts, associates, stratifies risk and compares with the others.
- evaluate$associationevaluate$association() asks how strongly each score is associated with each outcome: an odds ratio for a binary outcome, a regression coefficient for a continuous one, optionally adjusted for covariates.
- evaluate$compareevaluate$compare() asks how much better or worse each model predicts an outcome than one reference model does, as the paired difference in one performance metric.
- evaluate$incrementalevaluate$incremental() asks what each score adds to a model on the covariates alone.
- evaluate$performanceevaluate$performance() asks how well each score, on its own, predicts each outcome: how it separates cases from controls, or how much outcome variance it explains.
- evaluate$redundancyevaluate$redundancy() asks how much two models overlap: in the variants they use, or in the scores they give the same samples.
- evaluate$selectevaluate$select() ranks the models of an existing evaluation, per outcome, on one metric or a weighted composite of several.
- evaluate$stratificationevaluate$stratification() asks how the outcome changes across percentile bands of each model's score.
36 topics
visualize$
Plot functions for association and evaluation results
associations
- visualize$associations$forestPublication-style forest plot of the effect estimates in a PolyGeniusAssociation, laid out as names | table.left | forest | table.right.
- visualize$associations$heatmapDraws one cell per (row, column) combination of an association table, filled by color.by and starred by asterisk.by.
- visualize$associations$landscapeOne point per PRS-level association, grouped along the x-axis by group.by and placed at the -\log_{10} of the p-value column pval names.
- visualize$associations$survivalDraws the survival-curve artifacts carried on a summary-mode PolyGeniusAssociation from associate$regression.
- visualize$associations$variants.qqQuantile-quantile plot of a GWAS scan from associate$singleVariant: observed against expected -\log_{10}(p) under the null, with the 1:1 reference line and the genomic-inflation factor \lambda_{GC} annotated.
data
- visualize$data$embedding$samplesinternalScatters two dimensions of a stored embedding, one point per sample or per model.
- visualize$data$scores$distributionDraws one data$scores layer as a per-model distribution -- violin, density, histogram or boxplot -- optionally grouped and faceted by sample phenotypes.
- visualize$data$scores$distribution.heatmapReduces one score layer to a model-by-distribution heatmap: one row per model, one column per density grid point or histogram bin, on a score grid shared by every row so rows are comparable.
- visualize$data$scores$heatmapDraws one matrix from data$scores as a samples-by-models heatmap using ComplexHeatmap, or models-by-samples when flip.orientation = TRUE.
- visualize$data$similarity$samplesinternalDraws one stored pairwise matrix as a clustered heatmap, using ComplexHeatmap and circlize.
evaluate
- visualize$evaluate$associationForest plot of evaluate$association()'s per-model effect -- odds ratio for a binary outcome, beta for a continuous one -- faceted by outcome, with split.by strata dodged and colored within each model's row.
- visualize$evaluate$compareForest plot of evaluate$compare()'s per-model contrast against its reference model, faceted by outcome (each panel naming its reference).
- visualize$evaluate$incrementalForest plot of evaluate$incremental()'s deltas -- delta AUC, delta liability R-squared, delta R-squared -- one row per model, faceted by outcome and metric (columns) and stratum (rows).
- visualize$evaluate$performanceRenders evaluate$performance()'s curves from its own artifacts: a binary outcome draws ROC or precision-recall from the confusion counts, a continuous outcome draws binned score against outcome from...
- visualize$evaluate$profilePlots an evaluate$profile() result: one row per model, one column per metric, each point the stored estimate with its interval where one exists.
- visualize$evaluate$redundancyModel-by-model geom_tile of one evaluate$redundancy() method.
- visualize$evaluate$stratificationPlots evaluate$stratification()'s per-band contrast -- an odds ratio for a binary outcome, a mean difference for a continuous one -- by band, one point-and-interval series per model, faceted by outcome (columns) and...
execution
- visualize$execution$dashboardinternalRenders a self-contained HTML dashboard for one execution from its structured <exec_id>.jsonl event stream: run summary, a time-scrubber, the collapsed dynamic execution graph (type -> rule -> type), per-rule...
- visualize$execution$performanceinternalRenders the execution dashboard's performance charts as live ggplot2 objects from one execution's <exec_id>.jsonl event stream.
genome
- visualize$genome$attributioninternalRenders an attribution PolyGeniusGenomeSignal as a per-trait lane heatmap: one row per model, genome on x, each cell filled with the summed weight x single-variant effect that model attributes to the bin, on a...
- visualize$genome$concordanceinternalRenders a concordance PolyGeniusGenomeSignal as a signed genome track: one point per variant at the consensus S = value.num / value.den in [-1, 1], coloured on a diverging ramp and sized by n.models.
- visualize$genome$convergenceinternalRenders a convergence PolyGeniusGenomeSignal as a binned area track: per display bin, the summed |weight x outcome effect| the contributing models place there.
- visualize$genome$coverageCompanion coverage track to visualize$genome$effects: per genomic bin, the number of distinct models contributing a variant there.
- visualize$genome$cumulativeWeightinternalRenders a cumulative-weight PolyGeniusGenomeSignal as a binned area track: per display bin, the summed |beta| the model library places there.
- visualize$genome$effectsMiami-style genome track for exploring candidate pleiotropy across a model set.
- visualize$genome$lociA thin track of labelled rectangles, one per interval, for marking hotspots or candidate regions above the data tracks of a visualize$genome$stack.
- visualize$genome$manhattanManhattan-style genome track of per-variant association strength from associate$singleVariant.
- visualize$genome$overviewCompute the default genome signals for a model library and stack their tracks in one call: the applied shortcut for the compute-then-stack workflow, readable at any library size.
- visualize$genome$prsGenome track relating each model's variants to its PRS-level association with outcome.
- visualize$genome$reuseManhattan-style genome track of variant reuse: each point is a harmonized variant at its genomic position, its height the number of models carrying it, on a pseudo-log10 y-axis.
- visualize$genome$stackCompose genome tracks vertically onto one shared genome x-axis.
models
- visualize$models$reuseHistogram of how many models each harmonized variant appears in, one bar per integer count.
- visualize$models$sizesHistogram of the number of distinct variants per model across a PGS library, on a log10 x-axis.
- visualize$models$top.variantsBar chart of the variants that appear in the most models, most-reused first.
- visualize$models$uniquenessHistogram of the fraction of each model's variants that appear in no other model.
10 topics
workspace$
Workspace configuration and catalog access
- workspaceThe runtime context generate, compute, associate, evaluate and visualize all resolve against: configuration, resource catalogs, external-tool setup and execution internals.
- workspace$catalogsManifest-backed reference data, reached as workspace$catalogs.
- workspace$catalogs$genomeBuildsinternalRegistry of the genome builds PolyGenius supports.
- workspace$catalogs$LDblocksinternalInventory of approximately-independent LD block boundary sets (Berisa & Pickrell), keyed by population and genome build.
- workspace$catalogs$LDsinternalRead-only inventory of the built LD references available to PolyGenius.
- workspace$catalogs$liftoverChainsinternalInventory of the UCSC-style chain files used to convert coordinates between genome builds, keyed by (from, to) build pair.
- workspace$catalogs$referencePanelsinternalInventory of the reference panels available to PolyGenius, and the entry point that resolves one into the workspace cache.
- workspace$catalogs$variantSpacesinternalInventory of named variant-range sets, used to restrict a reference panel and as the default marker set for PCA and kinship (common20k).
- workspace$configThe runtime settings every generate, compute, associate, evaluate and visualize call resolves against.
- workspace$setupDownloads, registers and resolves the tools the pipeline shells out to: PLINK2, GCTB, the PRS-CS Python environment and the bigsnpr R stack.
97 topics
Classes & helpers
Classes, results and schemas, statistics, the execution engine and utilities
generate
visualize
PGS Library
- as.PGSCrosses from the shared-backbone side back to the owned side, explicitly: no code path materializes a view implicitly, so a caller always knows when it is paying for a copy of a model's variant table.
- comparePGSLibrariesCompares two libraries' model identity -- the ordered position -> (name, build) map -- without touching variant data.
Pgs objects
- PGSPGS is an R6 class holding one polygenic score model: a data.table of variants plus package- and user-defined metadata (GWAS metadata, generation logs, liftover logs, and whatever else a rule or import path...
- PGSLibraryPGSLibrary is an R6 container holding N polygenic score models.
- PGSLibraryViewinternalA PGSLibraryView is a non-owning (backbone, position) pointer to one model in a shared PolyGeniusBackbone, handed out by PGSLibrary$models.
- PolyGeniusBackboneinternalOne dictionary row per distinct variant identity (chr, position, nea, ea), shared by every model absorbed into it.
- PolyGeniusWeightStoreinternalOne (dict.slot, beta) pair-vector per absorbed model, indexed by model.slot -- a model's position in PolyGeniusBackbone's own $registry, assigned in absorb order.
Data Classes
- loadPolyGeniusReads a .pgd file and rebuilds the object it holds through real constructors, never by readRDS()-ing a live R6 object.
- PolyGenius-backendssavePolyGenius() and loadPolyGenius() write and read one kind of file: a saveRDS() of a version-stamped plain payload, suffix .pgd, whose class field is what the load side dispatches on.
- savePolyGeniusWrites study to a file (suffix .pgd) that loadPolyGenius() can read back, on this machine or any other.
Genome signals
- PolyGeniusGenomeSignalPolyGeniusGenomeSignal is the result object for a value defined over genome position: single-variant statistics, cross-model reductions (reuse, coverage, directional concordance, cumulative weight), per-trait...
- PolyGeniusGenomeTrackinternalAn S3 list returned by any visualize$...$*.track() builder: a plain spec (mark, data, params, positions, height, label, build) that holds no closures.
- PolyGeniusVariantFateinternalThe record scores.pack.variant.fate() retains on the score matrix: a list with $ids (the variant identities), $codes (one byte per variant per genotype: status in bits 0-2, flipped in bit 3,...
Genotypes
- as.bigsnprWrites the genotypes a GenotypeSource() points at into bigsnpr's file-backed .rds/.bk pair and attaches the result.
- GenotypeSourcePoints PolyGenius at a genotype fileset on disk and gives it a name, format and genome build.
- GenotypeSourceSetGroups several GenotypeSource objects into one set that fans a call out across all of them.
Study
- alignTags value so that study$samples$<name> <- align(value, by = ...) matches its rows to the study's samples by the IDs by gives.
- PolyGeniusFrameA tibble subclass holding fetched sample-side ("samples") or model-side ("models") tabular data together with its role metadata.
- PolyGeniusFrameSetThe container study$fetch(...) returns when one fetch resolves on both axes: a list holding the sample-side frame at $samples and the model-side frame at $models.
- PolyGeniusStudyBinds a PRS model library to a set of genotype filesets, producing the object every compute$(), associate$(), evaluate$() and visualize$() verb takes as its first argument.
- reductionTags value so that study$samples$<name> <- value files it as an embedding rather than as a plain table member.
- rolesroles() reads the named role mapping stored on a fetched PolyGeniusFrame.
Results
- PolyGeniusProvenanceConstructs one operation-provenance record for attachment to a PolyGenius object: what was run, what the caller supplied, and what the operation resolved or produced.
- PolyGeniusSchemaConstructs the column and artifact contract for the result table produced by one analysis type.
- schema-evaluationinternalEvaluate has no PolyGeniusSchema object.
Association schemas
- schema-comparisoninternalColumn and artifact contract for results of formal model, effect and group comparisons.
- schema-coxinternalColumn and artifact contract for results of Cox regression on a time-to-event outcome.
- schema-crrinternalColumn and artifact contract for results of Fine-Gray regression with a competing event.
- schema-glminternalColumn and artifact contract for results of logistic regression on a binary outcome.
- schema-kminternalColumn and artifact contract for results of grouped Kaplan-Meier and log-rank analyses.
- schema-lminternalColumn and artifact contract for results of linear regression on a continuous outcome.
- schema-mediationinternalColumn and artifact contract for mediation results, one row per decomposed effect.
- schema-metainternalColumn and artifact contract for pooled meta-analysis rows.
- schema-single-variantinternalColumn and artifact contract for single-variant scan results.
Result objects
- artifactsReads and writes the artifact tables an analysis emitted alongside its results.
- diagnosticsReads and writes the quality and coverage reports an operation emitted alongside its value.
- federateMakes a result safe to send out of a study site: returns a copy of x keeping only the $artifacts entries declared federation.safe.
- PolyGeniusAssociationPolyGeniusAssociation is the standard result object returned by association workflows.
- PolyGeniusEvaluationPolyGeniusEvaluation is the standard result object returned by the evaluate family of functions.
- PolyGeniusResultinternalPolyGeniusResult is the base every PolyGeniusAssociation and PolyGeniusEvaluation is built on.
- provenanceReads and writes the record of the call that produced a value: what was run, what the caller supplied, and what the operation resolved.
Statistics
Comparison families
- ComparisonFamilyinternalComparisonFamily dispatches the incremental, heterogeneity and contrast comparisons for associate$compare().
- CompetingRiskComparisoninternalComparison family for competing-risk outcomes (family key "crr").
- CoxComparisoninternalComparison family for time-to-event outcomes (family key "cox").
- LinearComparisoninternalComparison family for continuous outcomes (family key "lm").
- LogisticComparisoninternalComparison family for binary outcomes (family key "glm").
Outcome specs
Regression families
- CompetingRiskRegressioninternalFits Fine-Gray subdistribution-hazards models as a weighted Cox model on the Fine-Gray risk set. associate$regression instantiates this family for model = "crr", which model = "auto" also picks for a surv()...
- CoxRegressioninternalFits right-censored proportional-hazards models with survival::coxph(). associate$regression instantiates this family for model = "cox", which model = "auto" also picks for a surv() outcome that declares no...
- KaplanMeierRegressioninternalFits grouped Kaplan-Meier curves with survival::survfit() and tests them with survival::survdiff(). associate$regression instantiates this family only for an explicit model = "km"; "auto" never infers it.
- LinearRegressioninternalFits Gaussian linear models with stats::lm() for continuous outcomes. associate$regression instantiates this family for model = "lm", which is also what model = "auto" falls back to when the outcome is neither...
- LogisticRegressioninternalFits binomial generalized linear models with stats::glm(family = binomial()) for binary outcomes. associate$regression instantiates this family for model = "glm", which model = "auto" also picks for a logical...
- RegressionFamilyinternalRegressionFamily defines the shared interface implemented by each regression family used by associate$regression() and the evaluate engine.
Execution Engine
Algorithm rules
- BuildLDBigsnprRuleinternalReactive rule that materializes an ld.bigsnpr resource directly from a tidy reference.panel input using bigsnpr-native matrix operations.
- BuildLDBlocksRuleinternalBuildLDBlocksRule materializes a block-dense ld.blocks reference layout (per-block dense LD matrices in PRS-CS HDF5 format).
- ClumpVariantsRuleinternalClumpVariantsRule resolves one clumped.variants output by loading a GWAS resource and a matching reference panel, then running panel clumping.
- RunLassosum2RuleinternalRunLassosum2Rule resolves lassosum2 polygenius.model outputs.
- RunLdPred2RuleinternalRunLdPred2Rule resolves LDpred2 polygenius.model outputs.
- RunPrscsRuleinternalRunPrscsRule resolves one polygenius.model output for PRS-CS.
- ThresholdClumpedRuleinternalThresholdClumpedRule resolves one polygenius.model output for algorithm ClumpingThresholding.
Catalog rules
- ConvertReferencePanelRuleinternalConvertReferencePanelRule handles format conversion only.
- DownloadLDBlocksRuleinternalDownloadLDBlocksRule materializes an ld.block.boundaries resource by downloading the configured Berisa-Pickrell BED file into the resource store.
- DownloadLiftoverChainRuleinternalDownloadLiftoverChainRule materializes a liftover.chain resource by downloading the configured chain file into the resource store.
- DownloadReferencePanelRuleinternalDownloadReferencePanelRule handles panel requests where only retrieval is required.
- DownloadVariantSpaceRuleinternalDownloadVariantSpaceRule handles variant-space requests that are directly downloadable at the requested build -- no liftover needed.
- LiftoverLDBlocksRuleinternalLiftoverLDBlocksRule produces a block set for a build that is not published by lifting the same population's set from a build that is, then re-imposing the canonical partition.
- LiftoverReferencePanelRuleinternalLiftoverReferencePanelRule handles genome-build transformation only.
- LiftoverVariantSpaceRuleinternalLiftoverVariantSpaceRule handles variant-space requests that are only available at a different build: lifts a source variant space over to the requested build via the standard liftover-chain machinery.
- RestrictReferencePanelRuleinternalRestrictReferencePanelRule creates a reference-panel resource restricted to a named variant space.
Engine core
- executeinternalResolve one or more output resource specifications through the reactive execution engine.
- ExecutionEngineinternalInternal execution engine that orchestrates rule-based dependency resolution and scheduling for requested resources.
- ExecutionSchedulerinternalInternal controller for one reactive execution.
- ResourceResolverinternalLocked environment exposing $resolve() and $resolve.async() over an engine.
- resources.factoryinternalResource specification factory
- resources.typeinternalResource type registry
- ResourceSpecinternalLightweight resource descriptors used by the reactive execution framework.
- ResourceSpecSetList-like container for collections of specification objects, normally ResourceSpec.
- RuleinternalBase R6 class for reactive production rules.
- RuleRegistryinternalInternal registry of reactive rules.
- TaskExecutorinternalLocked environment over an engine with members $await(), $execute.status(), $execute.summary(), $execute.graph() and $last.execution(), and read-only max.cores, max.memory and requested.max.cores.
Setup rules
- ResolveGctbRuleinternalReactive rule that resolves the GCTB binary (used by SBayesR) into the workspace store.
- ResolvePlinkRuleinternalReactive rule that resolves the PLINK2 binary into the workspace store.
- SetupBigsnprRuleinternalReactive rule for the R package dependencies required by the LDpred2, lassosum2, and bigsnpr LD algorithms.
- SetupLiftoverRuleinternalReactive rule that ensures the Bioconductor packages required to lift variants between genome builds are installed in the active R library.
- SetupPrscsRuleinternalReactive rule that creates the PRS-CS Python environment.
Source rules
- FetchGWASCatalogRuleinternalFetchGWASCatalogRule resolves a single GWAS Catalog request into one gwas.sumstats resource.
- FetchOpenGWASRuleinternalFetchOpenGWASRule implements end-to-end OpenGWAS retrieval for gwas.sumstats outputs.
- FetchPGSCatalogRuleinternalFor each requested pgs.catalog.model spec the rule: - fetches full score metadata from GET /rest/score/{pgs_id}/; - selects the best available scoring file (GRCh38 harmonised -> GRCh37 harmonised -> original) and...
- LoadLocalGWASRuleinternalLoadLocalGWASRule resolves a local GWAS request into one gwas.sumstats resource.