PolyGenius
Contents

generate$from.pgs.catalog

Import PRS models from the PGS Catalog

generate$from.pgs.catalog() imports one or more polygenic score models from the PGS Catalog over its REST API and resolves them through the PolyGenius execution engine.

Usage

generate.from.pgs.catalog(
  ids,
  id.type = c("auto", "pgs", "publication", "trait")
)

Arguments

ArgumentDescription
idsCharacter vector of identifiers, interpreted according to id.type: "pgs" — PGS Catalog score accessions, e.g. "PGS002280". "publication" — PGS Catalog publication ids ("PGP000001") or PubMed ids (digits only, e.g. "30554720"). "trait" — EFO ids ("EFO_0007992") or free-text trait terms ("Alzheimer's disease"). "auto" — Per entry: ^PGS\\d+$ is a score, ^PGP\\d+$ or all digits is a publication, anything else is a trait term. Duplicates are dropped; an empty vector aborts.
id.typeOne of "auto" (default), "pgs", "publication", "trait". How to read ids.

Value

A PGSLibrary with one PGS per resolved score. Each model carries $gwas (trait, genome build, discovery-sample metadata) and $generation (PRS method, publication and scoring-file provenance, including weight.type, the scale the score was published on, and weight.transform, what PolyGenius applied to reach the log scale).

The set carries a record with name = "generate$from.pgs.catalog", the deparsed call, and a single sources entry recording the requested ids, the id.type used, and the pgs.ids they resolved to. algorithms is NULL, because these models were fitted by their authors and the method lives per model. See the Library-level provenance section of PGSLibrary.

Details

Identifiers are resolved to PGS score accessions first, over lightweight REST calls. Each accession then becomes a cached PGS resource: the scoring file is downloaded from the PGS Catalog FTP into the workspace resource cache, preferring harmonised coordinates (GRCh38, then GRCh37, then the original file), so re-running the same call is a cache hit.

One score failing does not discard the others

A score that cannot be fetched -- unreachable host, missing scoring file, unusable genome build -- is recorded in the returned set's provenance(set)$misc$failed.models (accession, status, reason) and warned about, while every other requested score is still imported. This is the contract generate$models() honours for a failed model.

The call aborts only when no requested score resolved at all. The abort carries the reasons in its message rather than in a returned object, and names the execution log holding the rest, so a total failure is diagnosed from the message, not from the record.

Identifier resolution is a weaker, separate stage: an ids entry matching no PGS Catalog record is dropped silently before any of the above applies, and never appears in failed.models.

A set whose scores span both GRCh37 and GRCh38 -- common for a trait query -- is returned without the multi-build advisory PGSLibrary() prints, because the collection path does not go through that constructor. Call liftover to converge it.

Weights are always on the log scale

$variants$beta is always a log-scale per-allele effect, whatever scale the score was published on. The conversion happens at import, so nothing downstream has to ask:

Published weight_type Applied
beta, log(OR), log(HR), ln(OR), ln(HR) nothing -- already log scale
OR, HR natural log
Log2(OR) rescaled by log(2)
NR or absent nothing, with a warning

Anything else -- Dosage, Unweighted, Odds Ratio over expected risk -- is refused, naming the value: those are not effect scales, and guessing would silently rescale every weight in the score. Build a PGS directly from the variant table to use such weights as they are.

A score whose scale is unreported (NR, the PGS Catalog default when the authors did not document it) is imported with its weights used as betas, and warns. If those weights are all positive with a median near 1 they are a ratio scale mislabelled, and the warning says so specifically -- check the publication before trusting the score.

A score whose scoring file repeats a variant (same chromosome, position, other and effect allele) loads it once: exact copies keep the first row, and a variant listed with different weights is dropped. Either case warns. See PGS().

Genome build

Import fails unless the resolved build is GRCh37/hg19 or GRCh38/hg38. A PGS Catalog score reports NR when no harmonised scoring file is available, and those coordinates cannot be placed against a cohort. Pick a score with a harmonised file, or construct the model directly if you know the assembly.

The returned model's $build is the canonical name ("hg19/GRCh37" or "hg38/GRCh38"), the same form PolyGenius's own producers (GWAS sources, liftover) store. $gwas$genome.build keeps the build the Catalog reported. A model written with $write.pgs.file() declares #genome_build=GRCh37 or GRCh38.

Examples

# Direct PGS IDs
models <- generate$from.pgs.catalog(c("PGS002280", "PGS001828"))

# Mixed auto-detected identifiers
models <- generate$from.pgs.catalog(
  c("PGS002280", "PGP000001", "30554720", "EFO_0007992", "Alzheimer's disease")
)

# By publication
models <- generate$from.pgs.catalog("PGP000001", id.type = "publication")

# By trait term
models <- generate$from.pgs.catalog("type 2 diabetes", id.type = "trait")

See Also

Aliases: generate-from-pgs-catalog, generate.from.pgs.catalog, generate$from.pgs.catalog