Contents
generate$from.pgs.catalog
Import PRS models from the PGS Catalog
generate$from.pgs.catalog() imports one or more polygenic score models from
the PGS Catalog over its REST API and resolves
them through the PolyGenius execution engine.
Usage
generate.from.pgs.catalog(
ids,
id.type = c("auto", "pgs", "publication", "trait")
)Arguments
| Argument | Description |
|---|---|
ids | Character vector of identifiers, interpreted according to id.type: "pgs" — PGS Catalog score accessions, e.g. "PGS002280". "publication" — PGS Catalog publication ids ("PGP000001") or PubMed ids (digits only, e.g. "30554720"). "trait" — EFO ids ("EFO_0007992") or free-text trait terms ("Alzheimer's disease"). "auto" — Per entry: ^PGS\\d+$ is a score, ^PGP\\d+$ or all digits is a publication, anything else is a trait term. Duplicates are dropped; an empty vector aborts. |
id.type | One of "auto" (default), "pgs", "publication", "trait". How to read ids. |
Value
A PGSLibrary with one PGS per resolved score. Each
model carries $gwas (trait, genome build, discovery-sample metadata) and
$generation (PRS method, publication and scoring-file provenance,
including weight.type, the scale the score was published on, and
weight.transform, what PolyGenius applied to reach the log scale).
The set carries a record with name = "generate$from.pgs.catalog",
the deparsed call, and a single sources entry recording the requested
ids, the id.type used, and the pgs.ids they resolved to. algorithms
is NULL, because these models were fitted by their authors and the method
lives per model. See the Library-level provenance section of
PGSLibrary.
Details
Identifiers are resolved to PGS score accessions first, over lightweight REST
calls. Each accession then becomes a cached PGS resource: the scoring file
is downloaded from the PGS Catalog FTP into the workspace resource cache,
preferring harmonised coordinates (GRCh38, then GRCh37, then the original
file), so re-running the same call is a cache hit.
One score failing does not discard the others
A score that cannot be fetched -- unreachable host, missing scoring file,
unusable genome build -- is recorded in the returned set's
provenance(set)$misc$failed.models (accession, status, reason) and warned about,
while every other requested score is still imported. This is the contract
generate$models() honours for a failed model.
The call aborts only when no requested score resolved at all. The abort carries the reasons in its message rather than in a returned object, and names the execution log holding the rest, so a total failure is diagnosed from the message, not from the record.
Identifier resolution is a weaker, separate stage: an ids entry matching no
PGS Catalog record is dropped silently before any of the above applies, and
never appears in failed.models.
A set whose scores span both GRCh37 and GRCh38 -- common for a trait query --
is returned without the multi-build advisory PGSLibrary() prints,
because the collection path does not go through that constructor. Call
liftover to converge it.
Weights are always on the log scale
$variants$beta is always a log-scale per-allele effect, whatever scale the
score was published on. The conversion happens at import, so nothing
downstream has to ask:
Published weight_type |
Applied |
|---|---|
beta, log(OR), log(HR), ln(OR), ln(HR) |
nothing -- already log scale |
OR, HR |
natural log |
Log2(OR) |
rescaled by log(2) |
NR or absent |
nothing, with a warning |
Anything else -- Dosage, Unweighted, Odds Ratio over expected risk -- is
refused, naming the value: those are not effect scales, and guessing would
silently rescale every weight in the score. Build a PGS directly from the
variant table to use such weights as they are.
A score whose scale is unreported (NR, the PGS Catalog default when the
authors did not document it) is imported with its weights used as betas, and
warns. If those weights are all positive with a median near 1 they are a ratio
scale mislabelled, and the warning says so specifically -- check the
publication before trusting the score.
A score whose scoring file repeats a variant (same chromosome, position, other and effect
allele) loads it once: exact copies keep the first row, and a variant listed with different
weights is dropped. Either case warns. See PGS().
Genome build
Import fails unless the resolved build is GRCh37/hg19 or
GRCh38/hg38. A PGS Catalog score reports NR when no harmonised scoring
file is available, and those coordinates cannot be placed against a cohort.
Pick a score with a harmonised file, or construct the model directly if you
know the assembly.
The returned model's $build is the canonical name ("hg19/GRCh37" or
"hg38/GRCh38"), the same form PolyGenius's own producers (GWAS sources,
liftover) store. $gwas$genome.build keeps the build the Catalog reported. A model written with
$write.pgs.file() declares #genome_build=GRCh37 or GRCh38.
Examples
# Direct PGS IDs
models <- generate$from.pgs.catalog(c("PGS002280", "PGS001828"))
# Mixed auto-detected identifiers
models <- generate$from.pgs.catalog(
c("PGS002280", "PGP000001", "30554720", "EFO_0007992", "Alzheimer's disease")
)
# By publication
models <- generate$from.pgs.catalog("PGP000001", id.type = "publication")
# By trait term
models <- generate$from.pgs.catalog("type 2 diabetes", id.type = "trait")