Getting a Model In and Out
Importing published scores from the PGS Catalog or a file, and writing your own back out
Importing from the PGS Catalog
generate$from.pgs.catalog() takes score accessions, publication identifiers or trait terms, and returns a PGS library.
Identifier type is detected automatically: a PGS accession is a score, a PGP accession or a bare number is a publication, and anything else is treated as a trait term.
A resolved model carries its variant table, the effect and other alleles, the weights, the genome build, and its provenance — which method produced it, which publication, and which scoring file it came from.
Scoring files are chosen in a fixed order: harmonised GRCh38, then harmonised GRCh37, then the original submission. The build you get comes from whichever file was used, which is why target.build matters.
Import goes through the resource store, so re-running the same accession is instant.
One caveat on breadth: a trait term can resolve to hundreds of scores, and resolution stops quietly at the first page it cannot fetch rather than erroring. There is no in-package preview of what a term will resolve to, so query the Catalog's REST interface first if the count matters.
Weights are always on the log scale
$variants$beta is always a log-scale per-allele effect, whatever scale the score was published on. The conversion happens here, at import, so nothing downstream has to ask.
Published weight_type |
What is applied |
|---|---|
beta, log(OR), log(HR), ln(OR), ln(HR) |
nothing — already log scale |
OR, HR |
natural log |
Log2(OR) |
rescaled by log(2) |
NR or absent |
nothing, with a warning |
The reason this table matters: log(OR) is a common real value and is already logged. A converter that matched loosely on "OR" would double-log the most common non-beta case.
A score whose scale is unreported imports with its weights used as betas, and warns. If those weights are all positive with a median near 1, they are a ratio scale that was mislabelled, and the warning says so specifically — check the publication before trusting the score.
Anything that is not an effect scale at all — Dosage, Unweighted, Odds Ratio over expected risk — is refused by name. Build a model directly from a variant table if you intend to use such weights as they are.
Builds are refused, not repaired
Only GRCh37/hg19 and GRCh38/hg38 are accepted. A score reporting NR, or an assembly PolyGenius has no chain for, is refused.
That refusal is deliberate: an unplaceable build used to travel silently all the way to scoring, where it produced a score with almost no variant overlap instead of an error.
The refusal costs you that one score, not the import. generate$from.pgs.catalog() returns every score it could fetch and records the ones it could not — accession, status and reason — in provenance(models)$misc$failed.models, warning once so you know to look. The same is true of a score whose file will not download. Only a call where nothing resolved errors out, since there would be no set worth returning; that message names the reasons and the run log. generate$from.pgs.file() is stricter and stops at the first unreadable file, because skipping one would break the position and fingerprint checks a written directory is verified against.
Reading a scoring file from disk
generate$from.pgs.file() reads any PGS-format file — one written by PolyGenius, downloaded from the Catalog, or produced by a third party.
The return type follows the input shape: a single file gives one model, a directory or a vector of paths gives a library. That is invisible until you index into the result.
Harmonised and original column layouts both normalise to the same canonical columns. Local file reads are plain I/O and are not cached.
Writing models out
write.pgs.file() writes one model, or a whole set into a directory with a manifest alongside it.
mode = "extended" adds PolyGenius metadata as extension header keys, so a round trip restores the model's own generation record — the algorithm and parameters that produced it, which is a per-model thing and distinct from provenance(), the record of a call. mode = "strict" writes only the standard header and columns, which is what an external repository wants.
A file PolyGenius writes always declares weight_type=beta, because the conversion already happened at import. Writing the original scale back would make the next read convert a second time.
Two limits on "lossless": metadata that was a data frame returns as a list of rows, and any model field whose name collides with a variant column is left out of the header.
You will most likely want this section after chapter 11, when you have generated candidates and selected one worth sharing.
Two things that surprise people
Imported variants are not filtered for usable coordinates the way GWAS sources are, so unmapped rows survive into the model and show up later as variants that could not be scored.
PGS scoring files carry no p-values, so pval is NA on an imported model and any p-value threshold applied to it does nothing.