Persistence and the Backbone
What a .pgd file physically is, why every object is rebuilt through its own constructor, and what a round trip does and does not keep
Your First Analysis ends by calling savePolyGenius() on a
finished study and leaves the mechanism unexplained on purpose — you do not need it to
use the function. This chapter is for the moment you do: a file is larger than you
expected, a load is slower than it should be, or you need to know whether something
you put on a study will still be there when you read it back.
What a .pgd file physically is
A .pgd file is one saveRDS() of one plain list. There is no zip container, no
manifest file, no database, and no dependency on workspace$.internal$store. A .pgd
file is a user-facing export: it is self-contained, it is portable to any machine with
the package installed, and nothing about a loaded object points back at the file it
came from.
The payload carries four fields beside the data itself — a format integer, the
class it holds, the package.version that wrote it, and a created timestamp.
loadPolyGenius() dispatches on class, and five are recognised: PolyGeniusStudy,
PGSLibrary, PGS, PolyGeniusAssociation and PolyGeniusEvaluation. One
verb pair covers all five — there is no savePolyGeniusAssociation() alongside
savePolyGeniusStudy() — because the payload, not the function name, is what tells a
load which reconstruction path to take.
loadPolyGenius() checks four things in order, and the order matters: the payload has
the shape of a .pgd at all, its class is one of the five, its format integer is
exactly the one this version reads, and only then does the reconstruction run. The
class check comes before the format check, so a file holding something this package
does not know is refused by name rather than by version. The format check is exact equality, not a
range — a mismatch names both versions and stops there.
That covers a file whose class or format changed. It does not cover every unreadable
file: a payload whose class and format are both still current but whose contents moved
— a renamed key inside the payload, say — passes both checks and fails inside the
reconstruction. Those cases are guarded individually where the reader would otherwise
read NULL, with a message naming the cause; the two checks above are the cheap outer
gate, not the whole of the validation.
Saving and loading
savePolyGenius(data, "cohort-analysis.pgd")
data2 <- loadPolyGenius("cohort-analysis.pgd")Every class is reconstructed through its own real constructor on load — never handed
back as a bare readRDS() result. This is the reason the format is a decomposed
payload rather than a serialized object, and it is not a stylistic preference.
PGS and PGSLibrary are R6 objects, and saveRDS() on a live R6 object
writes each method's body verbatim while recording the enclosing namespace by name.
A later readRDS() therefore reunites the old bodies with the new namespace:
rename an internal helper between save and load and the load becomes a bare "could not
find function" error, deep inside your own analysis rather than at the load call. It
takes two package versions to reproduce, so no single-session test can catch it.
A model library is therefore decomposed on save into plain tables — the variant dictionary's identities, each model's weight column, the registry, the column planes — and rebuilt through the backbone's own constructor on load, which re-derives the variant-key codec, the packed dictionary and the slot indices from the identities rather than trusting any of them from the file. A guard refuses to write a payload with a function or environment reachable from it at all, so this is enforced rather than merely intended.
What a round trip keeps
A model library round-trips in full: identities, weights, the registry, both column
planes, attached annotation sources, and the study-invariant model-pair definitions on
the backbone. A library's own name, provenance record and metadata fields come back
verbatim — its $provenance is a statement about how it was made, and surviving a
round trip is the point of recording it.
One thing is deliberately not kept, because a rebuild cannot honestly claim it: a
dictionary row's eaf/eaf.source. A rebuilt dictionary starts these unpopulated,
exactly as a freshly absorbed one does.
A study's own provenance record is kept, and this is the part worth knowing. It is
carried in the payload and reattached after the object is rebuilt, so
provenance(study)$version reports the version that made the study, not the one
reading it. The payload's own package.version field separately records what wrote the
file — a different question. A .pgd written before that record existed is still read
rather than refused; such a study reports whatever the load stamped, so treat a version
on an old file as the reading session's, not the writing one's.
Genotypes are never embedded. A study's GenotypeSource is saved as a pointer — name,
path, files, format, build and its sample table — and a loaded study reads genotypes from
wherever those paths point at load time. Move a .pgd file without moving the genotype
files it points to and everything that does not need genotypes still loads and works;
anything that would touch genotypes again does not.
The pointer is restored without reading the files, so neither a load nor a subset() needs
the genotype files or PLINK. A call that reads genotypes, such as compute$scores(), fails
when it is called. When the files have moved, give each fileset its new directory:
study <- loadPolyGenius("cohort1.pgd", genotypes = c(cohort = "/new/dir"))The fileset is rebuilt there from the same file stems, and PLINK reads its sample list. The load aborts unless those files hold exactly the recorded samples, in the recorded order, because the study's sample axis and samples table are aligned to that order.
Narrowing: a subset saves as a subset
A model-subset study — one built from subset(study, models = ...) or study[, j] —
does not carry its parent's full variant dictionary. A save writes exactly the variant
identities its retained models reference, so the loaded dictionary holds no row those
models do not use. A "full" study gets the same treatment, which also clears rows a
model's own filter.variants() left orphaned.
This is why a loaded object's $fingerprint is always a genuine computation over what
was actually saved, never a stale copy of a larger backbone's value.
What a load reads
loadPolyGenius() reads and validates the whole payload before it returns. A corrupt
or truncated file, a payload whose model registry names a genome build it carries no
backbone for, or one whose registry points past the end of its own weight store, all
fail at the load call rather than later on whatever line first touches the damage. The
returned object is fully resident in memory.
What this does not do
There is no cross-version migration. A .pgd file written in an older format version
fails to load with a clear message, and nothing in this package rewrites an old file
into the current format for you — the documented remedy is to regenerate it from source
with the current package version. This follows the project's stated priority:
forward-looking design over backwards compatibility, while PolyGenius is not yet in
production.
A .pgd file is not a cache. It is never read or written through the resource store's
SQLite index, so nothing about resource reuse, freshness checking or automatic
invalidation applies to it. You manage .pgd files yourself, the same way you would
manage any other file you hand to a collaborator.
Where to go next
Resources and catalogs covers the session-wide resource store, which is a separate mechanism from this one. The execution engine covers the block-store contract that governs weight access inside a live session, and how the engine schedules the work a loaded study still needs to do.