Contents
PolyGeniusStudy
Create a new PolyGeniusStudy object
Binds a PRS model library to a set of genotype filesets, producing the object every
compute$*(), associate$*(), evaluate$*() and visualize$*() verb takes as its first
argument. A study carries its samples, its models, the scores computed for them, and every
result later derived from them.
study[i, j] either subsets the study or fetches values out of it, decided by what i and j
resolve to. When both are supplied they must resolve to the same mode.
Shorthand for x$fetch(..., .context = "samples"). Resolves its arguments against the
sample-aligned planes only, and never subsets.
Resolves one sample-side and one model-side expression, validates the resulting indices, and
returns a new study. Unlike [, it never fetches: a side that resolves to data rather than to a
predicate aborts, naming x[i, j] and x$fetch() as the way to retrieve it.
Print stub for pre-rename PolyGeniusData objects; it always aborts.
Usage
PolyGeniusStudy(
name,
library,
genotypes,
samples = NULL,
sample.pairs = NULL,
uns = NULL
)S3 method for class 'PolyGeniusStudy'
x[i, j, drop = FALSE]S3 method for class 'PolyGeniusStudy'
x[[..., exact = TRUE]]S3 method for class 'PolyGeniusStudy'
subset(x, samples, models, ...)
print.PolyGeniusData(x, ...)Arguments
| Argument | Description |
|---|---|
name | Character scalar. Name for the study; must be non-empty after trimming. |
library | A PGS or a PGSLibrary, the model library the study is built against; a bare PGS is wrapped into a one-model library. Narrowed at construction into a library the study owns. |
genotypes | A GenotypeSource or a GenotypeSourceSet. Genotype filesets the study's samples come from; a bare GenotypeSource is wrapped into a one-element set. |
samples | For PolyGeniusStudy(): a table, a named list, or NULL (default), assigned as study$samples <- samples; see Assigning several values at once in the section Assigning sample-level data. For subset(): sample-side predicate or index (unquoted) -- a logical, numeric or character index that evaluates in the caller, a tidyselect selector, or an expression resolved through $fetch() to a logical vector of length n.samples. |
sample.pairs | Named list of sample-by-sample rectangular objects, or NULL (default). Each element is appended into study$sample.pairs under its own name. |
uns | Named list, or NULL (default). Each element is appended into study$uns under its own name. |
x | A PolyGeniusStudy. |
i | Unquoted sample-side index or expression(s). Missing keeps every sample in subsetting mode, and contributes nothing in fetching mode. |
j | Unquoted model-side index or expression(s). Same semantics as i, aligned to models. |
drop | Logical scalar, default FALSE. Ignored; present for [ signature compatibility. |
... | Unused. Present for S3 subset() compatibility. |
exact | Logical scalar, default TRUE. Ignored; present for [[ signature compatibility. |
models | For subset(): model-side predicate or index (unquoted), resolved as samples is but against model.names. |
Value
A new PolyGeniusStudy, already carrying study$samples$genotype -- a factor naming
each sample's genotype fileset. Aborts when name is empty, when library or genotypes is
the wrong class, when either side resolves to more than one genome build, or when the two
builds differ.
- Subsetting mode: a new
PolyGeniusStudy, with the parent unmodified. The child does not inheritassociations/evaluations/signalswhen either axis actually changed -- see Subsetting and unstructured slots insubset.PolyGeniusStudy(). - Fetching mode: that side's
PolyGeniusFramewhen only one side was supplied, or aPolyGeniusStudyFetchResult-- a list of$samplesand$modelsframes -- when both were. NULLwhen bothiandjare missing.
Whatever x$fetch(..., .context = "samples") returns: a PolyGeniusFrame, or a bare
vector when a single expression resolved without error.
A new PolyGeniusStudy, subsetted on samples and/or models. The parent is not modified.
Details
A study has two primary axes, samples and models, each with a pairwise counterpart. Every read and write lands on one named plane:
$scores$<layer>-- named sample-by-model score matrices (e.g.$scores$X), each carrying its own provenance (provenance()) and, when written bycompute$scores()/compute$standardize(), its owncoverage/transformdiagnostics (diagnostics()).$samples$<name>-- the sample axis: a primary table, table members, andreduction()-tagged embeddings, filed by the assigned value's own class. Barestudy$samplesreads the primary table.$sample.pairs$<name>-- sample-by-sample matrices, such as kinship.$models$<name>-- a per-study model-axis overlay, populated only by an explicit assignment.$model.pairs$<name>-- realized model-by-model statistics (score.pearson,score.spearman); the study-invariantsnp.*definitions live onstudy$library$backbone$model.pairs.$associations/$evaluations/$signals-- named result planes written byassociate$*(),evaluate$*()andcompute$genome$*().$uns$<name>-- free space independent of both axes.
library is narrowed into a library this study owns, reachable as study$library: a model's
variant table, weights, GWAS/algorithm annotations and snp.* model-pair definitions are all read
through it, and nothing the caller does to the library it passed in reaches the study afterwards.
Only what depends on this study's own samples -- scores, the model overlay, realized model
pairs, results -- lives on the study.
A study is an R6 object and is modified by reference: writing a plane changes the study every
binding points at. subset(), [, $subsetSamples() and $subsetModels() are the exception,
returning a new study and leaving the parent untouched.
See the Class reference section below for the field-by-field reference and what each plane validates.
A study is modified by reference, as every R6 object is: study$scores$X <- ... mutates the
study every binding points at, and no copy is made. Only subset(), [, $subsetSamples() and
$subsetModels() return a new study.
A side is in subsetting mode when its expression is a tidyselect selector, an index that
evaluates in the caller's environment (numeric, character or logical), or an expression that
x$fetch() resolves to a single aligned logical vector. Otherwise it is in fetching mode.
Subsetting samples while fetching models, or the reverse, aborts.
A single logical expression is therefore read as a predicate by default. Wrapping it in c()
forces fetching mode: study[age > 65] returns the subsetted study, study[c(age > 65)]
returns the logical vector.
Assigning sample-level data
study$samples$X <- value stores one value per study sample. Each item is placed on a sample in
one of two ways:
- By ID. Items that carry sample IDs are matched to
study$sample.names. - By order. Items without IDs are taken in
study$sample.namesorder: item i belongs to sample i.
Matching by ID works as a left join onto the study's samples. The result follows the study's
sample order. A sample that the value doesn't mention gets NA. An item whose ID is not a study
sample is dropped. A message reports how many samples matched, how many were filled with NA
and how many items were dropped.
| value | IDs come from | with no IDs, it needs |
|---|---|---|
| vector or factor | names(value) |
one item per sample |
| data.frame or matrix | rownames(value), but not R's automatic 1, 2, ... |
one row per sample |
align(table, by = <expr>) |
<expr>, evaluated using the table's columns |
not applicable |
| named list of vectors or tables | each element's own names or rownames | one item or row per sample in each element |
study$sample.names
#> [1] "S1" "S2" "S3" "S4"
# A single value is copied to every sample
study$samples$cohort <- "ERGO"
# No IDs: one value per sample, in sample order
study$samples$batch <- c(1, 1, 2, 2)
# Named: the order does not matter
study$samples$mmse <- sample(setNames(mmse$score, mmse$IID))
# Named, covering fewer samples: S2 and S4 get NA; S9 is not a study sample and is dropped
study$samples$apoe4 <- c(S3 = 1, S1 = 0, S9 = 2)
# Table with rownames, covering fewer samples: S1 and S3 get NA; S7 is dropped
pheno <- data.frame(age = c(71, 65, 80), sex = c("F", "M", "F"),
row.names = c("S4", "S2", "S7"))
study$samples$phenotypes <- pheno
# The same table with its IDs in a column
pheno <- data.frame(IID = c("S4", "S2", "S7"), age = c(71, 65, 80), sex = c("F", "M", "F"))
study$samples$phenotypes <- align(pheno, by = IID)
# IDs built from a column: the table stores 4, 2, 7; the study names are "S4", "S2", "S7"
ergo <- data.frame(ergoid = c(4, 2, 7), age = c(71, 65, 80))
study$samples$phenotypes <- align(ergo, by = paste0("S", ergoid))
# No rownames and no `by`: one row per sample is required; 3 rows for 4 samples aborts
study$samples$phenotypes <- data.frame(age = c(71, 65, 80))
# Named list: each element is matched on its own
study$samples$visits <- list(baseline = visit1, followup = visit2)
Studies with more than one genotype source
A sample ID is unique within one genotype source, because PLINK requires it. A
PolyGeniusStudy does not require IDs to be unique across sources: a participant genotyped on
two arrays appears once in each. study$sample.names lists the samples source by source, in
the order of study$genotypes$names.
Such a study accepts a value in either of two shapes:
- One value for the whole study, in any form above. By order, it needs one item per sample across all sources. By ID, each ID it carries must belong to only one source. An ID found in two sources aborts, because it could mean either sample.
- One value per source, as a list named by genotype source. Each element follows the rules above, matched against that source's samples only. By order, an element needs
study$n.samples.by.genotype[["<source>"]]items. A source left out of the list getsNAfor all its samples. A named list of values moves one level deeper.
The sources' values are combined into one value for the study. Tables take the union of their
columns: a column that one source's table lacks is NA for that source's samples. A column has
one type across sources. Integer and double count as one numeric type, and factors combine their
levels.
study$genotypes$names
#> [1] "gsa" "omni"
study$sample.names # S3 was genotyped on both arrays
#> [1] "S1" "S2" "S3" "S3" "S4"
# Aborts: S3 is in both sources
study$samples$mmse <- c(S1 = 28, S3 = 24)
# Works: every ID given is in one source only; S2 and both S3 samples get NA
study$samples$mmse <- c(S1 = 28, S4 = 30)
# Per source: S3 gets its value in both
study$samples$mmse <- list(gsa = c(S1 = 28, S3 = 24),
omni = c(S3 = 24, S4 = 30))
# Per source, from one table with a row per participant and array
study$samples$phenotypes <- lapply(split(pheno, pheno$array), \(x) align(x, by = IID))
# Only the omni samples had amyloid PET; the gsa samples get NA
study$samples$amyloid <- list(omni = setNames(pet$suvr, pet$IID))
# A named list of values, given per source
study$samples$visits <- list(gsa = list(baseline = gsa.v1, followup = gsa.v2),
omni = list(baseline = omni.v1, followup = omni.v2))
Assigning several values at once
study$samples <- x, or PolyGeniusStudy(samples = x), assigns everything x holds in one
step. x takes one of three forms:
- A table, plain or
align()-tagged. It is matched by the rules above, and its columns become columns of the primary table. - A list named by genotype source. Each source's element is either a table, whose columns become columns of the primary table, or a named list of values, each assigned as
study$samples$<name>. Each element is matched against its own source's samples only, and the sources' values are combined as above. - Any other named list. Each element is assigned as
study$samples$<name> <- element.
Every name x would write is checked before anything is written. A name that already exists in
study$samples, such as genotype, aborts. So does a name x would write twice: repeated
within one table or list, or a table column for one source and a named value for another. A
column or value given for several sources is one name, and is combined. An assignment that
aborts writes nothing.
# IID, age and sex become columns of the primary table
study$samples <- align(pheno, by = IID)
# gsa holds S1, S2, S3 and omni holds S3, S4. t1 has columns age and bmi; t2 has age and apoe.
# age takes each source's own values, bmi is NA for omni, and apoe is NA for gsa.
study$samples <- list(gsa = t1, omni = t2)
# study$samples$a is t1 with NA rows for omni; t3's columns join the primary table, NA for gsa
study$samples <- list(gsa = list(a = t1), omni = t3)
# study$samples$a holds t1's rows for gsa and t3's for omni, with the union of their columns
study$samples <- list(gsa = list(a = t1), omni = list(a = t3))
Assignment copies
study$samples is a value, like any R object: a plane read out with p <- study$samples is a
copy, and changing p leaves the study as it was until p is assigned back. The study itself is
not copied: study$samples$X <- value changes every binding to that study.
Removing a value
study$samples$X <- NULL removes X, whether it is a column of the primary table, a table
member or an embedding. Removing a name that is not there does nothing.
When an assignment aborts
- A value without IDs has the wrong number of items: the study's total, or that source's count for a per-source element.
- An ID appears twice within one value or one per-source element. Exact duplicate rows are removed first.
- A value has IDs, but none of them match the samples it is matched against. This usually means the wrong ID type, such as FID instead of IID.
- One value for the whole study carries an ID found in more than one source.
- A list mixes genotype source names with other names. This is almost always a misspelled source.
align()is given withoutby, and the value has no rownames or names.- A column has a different type in different sources' tables.
study$samples <- xwould write a name that already exists, or would write one name twice.study$samples <- xis given something other than a table or a named list, such as a vector; assign it withstudy$samples$<name> <- valueinstead.
Class reference
The R6 generator behind PolyGeniusStudy(); fields, validation, subset carry-forward below.
The sample and model axes, and the planes beside them
A study is built around two primary axes, each with a pairwise counterpart:
| axis | primary member | pairwise member | holds |
|---|---|---|---|
| sample | $samples |
$sample.pairs |
phenotypes, covariates, embeddings (PCA/UMAP); kinship, distances |
| model | $models |
$model.pairs |
per-study coverage/standardization overlay; realized score.pearson/score.spearman |
$scores sits across both axes rather than beside one: one named layer per entry (e.g.
$scores$X), n.samples rows by n.models columns. The variant axis is the third, and it lives
on the backbone rather than on the study: study$library$backbone owns the dictionary every
model's rows resolve against, and $variants$annotations there is the only per-variant overlay.
Each of these six members is validated at assignment against the study's current
n.samples/n.models, read fresh rather than cached, so a member sized against a stale count
fails at the assignment instead of at some later read. study$samples matches a value to the
samples by ID or by order, as Assigning sample-level data in PolyGeniusStudy() describes, and
then files it as a table column, a table member or an embedding by its own class. The other
members are positional: a value whose non-NULL dimnames disagree with the axis order aborts.
Score-layer column order is deliberately exempt, so compute$scores(models = <subset>) may
return a permutation of model.names.
What the backbone is
The model library a study holds is, underneath, a PolyGeniusBackbone: one variant
dictionary, one model registry and one weight store. The constructor narrows the
PGSLibrary it is handed into a library of the study's own, so nothing the caller does
to that set afterwards reaches the study, and nothing the study does reaches back. It spans
exactly one genome build,
and grows only by appending new variants and models, never shrinking or renumbering what is
already there, which is what keeps a stored weight column or a persisted result valid across
later generate$models() calls that extend the same library. A study reaches it only through
study$library$backbone; it is not one of the planes above.
Two kinds of quantity therefore never live on the study: a model's own variant table, weights
and pointer representation, and anything study-invariant the backbone computes once
(GWAS/algorithm annotations, snp.* model-pair definitions). Everything that depends on this
study's own samples does. For why the layer is split this way, see
.claude/references/objects-map.md, sections "The object graph, top to bottom" and "Why a
Results
$associations, $evaluations and $signals are named, unstructured slots holding
PolyGeniusResult objects produced by associate$*(), evaluate$*() and compute$genome$*()
respectively. Each depends on both axes, so a subset that removes any sample or any model drops
all three from the child -- see Subsetting and unstructured slots in
subset.PolyGeniusStudy(). The study axis is never a reduce axis for a signal: a pooled genome
signal is recomputed from a pooled association, never assembled by row-binding per-study signals.
Relation to other classes
- PGSLibrary / PGS -- the model library this study narrowed into one of its own at construction, reachable at
$library. - GenotypeSource / GenotypeSourceSet -- the genotype filesets this study was built from, kept at
$genotypes. - PolyGeniusResult -- the common shape of
$associations/$evaluations/$signalsentries.
Backends
savePolyGenius()/loadPolyGenius() persist a PolyGeniusStudy to a single, self-contained
.pgd file. A load reads and validates the whole payload in one pass, so a corrupt or missing
item surfaces at load time and the returned study is fully resident in memory. The file carries
the library in full -- the variant dictionary's identities, every weight column, the registry and
both column planes -- and never points into a resource store, so it can be read on a machine that
has none. Genotypes are the exception: they are stored as pointers and restored without reading
the files, so a study loads and subsets on a machine that does not hold them. See
loadPolyGenius() for moving them to a new directory.
See ?PolyGenius-backends for the on-disk layout each persisted class gets.
Subsetting and unstructured slots
Axis-aligned slots (samples, sample.pairs, scores, models, model.pairs) are sliced. The
unstructured slots cannot be, so the child inherits each one whole or not at all, by whether
every axis it depends on is intact:
| Slot | Depends on | Inherited by the child |
|---|---|---|
uns |
neither axis | always, except reserved .-prefixed keys, dropped whenever the subset actually changes either axis |
associations/evaluations/signals |
samples and models | only when nothing was removed on either axis |
A PolyGeniusAssociation or PolyGeniusEvaluation records the sample size it was fit on in
its own n column, and its adj.pval was pooled over the parent analysis's
outcome x family x stratum multiplicity family. A subset invalidates both, and filtering
the result repairs neither, so the child starts with none of them: re-run the analysis on it,
sub$associations$assoc <- associate$regression(sub, ...). Reading sub$associations$assoc
before that returns NULL, which fails at the plot rather than quietly plotting the parent's
numbers.
A uns key beginning with "." holds package-managed state whose owner is responsible for
propagating it -- a provenance ledger, say, whose honest child value is the parent's plus the
subset step, which the subset itself cannot write. Those keys are dropped; plain user keys are
inherited.
A subset that removes nothing -- subset(study), or a model side naming every model -- is a
no-op and keeps every slot. Dropped names are logged at info level.
Public fields
provenance — A PolyGeniusProvenance(). How this study was made, stamped at
construction and carried through a .pgd round trip; read it with
provenance(study). version is the package version that made the study, not the
one reading it.
Active bindings
name — Character scalar. Study name, set at construction. Read-only.
genotypes — The GenotypeSourceSet the study was built from. Read-only.
sample.names — Character vector of sample IDs, flattened across genotype filesets
in fileset order. This is the row order every sample-axis plane is checked against.
Read-only.
sample.names.by.genotype — Named list of character vectors, one per genotype
fileset. Read-only.
model.names — Character vector of model names, read through from this study's own
model library rather than captured once, so it always says what the library holds.
Read-only.
library — The study's own PGSLibrary, narrowed from the one supplied
at construction -- the single route to a model's own variant table, its
backbone-pointer representation, and its weight-store column. Read-only.
n.samples — Integer scalar. Total samples across all genotype filesets. Read-only.
n.samples.by.genotype — Integer vector of sample counts, one per genotype fileset.
Read-only.
n.models — Integer scalar, length(model.names). Read-only.
dim — Integer vector, c(n.samples, n.models). Read-only.
dim.by.genotype — List of c(n.samples, n.models) pairs, one per genotype fileset.
Read-only.
scores.keys — Character vector of score-layer names. Read-only.
scores — Named list of sample-by-model score matrices, one per layer. Assigning
replaces the whole list and validates every layer against n.samples rows, n.models
columns and sample.names row order; column order is deliberately not checked. A layer
written by compute$scores() carries a coverage diagnostics entry
(diagnostics(layer, "coverage"), one row per scored model), and one written by
compute$standardize() a transform entry (one row per model: center, scale,
scale.source). Both stay on the layer -- assigning a layer writes into no other slot.
Copy diagnostics(study$scores$X, "coverage") into study$models yourself if a
model-axis view is wanted for subset(models = ...)-style filtering.
samples.keys — Character vector of every name reachable under samples -- table
columns, then table members, then embeddings. Read-only.
sample.pairs.keys — Character vector of sample.pairs member names. Read-only.
samples — The sample-axis plane. Bare study$samples reads the primary table.
study$samples$<name> <- value first matches value to the samples, as Assigning
sample-level data in PolyGeniusStudy() describes. The result is then filed by class:
a vector becomes a table column, a data.frame, multi-column matrix or named list becomes
a table member, and a reduction()-tagged matrix becomes an embedding. A name lives in
exactly one of the three, so reassigning it with a different kind of value moves it, and
study$samples$<name> <- NULL removes it. study$samples <- x assigns every value a
table or named list x holds at once; assigning a PolyGeniusSamplesPlane of the same
height replaces the plane with a copy of it.
sample.pairs — Named list of square objects, n.samples by n.samples and checked
against sample.names on both axes (kinship, distances). Assigning replaces the whole
list.
models.keys — Character vector of per-study model-overlay member names. Read-only.
model.pairs.keys — Character vector of model.pairs member names. Read-only.
models — The per-study model-axis overlay: a named list of members with n.models
rows, checked against model.names. Nothing populates it implicitly -- a caller assigns
into it, typically to make a score layer's coverage filterable by
subset(models = ...). Study-invariant GWAS and algorithm annotations are not here;
they read through study$library$backbone$models$annotations.
model.pairs — Named list of square objects, n.models by n.models, holding
realized statistics only (score.pearson, score.spearman). The study-invariant
snp.* definitions read through study$library$backbone$model.pairs instead.
associations.keys — Character vector of associations names. Read-only.
evaluations.keys — Character vector of evaluations names. Read-only.
signals.keys — Character vector of signals names. Read-only.
associations — Named list of associate$*() results. Assigning checks only that
every entry is uniquely named. Depends on both axes, so a subset that removes any
sample or any model does not carry it to the child (see Subsetting and unstructured
slots in subset.PolyGeniusStudy()).
evaluations — Named list of evaluate$*() results. Same naming check and same
subsetting rule as associations.
signals — Named list of this study's own PolyGeniusGenomeSignal realizations.
Same naming check and same subsetting rule as associations; the study axis is never a
reduce axis for a signal.
uns.keys — Character vector of uns names, object scope first then sample scope.
Read-only.
uns — Free space, independent of both axes. Reading merges the object and sample
scopes into one named list; study$uns$<name> <- value always writes the object scope,
since no public accessor sets a per-entry scope today. A name beginning with "." is
reserved for package-managed state and is dropped by any subset that actually changes an
axis; every other name is inherited.
Methods
Public methods
PolyGeniusStudy.class$new()PolyGeniusStudy.class$print()PolyGeniusStudy.class$fetch()PolyGeniusStudy.class$subsetSamples()PolyGeniusStudy.class$subsetModels()PolyGeniusStudy.class$subsetObservations()PolyGeniusStudy.class$clone()
Method new()
Initialize a study, coupling a model library to genotype filesets. Called
by PolyGeniusStudy(), which first validates that both sides resolve to one shared
genome build; this method does not.
Usage
PolyGeniusStudy.class$new(name, library, genotypes)
Arguments
name — Character scalar. Study name.
library — A PGSLibrary. Narrowed into a library this study owns.
genotypes — A GenotypeSourceSet. Genotype filesets the samples come from.
Returns
A new PolyGeniusStudy whose samples plane holds an empty table sized to the
genotype sets and every other plane is empty.
Method print()
Print a compact summary: dimensions, genotype filesets, the member
names of every populated plane, and a closing line summarising $provenance.
Usage
PolyGeniusStudy.class$print(...)
Arguments
... — Unused. Present for S3 print() compatibility.
Details
A member name is followed by a grey superscript for each record it carries:
¹ for provenance(), ² for diagnostics(), ³ for artifacts(). A legend names the
ones in use. On a console that cannot show superscripts they read [1,2].
Returns
NULL, invisibly; called for its console output.
Method fetch()
Resolve unquoted expressions against the study's planes and return the values, packed by axis.
Usage
PolyGeniusStudy.class$fetch(
...,
.context = c("both", "samples", "models"),
.keep.nested.tables = FALSE,
.simplify = TRUE,
scores.layer = X
)
Arguments
... — Unquoted expressions to evaluate. An unnamed argument takes its de-parsed text
as the result column name.
.context — One of "both" (default), "samples", "models". Which planes are
searched, and which side(s) may be returned.
.keep.nested.tables — Logical scalar, default FALSE. Keep a fetched table as a
list-column instead of widening its inner columns into separate, name-prefixed result
columns. Only a table whose row count matches the sample or model dimension is ever
widened.
.simplify — Logical scalar, default TRUE. When exactly one column was produced and
no error was recorded, return that column as a bare vector instead of a one-column
frame.
scores.layer — Unquoted score-layer name, default X; a character scalar naming a
layer in scores.keys is also accepted. An expression that is a bare model name
resolves to that model's column in this layer, and errors if the name is duplicated.
Details
Each expression is parsed, its symbol-like nodes resolved to fully-qualified plane paths,
rewritten, evaluated, and packed by length into a sample-side and a model-side table.
uns is never searched. An unresolved or ambiguous symbol is reported on the result's
errors attribute rather than aborting.
Scope traversal. "samples" searches samples, sample.pairs and scores;
"models" searches models, model.pairs and scores; "both" searches the union.
Packing. A value of length n.samples, or a table with that many rows, goes to the
sample side; the same for n.models on the model side. A value matching neither is
recorded as an error rather than returned.
Errors. The errors attribute is a list holding any of unresolved (character
vector of symbols that matched nothing), ambiguous (named list mapping a symbol to
every matching path) and errors (named list of evaluation or packing failures, keyed by
requested column name). Printing the result summarizes them.
Duplicate model names. model.names may repeat. A bare lookup of a duplicated model
name is rejected as ambiguous; when one fetched score block unpacks to several columns
sharing a name, only the duplicated ones are disambiguated to <name>.<model.idx>, using
the model's global index in the study.
Returns
A PolyGeniusFrame when one side is populated, or a PolyGeniusFrameSet with
$samples and $models when .context = "both" and both sides are; a bare vector
when .simplify collapses a single-column result. Any of these may carry an errors
attribute.
Examples
\dontrun{
# Pull two columns; one from samples, one from models
out <- study$fetch(age, gwas$trait)
print(out) # pretty-prints, shows any ambiguity
# Restrict to sample scope (useful for predicates)
o <- study$fetch(age, sex, .context = "samples")
# Model-only fetch
m <- study$fetch(models, gwas$sample_size, .context = "models")
}
Method subsetSamples()
Subset the sample axis, returning a new study.
Usage
PolyGeniusStudy.class$subsetSamples(expr)
Arguments
expr — Unquoted sample-side predicate or index.
Details
expr is resolved first as a tidyselect selector, then by evaluating it in the
caller's environment (a numeric, character or logical index), and finally through
$fetch(.context = "samples"). A fetched value is accepted only when it is a single
aligned logical vector; anything else aborts, since this method never returns data --
use [ or $fetch() for that.
Returns
A new PolyGeniusStudy; the parent is not modified. If any sample is dropped the
child inherits no associations/evaluations/signals and no reserved .-prefixed
uns key -- see Subsetting and unstructured slots in subset.PolyGeniusStudy().
Method subsetModels()
Subset the model axis, returning a new study.
Usage
PolyGeniusStudy.class$subsetModels(expr)
Arguments
expr — Unquoted model-side predicate or index.
Details
Resolved exactly as $subsetSamples() resolves its own argument, but against
the model axis and through $fetch(.context = "models").
Returns
A new PolyGeniusStudy; the parent is not modified. If any model is dropped the
child inherits no associations/evaluations/signals and no reserved .-prefixed
uns key -- see Subsetting and unstructured slots in subset.PolyGeniusStudy().
The study's own backbone is narrowed to what the retained models reference.
Method subsetObservations()
Aborts, naming $subsetSamples() as the method to call instead.
Usage
PolyGeniusStudy.class$subsetObservations(...)
Arguments
... — Unused.
Returns
Never returns.
Method clone()
The objects of this class are cloneable with this method.
Usage
PolyGeniusStudy.class$clone(deep = FALSE)
Arguments
deep — Whether to make a deep clone.
Examples
geno <- GenotypeSource(name = "cohort", path = "data/plink",
format = "pfile", build = "GRCh38")
study <- PolyGeniusStudy(name = "cohort", library = my.models, genotypes = geno)
study$samples$phenotypes <- phenotypes # data.frame, one row per sample
study$scores$X <- compute$scores(study)
# Subsetting returns a new study; the parent is untouched.
older <- subset(study, samples = age > 65)
# The same brackets fetch values when the expression is not a predicate.
study[c(age, sex)]
```
## ------------------------------------------------
## Method `PolyGeniusStudy.class$fetch`
## ------------------------------------------------
```r
# Pull two columns; one from samples, one from models
out <- study$fetch(age, gwas$trait)
print(out) # pretty-prints, shows any ambiguity
# Restrict to sample scope (useful for predicates)
o <- study$fetch(age, sex, .context = "samples")
# Model-only fetch
m <- study$fetch(models, gwas$sample_size, .context = "models")See Also
PolyGeniusResult for the shape of $associations, $evaluations and
$signals entries.
Other study:
PolyGeniusFrame(),
PolyGeniusFrameSet(),
align(),
reduction(),
roles()