PolyGenius
Contents

PolyGeniusStudy

Create a new PolyGeniusStudy object

Binds a PRS model library to a set of genotype filesets, producing the object every compute$*(), associate$*(), evaluate$*() and visualize$*() verb takes as its first argument. A study carries its samples, its models, the scores computed for them, and every result later derived from them.

study[i, j] either subsets the study or fetches values out of it, decided by what i and j resolve to. When both are supplied they must resolve to the same mode.

Shorthand for x$fetch(..., .context = "samples"). Resolves its arguments against the sample-aligned planes only, and never subsets.

Resolves one sample-side and one model-side expression, validates the resulting indices, and returns a new study. Unlike [, it never fetches: a side that resolves to data rather than to a predicate aborts, naming x[i, j] and x$fetch() as the way to retrieve it.

Print stub for pre-rename PolyGeniusData objects; it always aborts.

Usage

PolyGeniusStudy(
  name,
  library,
  genotypes,
  samples = NULL,
  sample.pairs = NULL,
  uns = NULL
)

S3 method for class 'PolyGeniusStudy'

x[i, j, drop = FALSE]

S3 method for class 'PolyGeniusStudy'

x[[..., exact = TRUE]]

S3 method for class 'PolyGeniusStudy'

subset(x, samples, models, ...)

print.PolyGeniusData(x, ...)

Arguments

ArgumentDescription
nameCharacter scalar. Name for the study; must be non-empty after trimming.
libraryA PGS or a PGSLibrary, the model library the study is built against; a bare PGS is wrapped into a one-model library. Narrowed at construction into a library the study owns.
genotypesA GenotypeSource or a GenotypeSourceSet. Genotype filesets the study's samples come from; a bare GenotypeSource is wrapped into a one-element set.
samplesFor PolyGeniusStudy(): a table, a named list, or NULL (default), assigned as study$samples <- samples; see Assigning several values at once in the section Assigning sample-level data. For subset(): sample-side predicate or index (unquoted) -- a logical, numeric or character index that evaluates in the caller, a tidyselect selector, or an expression resolved through $fetch() to a logical vector of length n.samples.
sample.pairsNamed list of sample-by-sample rectangular objects, or NULL (default). Each element is appended into study$sample.pairs under its own name.
unsNamed list, or NULL (default). Each element is appended into study$uns under its own name.
xA PolyGeniusStudy.
iUnquoted sample-side index or expression(s). Missing keeps every sample in subsetting mode, and contributes nothing in fetching mode.
jUnquoted model-side index or expression(s). Same semantics as i, aligned to models.
dropLogical scalar, default FALSE. Ignored; present for [ signature compatibility.
...Unused. Present for S3 subset() compatibility.
exactLogical scalar, default TRUE. Ignored; present for [[ signature compatibility.
modelsFor subset(): model-side predicate or index (unquoted), resolved as samples is but against model.names.

Value

A new PolyGeniusStudy, already carrying study$samples$genotype -- a factor naming each sample's genotype fileset. Aborts when name is empty, when library or genotypes is the wrong class, when either side resolves to more than one genome build, or when the two builds differ.

  • Subsetting mode: a new PolyGeniusStudy, with the parent unmodified. The child does not inherit associations/evaluations/signals when either axis actually changed -- see Subsetting and unstructured slots in subset.PolyGeniusStudy().
  • Fetching mode: that side's PolyGeniusFrame when only one side was supplied, or a PolyGeniusStudyFetchResult -- a list of $samples and $models frames -- when both were.
  • NULL when both i and j are missing.

Whatever x$fetch(..., .context = "samples") returns: a PolyGeniusFrame, or a bare vector when a single expression resolved without error.

A new PolyGeniusStudy, subsetted on samples and/or models. The parent is not modified.

Details

A study has two primary axes, samples and models, each with a pairwise counterpart. Every read and write lands on one named plane:

  • $scores$<layer> -- named sample-by-model score matrices (e.g. $scores$X), each carrying its own provenance (provenance()) and, when written by compute$scores()/compute$standardize(), its own coverage/transform diagnostics (diagnostics()).
  • $samples$<name> -- the sample axis: a primary table, table members, and reduction()-tagged embeddings, filed by the assigned value's own class. Bare study$samples reads the primary table.
  • $sample.pairs$<name> -- sample-by-sample matrices, such as kinship.
  • $models$<name> -- a per-study model-axis overlay, populated only by an explicit assignment.
  • $model.pairs$<name> -- realized model-by-model statistics (score.pearson, score.spearman); the study-invariant snp.* definitions live on study$library$backbone$model.pairs.
  • $associations / $evaluations / $signals -- named result planes written by associate$*(), evaluate$*() and compute$genome$*().
  • $uns$<name> -- free space independent of both axes.

library is narrowed into a library this study owns, reachable as study$library: a model's variant table, weights, GWAS/algorithm annotations and snp.* model-pair definitions are all read through it, and nothing the caller does to the library it passed in reaches the study afterwards. Only what depends on this study's own samples -- scores, the model overlay, realized model pairs, results -- lives on the study.

A study is an R6 object and is modified by reference: writing a plane changes the study every binding points at. subset(), [, $subsetSamples() and $subsetModels() are the exception, returning a new study and leaving the parent untouched.

See the Class reference section below for the field-by-field reference and what each plane validates.

A study is modified by reference, as every R6 object is: study$scores$X <- ... mutates the study every binding points at, and no copy is made. Only subset(), [, $subsetSamples() and $subsetModels() return a new study.

A side is in subsetting mode when its expression is a tidyselect selector, an index that evaluates in the caller's environment (numeric, character or logical), or an expression that x$fetch() resolves to a single aligned logical vector. Otherwise it is in fetching mode. Subsetting samples while fetching models, or the reverse, aborts.

A single logical expression is therefore read as a predicate by default. Wrapping it in c() forces fetching mode: study[age > 65] returns the subsetted study, study[c(age > 65)] returns the logical vector.

Assigning sample-level data

study$samples$X <- value stores one value per study sample. Each item is placed on a sample in one of two ways:

  • By ID. Items that carry sample IDs are matched to study$sample.names.
  • By order. Items without IDs are taken in study$sample.names order: item i belongs to sample i.

Matching by ID works as a left join onto the study's samples. The result follows the study's sample order. A sample that the value doesn't mention gets NA. An item whose ID is not a study sample is dropped. A message reports how many samples matched, how many were filled with NA and how many items were dropped.

value IDs come from with no IDs, it needs
vector or factor names(value) one item per sample
data.frame or matrix rownames(value), but not R's automatic 1, 2, ... one row per sample
align(table, by = <expr>) <expr>, evaluated using the table's columns not applicable
named list of vectors or tables each element's own names or rownames one item or row per sample in each element
study$sample.names
#> [1] "S1" "S2" "S3" "S4"

# A single value is copied to every sample
study$samples$cohort <- "ERGO"

# No IDs: one value per sample, in sample order
study$samples$batch <- c(1, 1, 2, 2)

# Named: the order does not matter
study$samples$mmse <- sample(setNames(mmse$score, mmse$IID))

# Named, covering fewer samples: S2 and S4 get NA; S9 is not a study sample and is dropped
study$samples$apoe4 <- c(S3 = 1, S1 = 0, S9 = 2)

# Table with rownames, covering fewer samples: S1 and S3 get NA; S7 is dropped
pheno <- data.frame(age = c(71, 65, 80), sex = c("F", "M", "F"),
                    row.names = c("S4", "S2", "S7"))
study$samples$phenotypes <- pheno

# The same table with its IDs in a column
pheno <- data.frame(IID = c("S4", "S2", "S7"), age = c(71, 65, 80), sex = c("F", "M", "F"))
study$samples$phenotypes <- align(pheno, by = IID)

# IDs built from a column: the table stores 4, 2, 7; the study names are "S4", "S2", "S7"
ergo <- data.frame(ergoid = c(4, 2, 7), age = c(71, 65, 80))
study$samples$phenotypes <- align(ergo, by = paste0("S", ergoid))

# No rownames and no `by`: one row per sample is required; 3 rows for 4 samples aborts
study$samples$phenotypes <- data.frame(age = c(71, 65, 80))

# Named list: each element is matched on its own
study$samples$visits <- list(baseline = visit1, followup = visit2)

Studies with more than one genotype source

A sample ID is unique within one genotype source, because PLINK requires it. A PolyGeniusStudy does not require IDs to be unique across sources: a participant genotyped on two arrays appears once in each. study$sample.names lists the samples source by source, in the order of study$genotypes$names.

Such a study accepts a value in either of two shapes:

  • One value for the whole study, in any form above. By order, it needs one item per sample across all sources. By ID, each ID it carries must belong to only one source. An ID found in two sources aborts, because it could mean either sample.
  • One value per source, as a list named by genotype source. Each element follows the rules above, matched against that source's samples only. By order, an element needs study$n.samples.by.genotype[["<source>"]] items. A source left out of the list gets NA for all its samples. A named list of values moves one level deeper.

The sources' values are combined into one value for the study. Tables take the union of their columns: a column that one source's table lacks is NA for that source's samples. A column has one type across sources. Integer and double count as one numeric type, and factors combine their levels.

study$genotypes$names
#> [1] "gsa"  "omni"
study$sample.names                     # S3 was genotyped on both arrays
#> [1] "S1" "S2" "S3" "S3" "S4"

# Aborts: S3 is in both sources
study$samples$mmse <- c(S1 = 28, S3 = 24)

# Works: every ID given is in one source only; S2 and both S3 samples get NA
study$samples$mmse <- c(S1 = 28, S4 = 30)

# Per source: S3 gets its value in both
study$samples$mmse <- list(gsa  = c(S1 = 28, S3 = 24),
                           omni = c(S3 = 24, S4 = 30))

# Per source, from one table with a row per participant and array
study$samples$phenotypes <- lapply(split(pheno, pheno$array), \(x) align(x, by = IID))

# Only the omni samples had amyloid PET; the gsa samples get NA
study$samples$amyloid <- list(omni = setNames(pet$suvr, pet$IID))

# A named list of values, given per source
study$samples$visits <- list(gsa  = list(baseline = gsa.v1,  followup = gsa.v2),
                             omni = list(baseline = omni.v1, followup = omni.v2))

Assigning several values at once

study$samples <- x, or PolyGeniusStudy(samples = x), assigns everything x holds in one step. x takes one of three forms:

  • A table, plain or align()-tagged. It is matched by the rules above, and its columns become columns of the primary table.
  • A list named by genotype source. Each source's element is either a table, whose columns become columns of the primary table, or a named list of values, each assigned as study$samples$<name>. Each element is matched against its own source's samples only, and the sources' values are combined as above.
  • Any other named list. Each element is assigned as study$samples$<name> <- element.

Every name x would write is checked before anything is written. A name that already exists in study$samples, such as genotype, aborts. So does a name x would write twice: repeated within one table or list, or a table column for one source and a named value for another. A column or value given for several sources is one name, and is combined. An assignment that aborts writes nothing.

# IID, age and sex become columns of the primary table
study$samples <- align(pheno, by = IID)

# gsa holds S1, S2, S3 and omni holds S3, S4. t1 has columns age and bmi; t2 has age and apoe.
# age takes each source's own values, bmi is NA for omni, and apoe is NA for gsa.
study$samples <- list(gsa = t1, omni = t2)

# study$samples$a is t1 with NA rows for omni; t3's columns join the primary table, NA for gsa
study$samples <- list(gsa = list(a = t1), omni = t3)

# study$samples$a holds t1's rows for gsa and t3's for omni, with the union of their columns
study$samples <- list(gsa = list(a = t1), omni = list(a = t3))

Assignment copies

study$samples is a value, like any R object: a plane read out with p <- study$samples is a copy, and changing p leaves the study as it was until p is assigned back. The study itself is not copied: study$samples$X <- value changes every binding to that study.

Removing a value

study$samples$X <- NULL removes X, whether it is a column of the primary table, a table member or an embedding. Removing a name that is not there does nothing.

When an assignment aborts

  • A value without IDs has the wrong number of items: the study's total, or that source's count for a per-source element.
  • An ID appears twice within one value or one per-source element. Exact duplicate rows are removed first.
  • A value has IDs, but none of them match the samples it is matched against. This usually means the wrong ID type, such as FID instead of IID.
  • One value for the whole study carries an ID found in more than one source.
  • A list mixes genotype source names with other names. This is almost always a misspelled source.
  • align() is given without by, and the value has no rownames or names.
  • A column has a different type in different sources' tables.
  • study$samples <- x would write a name that already exists, or would write one name twice.
  • study$samples <- x is given something other than a table or a named list, such as a vector; assign it with study$samples$<name> <- value instead.

Class reference

The R6 generator behind PolyGeniusStudy(); fields, validation, subset carry-forward below.

The sample and model axes, and the planes beside them

A study is built around two primary axes, each with a pairwise counterpart:

axis primary member pairwise member holds
sample $samples $sample.pairs phenotypes, covariates, embeddings (PCA/UMAP); kinship, distances
model $models $model.pairs per-study coverage/standardization overlay; realized score.pearson/score.spearman

$scores sits across both axes rather than beside one: one named layer per entry (e.g. $scores$X), n.samples rows by n.models columns. The variant axis is the third, and it lives on the backbone rather than on the study: study$library$backbone owns the dictionary every model's rows resolve against, and $variants$annotations there is the only per-variant overlay.

Each of these six members is validated at assignment against the study's current n.samples/n.models, read fresh rather than cached, so a member sized against a stale count fails at the assignment instead of at some later read. study$samples matches a value to the samples by ID or by order, as Assigning sample-level data in PolyGeniusStudy() describes, and then files it as a table column, a table member or an embedding by its own class. The other members are positional: a value whose non-NULL dimnames disagree with the axis order aborts. Score-layer column order is deliberately exempt, so compute$scores(models = <subset>) may return a permutation of model.names.

What the backbone is

The model library a study holds is, underneath, a PolyGeniusBackbone: one variant dictionary, one model registry and one weight store. The constructor narrows the PGSLibrary it is handed into a library of the study's own, so nothing the caller does to that set afterwards reaches the study, and nothing the study does reaches back. It spans exactly one genome build, and grows only by appending new variants and models, never shrinking or renumbering what is already there, which is what keeps a stored weight column or a persisted result valid across later generate$models() calls that extend the same library. A study reaches it only through study$library$backbone; it is not one of the planes above.

Two kinds of quantity therefore never live on the study: a model's own variant table, weights and pointer representation, and anything study-invariant the backbone computes once (GWAS/algorithm annotations, snp.* model-pair definitions). Everything that depends on this study's own samples does. For why the layer is split this way, see .claude/references/objects-map.md, sections "The object graph, top to bottom" and "Why a

Results

$associations, $evaluations and $signals are named, unstructured slots holding PolyGeniusResult objects produced by associate$*(), evaluate$*() and compute$genome$*() respectively. Each depends on both axes, so a subset that removes any sample or any model drops all three from the child -- see Subsetting and unstructured slots in subset.PolyGeniusStudy(). The study axis is never a reduce axis for a signal: a pooled genome signal is recomputed from a pooled association, never assembled by row-binding per-study signals.

Relation to other classes

  • PGSLibrary / PGS -- the model library this study narrowed into one of its own at construction, reachable at $library.
  • GenotypeSource / GenotypeSourceSet -- the genotype filesets this study was built from, kept at $genotypes.
  • PolyGeniusResult -- the common shape of $associations/$evaluations/$signals entries.

Backends

savePolyGenius()/loadPolyGenius() persist a PolyGeniusStudy to a single, self-contained .pgd file. A load reads and validates the whole payload in one pass, so a corrupt or missing item surfaces at load time and the returned study is fully resident in memory. The file carries the library in full -- the variant dictionary's identities, every weight column, the registry and both column planes -- and never points into a resource store, so it can be read on a machine that has none. Genotypes are the exception: they are stored as pointers and restored without reading the files, so a study loads and subsets on a machine that does not hold them. See loadPolyGenius() for moving them to a new directory.

See ?PolyGenius-backends for the on-disk layout each persisted class gets.

Subsetting and unstructured slots

Axis-aligned slots (samples, sample.pairs, scores, models, model.pairs) are sliced. The unstructured slots cannot be, so the child inherits each one whole or not at all, by whether every axis it depends on is intact:

Slot Depends on Inherited by the child
uns neither axis always, except reserved .-prefixed keys, dropped whenever the subset actually changes either axis
associations/evaluations/signals samples and models only when nothing was removed on either axis

A PolyGeniusAssociation or PolyGeniusEvaluation records the sample size it was fit on in its own n column, and its adj.pval was pooled over the parent analysis's outcome x family x stratum multiplicity family. A subset invalidates both, and filtering the result repairs neither, so the child starts with none of them: re-run the analysis on it, sub$associations$assoc <- associate$regression(sub, ...). Reading sub$associations$assoc before that returns NULL, which fails at the plot rather than quietly plotting the parent's numbers.

A uns key beginning with "." holds package-managed state whose owner is responsible for propagating it -- a provenance ledger, say, whose honest child value is the parent's plus the subset step, which the subset itself cannot write. Those keys are dropped; plain user keys are inherited.

A subset that removes nothing -- subset(study), or a model side naming every model -- is a no-op and keeps every slot. Dropped names are logged at info level.

Public fields

provenance — A PolyGeniusProvenance(). How this study was made, stamped at construction and carried through a .pgd round trip; read it with provenance(study). version is the package version that made the study, not the one reading it.

Active bindings

name — Character scalar. Study name, set at construction. Read-only.

genotypes — The GenotypeSourceSet the study was built from. Read-only.

sample.names — Character vector of sample IDs, flattened across genotype filesets in fileset order. This is the row order every sample-axis plane is checked against. Read-only.

sample.names.by.genotype — Named list of character vectors, one per genotype fileset. Read-only.

model.names — Character vector of model names, read through from this study's own model library rather than captured once, so it always says what the library holds. Read-only.

library — The study's own PGSLibrary, narrowed from the one supplied at construction -- the single route to a model's own variant table, its backbone-pointer representation, and its weight-store column. Read-only.

n.samples — Integer scalar. Total samples across all genotype filesets. Read-only.

n.samples.by.genotype — Integer vector of sample counts, one per genotype fileset. Read-only.

n.models — Integer scalar, length(model.names). Read-only.

dim — Integer vector, c(n.samples, n.models). Read-only.

dim.by.genotype — List of c(n.samples, n.models) pairs, one per genotype fileset. Read-only.

scores.keys — Character vector of score-layer names. Read-only.

scores — Named list of sample-by-model score matrices, one per layer. Assigning replaces the whole list and validates every layer against n.samples rows, n.models columns and sample.names row order; column order is deliberately not checked. A layer written by compute$scores() carries a coverage diagnostics entry (diagnostics(layer, "coverage"), one row per scored model), and one written by compute$standardize() a transform entry (one row per model: center, scale, scale.source). Both stay on the layer -- assigning a layer writes into no other slot. Copy diagnostics(study$scores$X, "coverage") into study$models yourself if a model-axis view is wanted for subset(models = ...)-style filtering.

samples.keys — Character vector of every name reachable under samples -- table columns, then table members, then embeddings. Read-only.

sample.pairs.keys — Character vector of sample.pairs member names. Read-only.

samples — The sample-axis plane. Bare study$samples reads the primary table. study$samples$<name> <- value first matches value to the samples, as Assigning sample-level data in PolyGeniusStudy() describes. The result is then filed by class: a vector becomes a table column, a data.frame, multi-column matrix or named list becomes a table member, and a reduction()-tagged matrix becomes an embedding. A name lives in exactly one of the three, so reassigning it with a different kind of value moves it, and study$samples$<name> <- NULL removes it. study$samples <- x assigns every value a table or named list x holds at once; assigning a PolyGeniusSamplesPlane of the same height replaces the plane with a copy of it.

sample.pairs — Named list of square objects, n.samples by n.samples and checked against sample.names on both axes (kinship, distances). Assigning replaces the whole list.

models.keys — Character vector of per-study model-overlay member names. Read-only.

model.pairs.keys — Character vector of model.pairs member names. Read-only.

models — The per-study model-axis overlay: a named list of members with n.models rows, checked against model.names. Nothing populates it implicitly -- a caller assigns into it, typically to make a score layer's coverage filterable by subset(models = ...). Study-invariant GWAS and algorithm annotations are not here; they read through study$library$backbone$models$annotations.

model.pairs — Named list of square objects, n.models by n.models, holding realized statistics only (score.pearson, score.spearman). The study-invariant snp.* definitions read through study$library$backbone$model.pairs instead.

associations.keys — Character vector of associations names. Read-only.

evaluations.keys — Character vector of evaluations names. Read-only.

signals.keys — Character vector of signals names. Read-only.

associations — Named list of associate$*() results. Assigning checks only that every entry is uniquely named. Depends on both axes, so a subset that removes any sample or any model does not carry it to the child (see Subsetting and unstructured slots in subset.PolyGeniusStudy()).

evaluations — Named list of evaluate$*() results. Same naming check and same subsetting rule as associations.

signals — Named list of this study's own PolyGeniusGenomeSignal realizations. Same naming check and same subsetting rule as associations; the study axis is never a reduce axis for a signal.

uns.keys — Character vector of uns names, object scope first then sample scope. Read-only.

uns — Free space, independent of both axes. Reading merges the object and sample scopes into one named list; study$uns$<name> <- value always writes the object scope, since no public accessor sets a per-entry scope today. A name beginning with "." is reserved for package-managed state and is dropped by any subset that actually changes an axis; every other name is inherited.

Methods

Public methods

  • PolyGeniusStudy.class$new()
  • PolyGeniusStudy.class$print()
  • PolyGeniusStudy.class$fetch()
  • PolyGeniusStudy.class$subsetSamples()
  • PolyGeniusStudy.class$subsetModels()
  • PolyGeniusStudy.class$subsetObservations()
  • PolyGeniusStudy.class$clone()

Method new()

Initialize a study, coupling a model library to genotype filesets. Called by PolyGeniusStudy(), which first validates that both sides resolve to one shared genome build; this method does not.

Usage

PolyGeniusStudy.class$new(name, library, genotypes)

Arguments

name — Character scalar. Study name.

library — A PGSLibrary. Narrowed into a library this study owns.

genotypes — A GenotypeSourceSet. Genotype filesets the samples come from.

Returns

A new PolyGeniusStudy whose samples plane holds an empty table sized to the genotype sets and every other plane is empty.

Method print()

Print a compact summary: dimensions, genotype filesets, the member names of every populated plane, and a closing line summarising $provenance.

Usage

PolyGeniusStudy.class$print(...)

Arguments

... — Unused. Present for S3 print() compatibility.

Details

A member name is followed by a grey superscript for each record it carries: ¹ for provenance(), ² for diagnostics(), ³ for artifacts(). A legend names the ones in use. On a console that cannot show superscripts they read [1,2].

Returns

NULL, invisibly; called for its console output.

Method fetch()

Resolve unquoted expressions against the study's planes and return the values, packed by axis.

Usage

PolyGeniusStudy.class$fetch(
  ...,
  .context = c("both", "samples", "models"),
  .keep.nested.tables = FALSE,
  .simplify = TRUE,
  scores.layer = X
)

Arguments

... — Unquoted expressions to evaluate. An unnamed argument takes its de-parsed text as the result column name.

.context — One of "both" (default), "samples", "models". Which planes are searched, and which side(s) may be returned.

.keep.nested.tables — Logical scalar, default FALSE. Keep a fetched table as a list-column instead of widening its inner columns into separate, name-prefixed result columns. Only a table whose row count matches the sample or model dimension is ever widened.

.simplify — Logical scalar, default TRUE. When exactly one column was produced and no error was recorded, return that column as a bare vector instead of a one-column frame.

scores.layer — Unquoted score-layer name, default X; a character scalar naming a layer in scores.keys is also accepted. An expression that is a bare model name resolves to that model's column in this layer, and errors if the name is duplicated.

Details

Each expression is parsed, its symbol-like nodes resolved to fully-qualified plane paths, rewritten, evaluated, and packed by length into a sample-side and a model-side table. uns is never searched. An unresolved or ambiguous symbol is reported on the result's errors attribute rather than aborting.

Scope traversal. "samples" searches samples, sample.pairs and scores; "models" searches models, model.pairs and scores; "both" searches the union.

Packing. A value of length n.samples, or a table with that many rows, goes to the sample side; the same for n.models on the model side. A value matching neither is recorded as an error rather than returned.

Errors. The errors attribute is a list holding any of unresolved (character vector of symbols that matched nothing), ambiguous (named list mapping a symbol to every matching path) and errors (named list of evaluation or packing failures, keyed by requested column name). Printing the result summarizes them.

Duplicate model names. model.names may repeat. A bare lookup of a duplicated model name is rejected as ambiguous; when one fetched score block unpacks to several columns sharing a name, only the duplicated ones are disambiguated to <name>.<model.idx>, using the model's global index in the study.

Returns

A PolyGeniusFrame when one side is populated, or a PolyGeniusFrameSet with $samples and $models when .context = "both" and both sides are; a bare vector when .simplify collapses a single-column result. Any of these may carry an errors attribute.

Examples

\dontrun{
# Pull two columns; one from samples, one from models
out <- study$fetch(age, gwas$trait)
print(out)           # pretty-prints, shows any ambiguity

# Restrict to sample scope (useful for predicates)
o <- study$fetch(age, sex, .context = "samples")

# Model-only fetch
m <- study$fetch(models, gwas$sample_size, .context = "models")
}


Method subsetSamples()

Subset the sample axis, returning a new study.

Usage

PolyGeniusStudy.class$subsetSamples(expr)

Arguments

expr — Unquoted sample-side predicate or index.

Details

expr is resolved first as a tidyselect selector, then by evaluating it in the caller's environment (a numeric, character or logical index), and finally through $fetch(.context = "samples"). A fetched value is accepted only when it is a single aligned logical vector; anything else aborts, since this method never returns data -- use [ or $fetch() for that.

Returns

A new PolyGeniusStudy; the parent is not modified. If any sample is dropped the child inherits no associations/evaluations/signals and no reserved .-prefixed uns key -- see Subsetting and unstructured slots in subset.PolyGeniusStudy().

Method subsetModels()

Subset the model axis, returning a new study.

Usage

PolyGeniusStudy.class$subsetModels(expr)

Arguments

expr — Unquoted model-side predicate or index.

Details

Resolved exactly as $subsetSamples() resolves its own argument, but against the model axis and through $fetch(.context = "models").

Returns

A new PolyGeniusStudy; the parent is not modified. If any model is dropped the child inherits no associations/evaluations/signals and no reserved .-prefixed uns key -- see Subsetting and unstructured slots in subset.PolyGeniusStudy(). The study's own backbone is narrowed to what the retained models reference.

Method subsetObservations()

Aborts, naming $subsetSamples() as the method to call instead.

Usage

PolyGeniusStudy.class$subsetObservations(...)

Arguments

... — Unused.

Returns

Never returns.

Method clone()

The objects of this class are cloneable with this method.

Usage

PolyGeniusStudy.class$clone(deep = FALSE)

Arguments

deep — Whether to make a deep clone.

Examples

geno  <- GenotypeSource(name = "cohort", path = "data/plink",
                        format = "pfile", build = "GRCh38")
study <- PolyGeniusStudy(name = "cohort", library = my.models, genotypes = geno)

study$samples$phenotypes <- phenotypes      # data.frame, one row per sample
study$scores$X <- compute$scores(study)

# Subsetting returns a new study; the parent is untouched.
older <- subset(study, samples = age > 65)

# The same brackets fetch values when the expression is not a predicate.
study[c(age, sex)]

```


## ------------------------------------------------
## Method `PolyGeniusStudy.class$fetch`
## ------------------------------------------------


```r

# Pull two columns; one from samples, one from models
out <- study$fetch(age, gwas$trait)
print(out)           # pretty-prints, shows any ambiguity

# Restrict to sample scope (useful for predicates)
o <- study$fetch(age, sex, .context = "samples")

# Model-only fetch
m <- study$fetch(models, gwas$sample_size, .context = "models")

See Also

PolyGeniusResult for the shape of $associations, $evaluations and $signals entries.

Other study: PolyGeniusFrame(), PolyGeniusFrameSet(), align(), reduction(), roles()

Aliases: PolyGeniusStudy, PolyGeniusStudy.class, [.PolyGeniusStudy, [[.PolyGeniusStudy, subset.PolyGeniusStudy, print.PolyGeniusData