PolyGenius
Contents

GenotypeSource

Create a genotype information object

Points PolyGenius at a genotype fileset on disk and gives it a name, format and genome build. A GenotypeSource holds only that pointer plus the fileset's sample identifiers: it never loads genotypes into memory, and every operation on it shells out to PLINK2 and returns file paths or small summary tables.

Usage

GenotypeSource(
  name,
  path,
  files = NULL,
  files.pattern = ".*",
  format = c("vcf", "bfile", "pfile"),
  build,
  samples = NULL,
  logger = NULL,
  plink = NULL,
  nthreads = NULL,
  ...
)

Arguments

ArgumentDescription
nameCharacter scalar. Non-whitespace label for the source, used in messages and as the default output fileset stem.
pathCharacter scalar. Existing directory holding the genotype files.
filesCharacter vector or NULL (default). For "bfile" and "pfile", fileset stems without extension; for "vcf", full filenames relative to path. NULL discovers files with files.pattern.
files.patternCharacter scalar or NULL, default ".*". Perl regular expression matched against filenames when files is NULL.
formatOne of "vcf" (default), "bfile", "pfile". Layout of the fileset on disk.
buildCharacter scalar. Genome-build key or name resolvable by workspace$catalogs$genomeBuilds; see its $view() for the accepted values.
samplesCharacter vector of IIDs, or NULL (default). Initial working set; NULL includes every sample found. IIDs absent from the fileset are dropped silently.
loggerOptional PolyGeniusLogger, default NULL. Scopes the messages emitted during sample discovery; NULL creates one.
plinkCharacter scalar or NULL (default). Path to the PLINK2 executable; NULL resolves it through workspace$setup$get("plink").
nthreadsPositive integer scalar or NULL (default). PLINK --threads count used during sample discovery.
...Further named values, assigned as public fields on the returned object.

Value

A GenotypeSource R6 object. Aborts when name is blank, when path does not exist, when both files and files.pattern are NULL, when build is not supported, when any file required by format is missing, or when the fileset yields no samples or duplicate IIDs.

Details

Construction inspects the fileset. Each file is read once with PLINK2 --write-samples to build the FID/IID table behind $samples; no genotype data is read. Construction therefore costs one PLINK call per file, and needs a resolvable PLINK2 binary.

Class

The R6 generator behind GenotypeSource(), a facade over internal services that own shared state, sample identity, PLINK execution, artifact staging, dosage and frequency extraction, variant operations and PRS scoring. The object stores a path, a file list, a format, a build and a sample table, never genotypes.

Methods that transform the fileset ($tidy(), $lift(), $merge(), $variants.*(), $samples.extract(), $samples.exclude()) write new files with PLINK2 and return a new GenotypeSource pointing at them. The receiver is left unchanged, except for $variants.updateIDs(), which rewrites this object's own files in place. Methods that read ($dosages(), $variants.LD(), $variants.frequencies(), $variants.associations()) return file paths by default and only materialise a table when asked with load = TRUE.

Active bindings

name — Character scalar, the source's name. Read-only.

path — Character scalar, the directory holding the genotype files. Read-only.

files — Character vector of fileset stems ("bfile"/"pfile") or VCF filenames, relative to $path. Read-only.

format — Character scalar, one of "vcf", "bfile", "pfile". Read-only.

build — Character scalar, the resolved genome-build name. Read-only.

samples — Character vector of the working set's sample IIDs. Read/write.

samples.full — Character vector of every sample IID in the fileset, ignoring the working set. Read-only.

n.samples — Integer scalar, size of the current working set. Read-only.

n.samples.full — Integer scalar, size of the full fileset. Read-only.

n.variants — Integer scalar, variants across every file. Read-only, and computed on each access by reading the variant metadata of every file.

Methods

Public methods

  • GenotypeSource.class$new()
  • GenotypeSource.class$print()
  • GenotypeSource.class$dosages()
  • GenotypeSource.class$tidy()
  • GenotypeSource.class$lift()
  • GenotypeSource.class$merge()
  • GenotypeSource.class$clump()
  • GenotypeSource.class$samples.extract()
  • GenotypeSource.class$samples.exclude()
  • GenotypeSource.class$samples.kinship()
  • GenotypeSource.class$samples.PCA()
  • GenotypeSource.class$samples.PCAProject()
  • GenotypeSource.class$variant.IDs()
  • GenotypeSource.class$variants()
  • GenotypeSource.class$variants.exclude()
  • GenotypeSource.class$variants.extract()
  • GenotypeSource.class$variants.filterMAF()
  • GenotypeSource.class$variants.updateIDs()
  • GenotypeSource.class$variants.LD()
  • GenotypeSource.class$variants.associations()
  • GenotypeSource.class$variants.frequencies()
  • GenotypeSource.class$scores()
  • GenotypeSource.class$sample.table()
  • GenotypeSource.class$clone()

Method new()

Setting restricts the working set to the IIDs given; those absent from the fileset are dropped silently. Setting NULL restores every sample. Only the in-memory working set changes -- the files on disk are untouched. Aborts on a value that is neither NULL nor a character vector.

Create a GenotypeSource object, discovering its samples with one PLINK2 --write-samples call per file. No genotypes are read.

Usage

GenotypeSource.class$new(
  name,
  path,
  files = NULL,
  files.pattern = ".*",
  format = c("vcf", "bfile", "pfile"),
  build,
  samples = NULL,
  logger = NULL,
  plink = NULL,
  nthreads = NULL,
  sample.table = NULL,
  ...
)

Arguments

name — Character scalar. Non-whitespace label for the source.

path — Character scalar. Existing directory holding the genotype files.

files — Character vector or NULL (default). Fileset stems without extension for "bfile"/"pfile", full filenames for "vcf"; NULL discovers them with files.pattern.

files.pattern — Character scalar or NULL, default ".*". Perl regular expression matched against filenames when files is NULL.

format — One of "vcf" (default), "bfile", "pfile". Layout of the fileset.

build — Character scalar. Genome-build key or name resolvable by workspace$catalogs$genomeBuilds.

samples — Character vector of IIDs, or NULL (default). Initial working set; NULL includes every sample found.

logger — Optional PolyGeniusLogger, default NULL. Scopes discovery messages.

plink — Character scalar or NULL (default). PLINK2 executable path; NULL resolves it through workspace$setup$get("plink").

nthreads — Positive integer scalar or NULL (default). PLINK --threads count for sample discovery.

sample.table — data.table with FID, IID and included, or NULL (default). Internal restore branch, reached only through genotype.source.hydrate(): installs this table as the sample roster and skips the directory check, file resolution, sample discovery and PLINK lookup; files.pattern, samples, logger and nthreads are then unused.

... — Further named values, assigned as public fields on the object.

Returns

A new GenotypeSource object.

Method print()

Print the source's name, file count, format, build, path and working versus full sample counts.

Usage

GenotypeSource.class$print(...)

Arguments

... — Unused. Present for S3 method compatibility.

Returns

self, invisibly. Called for its console output.

Method dosages()

Export dosages for selected variants with PLINK2 --export, one call per file.

Usage

GenotypeSource.class$dosages(
  variants,
  format = c("sample", "variant"),
  samples = NULL,
  load = FALSE,
  merge = FALSE,
  logger = NULL,
  nthreads = NULL
)

Arguments

variants — Character vector, data frame, or bed1-like table. Variants to extract; written to a temporary bed1 file for PLINK --extract bed1.

format — One of "sample" (default, --export A) or "variant" (--export Av). Orientation of the exported table.

samples — Character vector of IIDs, or NULL (default). Samples to export; NULL uses the current working set.

load — Logical scalar, default FALSE. When TRUE, read the exported files into R and delete them.

merge — Logical scalar, default FALSE. With load = TRUE, combine the per-file tables into one.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count.

Returns

With load = FALSE, a character vector of export paths named by source file, empty outputs dropped. With load = TRUE, merge = FALSE, a named list of data.tables. With merge = TRUE, one data.table: joined on IID for format = "sample", row-bound with a file column for format = "variant".

Method tidy()

Normalize the fileset: convert format, recode chromosomes, standardize variant IDs, and optionally merge split input first.

Usage

GenotypeSource.class$tidy(
  to = NULL,
  output = NULL,
  merge = FALSE,
  standardize.ids = TRUE,
  chromosome.format = "26",
  logger = NULL,
  nthreads = NULL
)

Arguments

to — One of "vcf", "bfile", "pfile", or NULL (default) to keep the input format.

output — Character scalar or NULL (default). Output directory, created if absent; NULL writes into $path, overwriting the current fileset.

merge — Logical scalar, default FALSE. When TRUE and the source holds more than one file, merge them first and tidy the result.

standardize.ids — Logical scalar, default TRUE. Rewrite variant IDs to the PolyGenius standard.

chromosome.format — One of "26" (default), "M", "MT", "0M", "chr26", "chrM", "chrMT", or NULL to leave codes untouched. PLINK --output-chr value.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count.

Returns

A new GenotypeSource pointing at the tidied fileset. Aborts on an unsupported to or chromosome.format.

Method lift()

Lift variant coordinates to another genome build with rtracklayer::liftOver(), then rewrite the fileset's coordinates with PLINK2.

Usage

GenotypeSource.class$lift(
  to.build,
  output = NULL,
  standardize.ids = TRUE,
  chromosome.format = "26",
  chain.path = NULL,
  logger = NULL,
  nthreads = NULL
)

Arguments

to.build — Character scalar. Target genome-build key or name.

output — Character scalar or NULL (default). Output directory, created if absent; NULL writes into $path, overwriting the current fileset.

standardize.ids — Logical scalar, default TRUE. Rewrite variant IDs after lifting.

chromosome.format — One of "26" (default), "M", "MT", "0M", "chr26", "chrM", "chrMT". PLINK --output-chr value for the lifted fileset.

chain.path — Character scalar or NULL (default). Explicit chain-file path; NULL resolves and downloads one through workspace$catalogs$liftoverChains.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count.

Returns

A new GenotypeSource on to.build, carrying two extra public fields: unlifted and ambiguous, data.tables of the variants excluded from the output, each with a leading file column. Aborts when no chain covers the build pair or chromosome.format is invalid.

Method merge()

Merge every file of this source into one fileset with PLINK2 --pmerge-list.

Usage

GenotypeSource.class$merge(
  output,
  format = NULL,
  output.name = NULL,
  variants = NULL,
  logger = NULL,
  nthreads = NULL,
  memory = NULL
)

Arguments

output — Character scalar. Output directory, created if absent.

format — One of "vcf", "bfile", "pfile", or NULL (default) to keep the input format.

output.name — Character scalar or NULL (default). Stem of the merged fileset; NULL uses $name.

variants — Character scalar path to a bed1 file, a data frame with chr, start, end columns, or NULL (default). When given, each file is slimmed to those variants before merging.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count.

memory — Positive numeric scalar (MB) or NULL (default). PLINK --memory cap; NULL leaves PLINK's default.

Returns

A new GenotypeSource holding the single merged fileset. Aborts on a blank output.name.

Method clump()

LD-clump a variant table against this fileset with PLINK2 --clump.

Usage

GenotypeSource.class$clump(
  variants,
  clump.p1 = 1e-04,
  clump.r2 = 0.5,
  clump.kb = 250,
  chromosome.column = "chr",
  position.column = "position",
  a1.allele.column = "ea",
  a2.allele.column = "nea",
  pval.column = "pval",
  label = NULL,
  logger = NULL,
  nthreads = NULL
)

Arguments

variants — Data frame of variants to clump, one row per variant.

clump.p1 — Numeric scalar, default 1e-4. Index-variant p-value threshold.

clump.r2 — Numeric scalar, default 0.5. LD r2 threshold.

clump.kb — Numeric scalar, default 250. Clump radius in kilobases.

chromosome.column, position.column — Character scalars, defaults "chr" and "position". Columns of variants holding the locus.

a1.allele.column, a2.allele.column — Character scalars, defaults "ea" and "nea". Columns of variants holding the alleles.

pval.column — Character scalar, default "pval". Column of variants holding the p-value.

label — Character scalar or NULL (default). Name of the clumped source (e.g. a GWAS id) used to attribute this call's log messages.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count.

Returns

variants filtered to the index variants, columns unchanged, carrying attr(x, "misc"): paths to PLINK's .clumps, .clumps.missing_id and .clumps.missing_allele files where written. An empty input, or one whose chromosome codes are all unmappable, returns a zero-row table of the same shape. Aborts when the source holds more than one file.

Method samples.extract()

Write a new fileset holding only the given samples (PLINK2 --keep).

Usage

GenotypeSource.class$samples.extract(
  samples,
  format = c("vcf", "bfile", "pfile"),
  output = NULL,
  logger = NULL,
  nthreads = NULL
)

Arguments

samples — Character vector of IIDs to keep.

format — One of "vcf" (default), "bfile", "pfile". Output format.

output — Character scalar or NULL (default). Output directory; NULL uses a temporary directory.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count.

Returns

A new GenotypeSource pointing at the written fileset. This object is unchanged.

Method samples.exclude()

Write a new fileset with the given samples removed (PLINK2 --remove).

Usage

GenotypeSource.class$samples.exclude(
  samples,
  output = NULL,
  logger = NULL,
  nthreads = NULL
)

Arguments

samples — Character vector of IIDs to drop.

output — Character scalar or NULL (default). Output directory; NULL uses a temporary directory.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count.

Returns

A new GenotypeSource pointing at the written fileset, in this source's own format. This object is unchanged.

Method samples.kinship()

Compute KING-robust kinship for sample pairs at or above a threshold (PLINK2 --make-king-table).

Usage

GenotypeSource.class$samples.kinship(
  threshold = 0.0884,
  variants = NULL,
  samples = NULL,
  logger = NULL,
  nthreads = NULL,
  memory = NULL
)

Arguments

threshold — Numeric scalar, default 0.0884 (the second-degree cut point). Minimum kinship coefficient to report, compared inclusively; must be finite and no greater than 0.5.

variants — Character vector, data frame, or bed1-like table, or NULL (default). Variants to restrict the computation to.

samples — Character vector of IIDs, or NULL (default). Samples to use; NULL uses the current working set.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count.

memory — Positive numeric scalar (MB) or NULL (default). PLINK --memory cap.

Returns

A data.table of related sample pairs with columns iid1, iid2, n.snp, hethet, ibs0, kinship; zero rows when nothing meets threshold or fewer than two samples are selected. IIDs are returned exactly as the .psam holds them. A pair with no marker called in both samples has no estimate and is kept with NaN kinship. Aborts on a threshold above 0.5.

Method samples.PCA()

Compute an in-sample PCA over the selected samples and variants (PLINK2 --pca). A multi-file source is merged first, since one eigendecomposition needs one fileset.

Usage

GenotypeSource.class$samples.PCA(
  npcs,
  variants = NULL,
  samples = NULL,
  approx = TRUE,
  logger = NULL,
  nthreads = NULL,
  memory = NULL
)

Arguments

npcs — Integer scalar. Number of principal components to compute.

variants — Character vector, data frame, or bed1-like table, or NULL (default). Variants to restrict the decomposition to.

samples — Character vector of IIDs, or NULL (default). Samples to use; NULL uses the current working set.

approx — Logical scalar, default TRUE. Pass PLINK's approx modifier.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count.

memory — Positive numeric scalar (MB) or NULL (default). PLINK --memory cap; NULL leaves PLINK's default.

Returns

A list with embedding (data frame, samples x npcs, rows named and ordered by IID) and eigenvalues (numeric vector). NULL when PLINK produced no output.

Method samples.PCAProject()

Project this source's samples onto a principal-component space eigendecomposed on another GenotypeSource (PLINK2 --pca on the anchor, --score on the query).

Usage

GenotypeSource.class$samples.PCAProject(
  project.onto,
  npcs,
  variants = NULL,
  samples = NULL,
  approx = TRUE,
  logger = NULL,
  nthreads = NULL,
  memory = NULL
)

Arguments

project.onto — A GenotypeSource. Anchor whose eigenvectors define the space.

npcs — Integer scalar. Number of principal components.

variants — Character vector, data frame, or bed1-like table, or NULL (default). Variants to restrict both anchor and query to.

samples — Character vector of IIDs, or NULL (default). Query samples; NULL uses the current working set.

approx — Logical scalar, default TRUE. Pass PLINK's approx modifier on the anchor decomposition.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count.

memory — Positive numeric scalar (MB) or NULL (default). PLINK --memory cap; NULL leaves PLINK's default.

Returns

A list with embedding (data frame, samples x npcs, rows named and ordered by IID, on PLINK's per-allele average scale) and, when PLINK reports them, variants, the variant IDs the projection used. Aborts when project.onto is not a GenotypeSource.

Method variant.IDs()

Read every file's variant IDs from its .bim/.pvar metadata. No genotypes are read.

Usage

GenotypeSource.class$variant.IDs(logger = NULL)

Arguments

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

Returns

Character vector of variant IDs, every file concatenated in file order.

Method variants()

Read every file's variant metadata. No genotypes are read.

Usage

GenotypeSource.class$variants(merged = FALSE, logger = NULL)

Arguments

merged — Logical scalar, default FALSE. When TRUE, row-bind the per-file tables into one.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

Returns

A named list of data.tables with columns chr, position, ID, ref, alt, one per file; with merged = TRUE, one data.table of those row-bound under a file column.

Method variants.exclude()

Write a new fileset with the given variants removed (PLINK2 --exclude).

Usage

GenotypeSource.class$variants.exclude(
  variants,
  format = NULL,
  output = NULL,
  logger = NULL,
  nthreads = NULL
)

Arguments

variants — Variants to drop, in one of three shapes: a character vector of variant ids as PLINK writes them (chr:pos:a1:a2, chromosome in --output-chr 26 style); a data frame of at least three columns read as 1-based fully-closed chr/begin/end intervals; or a path to an existing file of such intervals, which is used as given and never removed. An id file is not a supported shape -- read it in and pass the vector.

format — One of "vcf", "bfile", "pfile", or NULL (default) to keep the input format.

output — Character scalar or NULL (default). Output directory; NULL uses a temporary directory.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count.

Returns

A new GenotypeSource pointing at the written fileset. This object is unchanged.

Method variants.extract()

Write a new fileset holding only the given variants (PLINK2 --extract).

Usage

GenotypeSource.class$variants.extract(
  variants,
  format = NULL,
  output = NULL,
  logger = NULL,
  nthreads = NULL
)

Arguments

variants — Variants to keep, in one of three shapes: a character vector of variant ids as PLINK writes them (chr:pos:a1:a2, chromosome in --output-chr 26 style); a data frame of at least three columns read as 1-based fully-closed chr/begin/end intervals; or a path to an existing file of such intervals, which is used as given and never removed. An id file is not a supported shape -- read it in and pass the vector.

format — One of "vcf", "bfile", "pfile", or NULL (default) to keep the input format.

output — Character scalar or NULL (default). Output directory; NULL uses a temporary directory.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count.

Returns

A new GenotypeSource pointing at the written fileset. This object is unchanged.

Method variants.filterMAF()

Write a new fileset keeping only variants at or above a minor-allele frequency (PLINK2 --maf).

Usage

GenotypeSource.class$variants.filterMAF(
  threshold,
  samples = NULL,
  format = NULL,
  output = NULL,
  logger = NULL,
  nthreads = NULL
)

Arguments

threshold — Numeric scalar in [0, 0.5]. Minimum minor allele frequency.

samples — Character vector of IIDs, or NULL (default). Samples the frequency is computed over; NULL uses the current working set.

format — One of "vcf", "bfile", "pfile", or NULL (default) to keep the input format.

output — Character scalar or NULL (default). Output directory; NULL writes into $path.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count.

Returns

A new GenotypeSource pointing at the written fileset. This object is unchanged.

Method variants.updateIDs()

Rewrite this source's variant IDs in place, overwriting the files on disk. The only method that mutates the fileset it is called on.

Usage

GenotypeSource.class$variants.updateIDs(
  pattern = "@:#:$1:$2",
  logger = NULL,
  nthreads = NULL
)

Arguments

pattern — Character scalar, default "@:#:$1:$2". PLINK --set-all-var-ids template.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count.

Returns

self, invisibly, for chaining.

Method variants.LD()

Compute linkage disequilibrium with PLINK2 --r2-phased and friends, one call per file. Any of window, window.kb, window.r2 selects sparse output; all three NULL selects a dense square matrix.

Usage

GenotypeSource.class$variants.LD(
  variants = NULL,
  measure = "r2",
  phased = TRUE,
  maf = NULL,
  window = NULL,
  window.kb = 1000,
  window.r2 = 0.2,
  inter.chr = FALSE,
  load = TRUE,
  merge = TRUE,
  logger = NULL,
  nthreads = NULL,
  announce = TRUE,
  log.output = TRUE
)

Arguments

variants — Bed1-style ranges, or NULL (default). Restricts the computation.

measure — One of "r2" (default) or "r". LD measure.

phased — Logical scalar, default TRUE. Use PLINK's phased estimator.

maf — Numeric scalar or NULL (default). PLINK --maf floor.

window — Integer scalar or NULL (default). PLINK --ld-window, in variants.

window.kb — Numeric scalar, default 1000. PLINK --ld-window-kb.

window.r2 — Numeric scalar, default 0.2. PLINK --ld-window-r2 reporting floor.

inter.chr — Logical scalar, default FALSE. Allow inter-chromosomal pairs.

load — Logical scalar, default TRUE. Read the PLINK output into R.

merge — Logical scalar, default TRUE. With load = TRUE and sparse output, row-bind the per-file tables.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count.

announce — Logical scalar, default TRUE. Emit nested LD progress messages.

log.output — Logical scalar, default TRUE. Forward successful PLINK stdout/stderr to the logger; failed commands always emit their captured output.

Returns

With load = FALSE, a named list of output paths per file, files that produced nothing dropped. With load = TRUE: sparse output gives a named list of data.tables, or one row-bound data.table with a file column when merge = TRUE; dense output gives a named list of square numeric matrices with variant IDs as dimnames, whatever merge is. Aborts on an unsupported measure.

Method variants.associations()

Run single-variant association tests with PLINK2 --glm, one call per file.

Usage

GenotypeSource.class$variants.associations(
  variants,
  responses,
  control.for = NULL,
  samples = NULL,
  conf.level = 0.95,
  maf = NULL,
  mac = NULL,
  vif = NULL,
  reference.levels = NULL,
  load = FALSE,
  merge = TRUE,
  logger = NULL,
  nthreads = NULL,
  memory = NULL
)

Arguments

variants — Data frame of variants with chromosome and position columns. Restricts the tests.

responses — Data frame of phenotypes, one row per sample.

control.for — Data frame of covariates, or NULL (default). One row per sample.

samples — Character vector of IIDs, or NULL (default). Samples to test; NULL uses the current working set.

conf.level — Numeric scalar in (0, 1), default 0.95. Confidence level for the reported effect intervals.

maf, mac — Numeric scalars or NULL (default). PLINK --maf and --mac floors.

vif — Numeric scalar or NULL (default). PLINK --vif covariate variance-inflation ceiling.

reference.levels — Named list or NULL (default). Reference level per binary phenotype; NULL lets the encoder choose.

load — Logical scalar, default FALSE. Read the PLINK output into R.

merge — Logical scalar, default TRUE. With load = TRUE, concatenate the per-file tables.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count.

memory — Positive numeric scalar (MB) or NULL (default). PLINK --memory cap.

Returns

With load = FALSE, a character vector of .glm.* paths; with load = TRUE, a list of association tables or one concatenated table. Either way the result carries a reference.levels attribute recording the encoding actually used.

Method variants.frequencies()

Compute allele frequencies with PLINK2 --freq, one call per file.

Usage

GenotypeSource.class$variants.frequencies(
  variants,
  samples = NULL,
  load = FALSE,
  merge = FALSE,
  logger = NULL,
  nthreads = NULL
)

Arguments

variants — Character vector, data frame, or bed1-like table. Variants to restrict to.

samples — Character vector of IIDs, or NULL (default). Samples the frequencies are computed over; NULL uses the current working set.

load — Logical scalar, default FALSE. Read the .afreq files into R.

merge — Logical scalar, default FALSE. With load = TRUE, concatenate the loaded tables.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count.

Returns

With load = FALSE, a character vector of .afreq paths named by and aligned to the source files, NA where a file produced no output. With load = TRUE, a list of frequency tables, or one concatenated table when merge = TRUE.

Method scores()

Compute polygenic scores with PLINK2 --score.

Usage

GenotypeSource.class$scores(
  models,
  samples = NULL,
  maf.threshold = 0,
  max.cells.per.batch = 2e+08,
  models.per.group = Inf,
  logger = NULL,
  nthreads = NULL,
  memory = NULL,
  file.backed.threshold = 2.5e+08,
  model.filter.columns = NULL
)

Arguments

models — A PGS, a PGSLibrary, a data frame of variants, a list of data frames, or a list of PGS objects. The models to score.

samples — Character vector of IIDs, or NULL (default). Samples to score; NULL uses the current working set.

maf.threshold — Numeric scalar in [0, 0.5], default 0. Minor-allele-frequency floor for scored variants.

max.cells.per.batch — Numeric scalar, default 200e6 (about 800 MB in PLINK, 1.6 GB in R). Cap on variant x model cells per PLINK --score call; controls peak memory.

models.per.group — Positive numeric scalar, default Inf. Models scored per group (model-axis streaming); a finite value lowers peak records at the cost of re-reading the genotype once per group. Scores are identical either way.

logger — Optional PolyGeniusLogger, default NULL. Scopes this call's messages.

nthreads — Positive integer scalar or NULL (default). PLINK --threads count. Files and batches are scored sequentially; this parallelises within one PLINK call.

memory — Positive numeric scalar (MB) or NULL (default). PLINK --memory cap; NULL leaves PLINK's default.

file.backed.threshold — Numeric scalar, default 2.5e8. Backbone-backed input only: the dense backbone$n.variants() x M cell count above which the scored union is read out-of-core from the weight store rather than from the in-memory backbone.

model.filter.columns — Internal only, not user-facing. NULL (default), or a model.filter mask pushdown -- see GenotypeSourceScores$compute()'s own @param.

Returns

A list with scores (numeric matrix, samples x models, columns labelled positionally m1..mM), given.names (those labels, in model order -- the caller reorders on these and applies its own model names, which may repeat), variant.fate (per variant, what became of it), strategy ("fused" or "via.extract", the regime actually chosen) and n.batches (number of PLINK --score calls made).

Method sample.table()

Return the fileset's sample table, as discovered at construction.

Usage

GenotypeSource.class$sample.table()

Returns

A data.table copy with columns FID, IID and included, one row per sample in the full fileset; included marks the current working set. Modifying the copy does not change the object.

Method clone()

The objects of this class are cloneable with this method.

Usage

GenotypeSource.class$clone(deep = FALSE)

Arguments

deep — Whether to make a deep clone.

Examples

geno <- GenotypeSource(
  name   = "cohort",
  path   = "/data/genotypes",
  format = "pfile",
  build  = "GRCh38"
)
geno$n.samples

# Restrict the working set; the files on disk are untouched.
geno$samples <- head(geno$samples.full, 100)

See Also

Aliases: GenotypeSource, GenotypeSource.class