PolyGenius
Contents

as.bigsnpr

Convert a GenotypeSource to bigsnpr file-backed format

Writes the genotypes a GenotypeSource() points at into bigsnpr's file-backed .rds/.bk pair and attaches the result. This is what makes a PolyGenius genotype fileset usable by bigsnpr-based algorithms such as LDpred2 and lassosum2.

Usage

as.bigsnpr(
  genotype.source,
  output = NULL,
  output.name = NULL,
  ind.row = NULL,
  ind.col = NULL,
  ncores = 1,
  nthreads = NULL,
  logger = NULL
)

Arguments

ArgumentDescription
genotype.sourceA GenotypeSource. Fileset to convert; its $samples working set selects the samples read.
outputCharacter scalar or NULL (default). Directory the .rds and .bk files are written to, created if absent; NULL uses tempdir().
output.nameCharacter scalar or NULL (default). Base name for the output files, without extension; NULL uses genotype.source$name.
ind.rowInteger vector or NULL (default). Row (sample) indices passed to bigsnpr's snp_readBed2(), applied to the fileset that remains after the working-sample filter; NULL reads every sample in genotype.source$samples.
ind.colInteger vector or NULL (default). Column (variant) indices passed to snp_readBed2(); NULL reads every variant.
ncoresInteger scalar, default 1. Cores bigsnpr uses for the conversion step.
nthreadsPositive integer scalar or NULL (default). PLINK --threads count for the preparation steps; NULL passes no --threads flag.
loggerOptional PolyGeniusLogger, default NULL. Scopes the progress messages; NULL creates one.

Value

A named list: obj — The attached bigSNP object.

rds — Character scalar path to the written .rds file.

map — Data frame of variants aligned to the genotype columns, with chr (character), pos (integer), a0 (reference allele), a1 (effect allele), marker.id (standardized PolyGenius ID) and genetic.dist.

fam — Data frame of sample information, as bigsnpr reports it.

genotypes — The file-backed matrix (FBM), samples x variants.

build — Character scalar genome build, copied from genotype.source. Aborts when genotype.source is not a GenotypeSource, when nthreads is neither NULL nor a positive integer, or when bigsnpr is not installed.

Details

The input is prepared with PLINK2 before bigsnpr sees it: converted to bfile format with $tidy() when it is not already, merged into one fileset with $merge() when it holds more than one, and filtered with --keep when the working sample set is smaller than the fileset or when ind.row is supplied. Those intermediates are written to a temporary directory that is removed when the call returns; only the .rds and .bk files land in output.

Variant identifiers in map$marker.id follow the PolyGenius standard, chr:pos:a0:a1 with the alleles in lexicographic order, so they match the IDs used elsewhere in the package (for example when clumping). Chromosome codes are normalised to PLINK numeric coding; a code that cannot be resolved is left verbatim rather than dropped, because map rows are positionally aligned with the genotype matrix columns.

Requires the bigsnpr package, which is a soft dependency: install it with install.packages("bigsnpr").

Examples

# Convert a reference panel to bigsnpr format
ref.panel <- workspace$catalogs$referencePanels$get("EUR", "hg19", format = "pfile")
bigsnp.obj <- as.bigsnpr(ref.panel)

# The file-backed genotype matrix, samples x variants
G <- bigsnp.obj$genotypes
dim(G)

# Keep the conversion outside tempdir()
bigsnp.perm <- as.bigsnpr(
  ref.panel,
  output = "/path/to/permanent/storage",
  output.name = "EUR.ref.hg19"
)

See Also