Code Organisation and Contributing
How the source tree is arranged, the conventions that hold across it, and what a change is expected to include
One file, one domain
Every source file belongs to exactly one domain, and the domain is in its name. This is the first thing to internalise, because the boundaries are enforced socially rather than mechanically and crossing one silently is how a codebase turns into a single tangle.
| Prefix | Responsibility |
|---|---|
generate-* |
model specification, GWAS retrieval, construction algorithms |
compute-* |
scores, population structure, relatedness, embeddings |
associate-* |
regression, single-variant, comparison, mediation, pooling |
evaluate-* |
performance, incremental value, association, stratification, comparison, redundancy, selection |
visualize-* |
plots, and nothing else |
data-classes-* |
the core objects and result classes |
statistics-* |
regression families, schemas, outcome specifications |
genome-* |
positioned-signal reduction shared by compute and visualize |
catalogs-* |
reference panels, chains, variant spaces, LD blocks, tool discovery |
internals-* |
the execution engine, scheduler and resource store |
genotype-info-* |
PLINK file inspection and per-genotype operations |
Three rules follow from this and are worth stating as rules:
visualize-* never computes a statistic. If a figure needs a quantity, the analysis
emits it.
associate-* and evaluate-* never silently run compute-*. An analysis that needs
scores requires them to exist; neither computes or standardises them behind your back.
Each domain's output is a typed object, not a plain list.
Naming within a domain
<domain>-zz-<domain>.R— the user-facing accessors. New public functions go here.<domain>-engine.R— the computation core, internal.<domain>-schemas.R— schema and contract definitions.<domain>-<feature>.R— a specific feature.
The zz- prefix exists so those files collate last, after everything they reference.
Conventions
Separators are dots, never underscores. my.function, result.table, n.events.
This is consistent across the whole package and mixing in underscores is the most common
review comment on a first contribution.
Imports are declared, not inlined. Hard dependencies are imported in roxygen; do not
write pkg::fn in the body. Packages used pervasively are imported wholesale.
Optional dependencies are checked at the point of use, with a clear message naming what to install and which function needed it.
Comments explain why, not what. A comment restating the code is noise; a comment recording why a non-obvious choice was made is the most valuable thing in a file. Non-obvious invariants, magic constants and domain logic that would not be guessable should carry one.
Private functions use plain comment blocks, not roxygen — a roxygen block on an unexported function generates a broken documentation entry.
No premature abstraction. Three similar lines are fine. A helper that exists to avoid repeating two lines makes both call sites harder to read.
Validate at boundaries. User input and external API responses get checked; internal calls are trusted. Validation everywhere is noise that hides the checks that matter.
External tools go through one helper. Never call out to the shell directly — the helper handles logging, error capture and the conventions the rest of the package relies on.
The user-facing surface
The package exposes namespaced environments rather than a flat list of functions:
generate$models(sources = ..., algorithms = ...)
compute$scores(data)
associate$regression(data, outcomes = ...)
evaluate$profile(data, outcomes = ...)
visualize$data$scores$distribution(data)Very little should be reachable outside those environments plus workspace. A few things
have no accessor home and are called directly — the outcome helpers, liftover, and the
object constructors.
Note that some accessor functions are also exported under their bare names. Browsing the reference index you will see both forms. The accessors are the API the documentation teaches; the bare forms are an artefact and not the recommended surface.
Tests
Four tiers, and putting a test in the wrong one is a real cost — it either does not run or it makes the fast suite slow.
Unit (tests/testthat/) — mocked. No PLINK, no network, no engine workers. This is
where declarations, contracts, invariances and analytic limiting cases are pinned. It
must run under devtools::load_all(), because the helpers use package internals;
running it against an installed package fails in ways that are not real failures.
Integration (tests/integration/) — real PLINK2 against synthetic fixtures, never a
real download. Assert structure, determinism, cache identity and error paths. Never
assert statistical quality: the genotypes are random, so any quality claim is noise.
Slow (tests/slow/) — real resources and internet. This is where statistical
correctness and agreement with an upstream implementation belong. Must skip rather than
pass silently when a resource is absent.
Benchmark (tests/benchmark/) — engine throughput, fixture-backed, opt-in.
Two habits that matter more than the tier rules:
A test that passes before and after your change proves nothing. Write the test first, watch it fail, then fix. Several real defects in this package survived behind tests that were structurally incapable of failing — a mock that satisfied two contradictory contracts, an assertion that checked a field name rather than its value.
Propose slow-tier tests whenever statistical correctness is at stake, and say plainly if they are deferred for want of resources. A deferred oracle leaves a stated gap; an unmentioned one leaves a false impression.
What a change is expected to include
Code, tests, and documentation, in one change rather than three.
Documentation means the roxygen content — parameters, return values, the sections that describe behaviour — regenerated so the manual pages match. A changed signature with a stale manual page is an incomplete change.
If user-facing behaviour changed, the website source changes too.
Reference pages
The repository carries a set of internal reference pages describing each subsystem. They exist so that a contributor — human or otherwise — does not re-derive a subsystem's contracts from source every time.
Three rules govern them, and they exist because a confidently wrong page is worse than a stale one:
If your change makes a page wrong, fix it in the same change and say so in your summary. The old text is wrong by construction and leaving it guarantees drift.
If a page disagrees with code you did not touch, report it — do not reconcile it. Give both readings, say which you believe was intended, and get a decision. Editing the page to match the code quietly ratifies whatever drift happened, and makes the discrepancy invisible because the two now agree on the wrong thing.
Never delete a recorded fact to make a page agree with new code. Mark it superseded, with the reason.
Cite a symbol and a file, never a line number — line numbers rot on every edit above them.
Working practice
Prefer several small, independently verifiable changes over one large one. Each should leave the tree working, with its own tests.
The full contributor guide in the repository has the current details on branching, gates and the phased process for larger work. It is the authority; this chapter is the map.