Extending PolyGenius
Adding a construction algorithm, a GWAS source, an analysis or a plot — and the contracts each must honour
Most extensions are one of four things: a new way to construct a model, a new place to get summary statistics from, a new analysis of scores, or a new figure. Each plugs into a different seam, and each has a contract that exists because breaking it produces plausible wrong answers rather than errors.
This chapter is the map of those seams. It assumes you have read data architecture and the execution engine.
Adding a construction algorithm
This is the largest of the four, because a construction method touches the resource system, the execution engine and the model contract at once.
The two halves
A specification constructor is user-facing. It validates parameters and returns a set of resource specifications — one per model the declaration implies. It does no work.
A rule produces the resource. It declares what it needs, what it costs, and how to run.
Keeping them separate is what makes declaration cheap: you can construct and combine specifications for hundreds of models without touching a file.
What the constructor owes
Validate every parameter at construction, not at run time. A typo in a parameter name should fail when the user writes it, not forty minutes into a job.
Decide which parameters are part of the model's identity. Anything that changes the output is; anything that only changes how the computation runs is not. Thread counts and interpreter paths belong in the runtime bucket, everything else in identity.
Expand grids at the point that keeps the work amortised. If your method can fit many parameter values from one pass over the data, emit one specification and let the rule produce several models. If each value is an independent fit, emit one specification per value — but understand that you are declaring N independent runs, and the guide will say so.
What the rule owes
Declare your inputs honestly. If you need a panel in a particular format, declare that format; the engine will insert the conversion.
Declare your cost. Under-declaring memory means the scheduler over-commits the machine.
Return models that satisfy the model contract:
- Weights are log-scale per-allele effects. If your method's natural output is on another scale, convert before returning.
- The build is a supported build, and it matches what your inputs were on.
- Provenance is complete: which GWAS, which algorithm, which parameters. This is what makes a PGS library self-describing six months later.
Do not write outside your resource directory. The engine wipes that directory before each attempt, which is what prevents a half-written payload from being mistaken for a complete one — a file written elsewhere escapes that guarantee.
The sample-size contract
This one deserves its own note because it has caused real bugs.
Methods that shrink effect estimates need the effective sample size, since it determines each estimate's variance. For a quantitative trait that equals the total; for a case/control trait it is the standard effective-N formula.
There is one shared resolution helper for this, and new algorithms should use it rather than reading a column directly. It prefers an explicit per-variant effective size, then derives one from case and control counts, then falls back to a total with a warning. Reading a column named for a total and treating it as effective is exactly the mistake the helper exists to prevent.
Gating
Adding an algorithm is phased and gated, not a single change. The process — literature reading, resource alignment, phased implementation with per-phase tests, and three independent reviews consolidated against the method's own publication — is documented in the repository's contributor guide rather than here, because it is a development process rather than an API.
Adding a GWAS source
Smaller, and the same two-part shape: a constructor that declares, a rule that fetches.
The output contract is a variant table with chromosome, position, effect and other allele, effect size and p-value, plus normalised metadata.
Four things a source must do:
Normalise the build label before it reaches identity. The identity layer hashes what it is given, so two spellings of one build would otherwise be two resources.
Coerce coordinate types and drop rows without a usable coordinate — but count them and log the count. Silently discarding rows is how variant counts stop adding up.
Emit the canonical sample-size fields. If your source knows case and control counts, emit them; the shared resolution helper will do the right thing. Emitting only a total means every downstream consumer falls back and warns.
Keep the raw provider payload alongside the normalised metadata. Users ask questions the normalisation did not anticipate.
One thing a source must not do: assume it knows what threshold will be needed. The effective fetch threshold is the loosest one across every algorithm in the call, resolved by the generation engine, not by the source.
Adding an analysis
An analysis takes a cohort object and returns a result. The contracts here are about the result rather than the computation.
If it belongs in associate
Declare a schema: the columns you always produce, the ones you may, which artifacts are required and which optional, which plots apply, and which columns identify a poolable cell.
Use the shared core columns wherever your analysis fits them, because that is what lets existing plots and the pooling engine work on your output without modification. Diverge only when the shared vocabulary would be actively misleading — the single-variant schema diverges for exactly that reason.
Declare an effect scale from the existing closed vocabulary. If your estimate is on a scale that is not in it, that is a conversation before it is a code change: the vocabulary is what lets the pooling engine refuse to combine incompatible quantities.
Declare whether your results can be pooled by giving pooling keys, or declare that they cannot by leaving them empty. An analysis whose results should not be meta-analysed — because it produces a test statistic rather than an effect with a standard error — should say so structurally rather than in prose.
Emit artifacts for anything a plot will need. This is the boundary that keeps figures and tables in agreement, and it is your responsibility as the analysis author: a plot cannot compute what you did not emit.
If it belongs in evaluate
Register your metrics. The registry, not the producer, defines each metric's direction, its ideal value, its role, and whether it participates in model selection. A metric with no registry entry has no direction and cannot be ranked or labelled.
Be deliberate about the role. A diagnostic p-value or an effect size should not be able to pick a model, and the registry is where you say so.
Where multiplicity lives
Decide your adjustment grouping and document it on the analysis. There is no global correction, and adjustment does not compose across calls — so your grouping is a real design decision, not a default you inherit.
Adding a plot
The smallest extension, with the strictest rule.
A plot must not compute a statistic it reports. It may derive anything cosmetic: ordering, display binning, an axis transform, label placement. It may not fit a model or compute a summary that appears as a number in the figure.
If a plot needs a quantity, the analysis must emit it as an artifact. That is a change to the analysis, not a workaround in the plot.
Beyond that: register under the accessor tree so it is discoverable, take the standard filtering arguments so it composes with filtered results, resolve colour through the palette system rather than hard-coding, and document your return type — the four shapes behave differently under composition and users need to know which they have.
Fail with a message naming the missing artifact rather than rendering something empty.
What extension does not currently support
Two honest gaps.
Plot registration is advisory. Schemas declare which plots apply, but nothing reads that declaration — only the default plot is consulted. A plot not declared by any schema is still reachable, and a schema can declare a plot that does not exist. Do not rely on the declaration to constrain anything.
There is no plugin discovery. Extensions are additions to the package, not separately-installable units. A new rule has to be registered in the package's own rule list.
Where to go next
Code organisation covers where each of these lives in the source tree and what the review expectations are.