Running and Debugging a Generation Job
Scaling a run across sources and algorithms, reading what came back, and what caching does for you
Many sources, many algorithms
A generation call pairs every source with every algorithm, so two GWAS and three algorithms is six declarations before any grid expansion.
Count before you pay: ask for the plan without executing it and see how many models it implies. Note that the preview does not check whether the plan fits your core budget, so it can show you something that aborts on execution.
Grid mechanisms differ in cost, which chapter 09 covers. What matters here is the tally: candidates multiply into score columns and evaluation rows downstream.
Naming what comes back
Model names are generated to distinguish siblings: the trait when more than one GWAS is involved, the algorithm unless it is implied, and only the parameters that actually vary.
Grid models routinely share a display name, so residual duplicates get a numeric suffix. You will see those suffixes in plot legends, by design.
"It ran" is not "it all ran"
A generation call does not error when a model fails. It returns the models that succeeded, together with a record of every failure — and if nothing succeeded, that record comes back with an empty set rather than an exception.
So the first thing to do after any run is compare what you got against what you asked for.
The failure record carries, per failed model, the reason and the parameters that defined it — so you know both what went wrong and which model it was.
Alongside it are the paths to the run's log and event stream, so you never have to guess which run you are looking at.
Triage, in order
- Count what you got against what you asked for. If they match, your problem is not generation.
- Read the failure record. The reason names the cause, on the same row as the model's parameters.
- Classify before investigating. A budget message is a configuration problem, not a data problem. A dependency message means this model is collateral — find the one failure that is not a dependency and fix only that.
- Read the resource log for the failing model, which carries the underlying tool's own output plus everything upstream of it.
Those four steps resolve most runs in a couple of minutes.
What caching does for you
Every intermediate is stored and reused: a retrieved GWAS, a reference panel, an LD build, a clumped variant set, each model. Re-running the same call recomputes none of it.
What defines a stored result is the parameters you asked for. Every algorithm argument counts, except the ones that only affect how the work runs rather than what it produces.
A fully cached re-run looks like nothing happened. It finishes almost instantly, the progress bar goes straight to complete, and the run dashboard is empty because no work was dispatched. That is a cache hit, not a fault.
One reassurance about thresholds, since it is easy to worry about the wrong thing: tightening a p-value threshold does produce a new model. What gets reused is the download — summary statistics fetched at a looser threshold satisfy a stricter request, so you are not re-downloading.
Two ways to get a stale answer
Your own edited input. A local GWAS's contents are not part of what identifies it. Change the file, keep the same identifier, re-run, and you silently get the model built from the old data. Version the identifier, or generate into a fresh store.
An upgrade. What identifies a model is what you asked for, not the version of PolyGenius that built it. That is what makes caching reproducible within a project — and it means an upgrade does not invalidate anything. When a release note says a construction algorithm changed, generate into a fresh store so the new code actually runs, and record your package version alongside a result set for the same reason you record your reference panel.
There is no supported way to force a re-run: change a parameter, point the store elsewhere, or remove the stored result from disk.
Configuration worth setting
The store root, which is where everything above lives.
The memory budget, which defaults to unlimited. That means memory never limits how many tasks run at once, while each LD-based task declares a substantial requirement — so a dozen workers can commit far more than your machine has. Set it to your allocation. Setting it too low does not queue the work; it fails those tasks with a message saying the requirements exceed the budget.
The core count is a ceiling request rather than a guarantee: PolyGenius caps it just below the cores it can detect. Print the workspace to see what you actually got — if it says three on a machine you believe has sixty-four, you are not on the node you think you are.
The store, briefly
Generated results live under the store root, one directory per result. Models are tiny; what fills a disk is reference panels and LD resources, and those are shared by every model that used them, so do not delete them casually.
Deleting a stored result's directory is safe — the next run notices it is gone and recomputes it.
Layout, indexing and the deeper mechanics are in the execution engine chapter.