Benchmark

Consensus V2 selection: candidate E vs the V8 decoder

Simulated

Consensus V2 selection: candidate E vs the V8 decodercov 5 · v7-lowcov V8: 61; cov 5 · v7-lowcov E: 89; cov 10 · v4-balanced V8: 79; cov 10 · v4-balanced E: 94; cov 10 · v7-lowcov V8: 100; cov 10 · v7-lowcov E: 100; cov 5 · v4-balanced V8: 0; cov 5 · v4-balanced E: 0; cov 3 · both V8: 0; cov 3 · both E: 0EXACT of 1006189cov 5 · v7-lowcov7994cov 10 · v4-balanced100100cov 10 · v7-lowcov00cov 5 · v4-balanced00cov 3 · bothCell (coverage, profile)
SimulatedV8EMethodology and limitations
Result
EXACT 283 / 600 vs 240 / 600; pooled ORE 0.575 vs 0.488; Holm-adjusted exact McNemar p = 5.2e-12; 0 false success (archives recovered exactly)
Classification
Software-generated DNA strands passed through a software channel model. Not a laboratory result.
Methodology
Pre-registered before any evaluation decode (docs/V9_PREREGISTRATION.md, amendments A1, A2). Each candidate evaluated once on the same read pools (reads_sha256 checked). Paired exact McNemar test, Holm correction over candidates; Wilson 95 % intervals for pooled ORE.
Dataset
600 paired EVAL cases: seeds 91000-91099 in six cells (coverage 3, 5, 10 × profiles v4-balanced, v7-lowcov) under the D13-F1 stress channel.
Configuration
Candidate E = ClusterConfig(consensus_template="full", c_indel=4, fill="template"), opt-in. Baseline = V8 production decoder.
Environment
Development host: shared VPS, Intel Xeon Gold 6240 @ 2.60 GHz, 4 cores / 8 threads, Ubuntu 24.04, Python 3.12, CPU only.
Limitations
D13-F1 is a stress channel, not a validated model of nanopore reads (its successor G1 was rejected). At coverage 3 no candidate recovers any archive. E is opt-in; the library default is unchanged.
Version
V9
Source
v9.0.0/experiments/v9/consensus/results/selection.json
Reproduce
experiments/v9/reproduce.sh decoder # full; `spot` re-runs a pre-specified subset · how to reproduce