Benchmark
Consensus V2 selection: candidate E vs the V8 decoder
Simulated
- Result
- EXACT 283 / 600 vs 240 / 600; pooled ORE 0.575 vs 0.488; Holm-adjusted exact McNemar p = 5.2e-12; 0 false success (archives recovered exactly)
- Classification
- Software-generated DNA strands passed through a software channel model. Not a laboratory result.
- Methodology
- Pre-registered before any evaluation decode (docs/V9_PREREGISTRATION.md, amendments A1, A2). Each candidate evaluated once on the same read pools (reads_sha256 checked). Paired exact McNemar test, Holm correction over candidates; Wilson 95 % intervals for pooled ORE.
- Dataset
- 600 paired EVAL cases: seeds 91000-91099 in six cells (coverage 3, 5, 10 × profiles v4-balanced, v7-lowcov) under the D13-F1 stress channel.
- Configuration
- Candidate E = ClusterConfig(consensus_template="full", c_indel=4, fill="template"), opt-in. Baseline = V8 production decoder.
- Environment
- Development host: shared VPS, Intel Xeon Gold 6240 @ 2.60 GHz, 4 cores / 8 threads, Ubuntu 24.04, Python 3.12, CPU only.
- Limitations
- D13-F1 is a stress channel, not a validated model of nanopore reads (its successor G1 was rejected). At coverage 3 no candidate recovers any archive. E is opt-in; the library default is unchanged.
- Version
- V9
- Source
- v9.0.0/experiments/v9/consensus/results/selection.json
- Reproduce
experiments/v9/reproduce.sh decoder # full; `spot` re-runs a pre-specified subset· how to reproduce