Benchmark

Whole-archive recovery vs archive size under a stress channel

Simulated

Result
20 KB: 28 / 30 EXACT; 1 MiB: 2 / 30 EXACT; 0 false success at both sizes (archives recovered exactly)
Classification
Software-generated DNA strands passed through a software channel model. Not a laboratory result.
Methodology
Per-strand and per-row success measured; Wilson 95 % intervals (20 KB 0.79-0.98, 1 MiB 0.02-0.21). Larger sizes only by an analytic row model, labelled as extrapolation.
Dataset
Seeds 93000-93029; D13-F1 channel, mean coverage 10, NB dispersion 4, profile v4-balanced.
Configuration
Decoder E, 4 workers. A size is decoded only if a single decode is projected at ≤ 2 h on this host; 10 MB (projected 3.5 h) and larger were not decoded.
Environment
Development host: shared VPS, Intel Xeon Gold 6240 @ 2.60 GHz, 4 cores / 8 threads, Ubuntu 24.04, Python 3.12, CPU only. Host load 1.0-4.3 during the 1 MiB decodes.
Limitations
Per-row success ≈ 0.995 is enough for 9 rows but not for 411, so failures compound with size. The analytic model gives ≈ 0 at 100 MB in this setting: large archives under this channel need more coverage or parity. This is a known V9 limitation and a V10 work item.
Version
V9
Source
v9.0.0/experiments/v9/scale/results/noisy.jsonl
Reproduce
experiments/v9/reproduce.sh scale · how to reproduce