Why Runtime Variance Can Invalidate an Otherwise Correct Benchmark
A narrow note on separating valid output from publishable performance measurement.
Note identity
- Note ID
- FRN-2026-002
- Schema
- FARPY-RESEARCH-NOTE-V1
- Publication date
- 2026-07-21
- Revision
- 1
- Source report
- FRS-2026-001
Question
Can visually valid benchmark output coexist with runtime measurements that are too unstable to publish?
Finding
Yes, under the tested Octane contract and configuration. The decoded pixel payloads matched, but timing did not satisfy the frozen acceptance rule. The runtime coefficient-of-variation gate was at most 3%. The best consecutive three-pass CV was 15.579055%, and additional warmups did not stabilize the measurements. FARPY therefore rejected performance publication and published no verified Octane benchmark result.
Evidence
FRS-2026-001 separates output validation from performance acceptance. Its canonical record shows matching decoded pixels, a frozen runtime CV gate of ≤3%, a best consecutive three-pass CV of 15.579055%, and unsuccessful extra warmups. Those facts support a research finding about the rejection decision, not a verified score or runtime claim.
Interpretation
Visual correctness is necessary for this workload, but it is not sufficient for a defensible performance benchmark. A benchmark runtime is useful only when the measurement process meets its predeclared stability rule. Publishing an unstable timing result would attach false precision to a measurement that failed its own gate. Rejecting the performance result preserves the distinction between a correct-looking output and a reliable performance estimate.
A rejected benchmark can still be useful research evidence. It documents the failure mode, shows which gate prevented publication, and identifies where additional controlled work would be needed. Rejection is an outcome of the protocol, not missing data to be silently promoted.
Limitation
This note applies only to the tested workload contract and identified configuration. It does not generalize the observed variance to other Octane versions, scenes, GPUs, drivers, or hosts. It does not claim that Octane cannot be benchmarked, and it does not create a verified Octane performance result.
Citation
Farpy Research Station. "Why Runtime Variance Can Invalidate an Otherwise Correct Benchmark." FRN-2026-002, revision 1, 2026.
Related full report
FRS-2026-001: Octane Pixel Determinism Versus Runtime Instability
Canonical URL: https://farpy.com/labs/research/frn-2026-002-runtime-variance-invalid-benchmark