On this page
A number is not a benchmark
A useful public benchmark records the exact candidate, workload, model and data identities, scale, hardware, runtime, affinity, warm-up, sample count, order/randomisation, aggregation, uncertainty, failures, exclusions, endpoint, units, source artifact, and claim boundary.
Four result classes
- Component engineering result: measures one named code path and cannot imply end-to-end behavior.
- Scout: explores direction or feasibility and cannot support a public performance or quality claim.
- Matched systems result: holds relevant semantics, workload, and environment constant and measures the actual endpoint repeatedly.
- Claim-grade result: adds frozen protocol, quality parity or materiality, sufficient repetition, statistics, full lineage, independent review, and claim closure.
Published historical engineering family
The extreme-context ladder records three CPU-only synthetic candidates, 20 decode trials per context and mode, 10 cache-transition trials per context and profile, 10,000 bootstrap resamples, and explicit resource admission. Tiny and Small completed the 512-to-524,288-token ladder; Medium completed every safely admitted point through 131,072 tokens. Across 4,991 ordered metric rows, the observatory reported zero errors.
The all-zero K/V history does not measure language quality. Turbo paths are semantic-changing experiments, fixed-shortlist failures remain negative evidence, and the recorded cache values are tensor representations rather than whole-process RSS. The family supports only its current-host protocol, exact storage counts, semantic diagnostics, and admission observations. Inspect the normalized evidence.
Published historical component result
Direct block-region counting retained the tested sparse-plan identity and reduced mean 65,536-token plan-construction wall time from 784.49 ms to 7.99 ms across 20 measurements per arm on one disclosed CPU host—a 98.98% component reduction.
The arms ran sequentially in one process and were not randomised or isolated. The result supports the exact component and host-local wording only. It does not support end-to-end training or generation speedup, model quality, memory or energy savings, cross-host performance, or superiority. Inspect the source-linked evidence.
New held evidence package
The 28 July sanitised handoff adds a 50-replicate Tiny CPU candidate through 2,048K synthetic all-zero K/V context. At 2,048K, full-vocabulary decode averaged 357.8 tok/s. Turbo averaged 441.9 tok/s, but exact sequence agreement was 82% and first-step logits were not bitwise equal, so it is retained as a semantic-changing diagnostic rather than a promoted performance result. The resource guard refused 4,096K before allocation.
The handoff also records two negative sparse-geometry component results. K8 missed the frozen approximation gate; at equal active width 1,152, both registered geometries failed, and the finer geometry first passed only at 2,880 of 3,072 active neurons. These findings inform refracture and recovery. They are not language-quality, serving, or end-to-end benchmarks.
Source-family CI passed. Final assembly CI did not execute because of an external billing limit, so the terminal handoff remains held from current-claim promotion.
Historical reconstruction anchor
Complete K32 reconstruction matched the pinned Qwen3 donor on five frozen engineering cases for first logits, first-token top-five ordering, and full greedy generation, with 0.0 maximum observed first-token logit error.
This isolates future sparse approximation from basic fracture or loading error. Five cases do not establish broad equivalence, and K32 activates the full expert bank. It says nothing yet about material sparse quality or savings.
The current evidence refresh remains held. These records retain value only at their published source boundaries and must not be presented as current-stack benchmark results.
Scout promotion rule
A scout does not reach the public site as a quality or performance claim unless:
- the exact candidate and source artifact are pushed and pinned;
- semantic identity or a pre-agreed tolerance passes for the measured change;
- the actual claim endpoint is repeated under a frozen method;
- quality is measured before or alongside systems benefit;
- failures, warm-up, order, host, and uncertainty remain visible;
- a matched control and all relevant resource costs are included; and
- the claim owner approves exact wording and prohibited inferences.
The site may publish a bounded component result before end-to-end proof only when the component is named prominently and broader inferences are explicitly blocked.
Next decisive conversion ladder
Hold donor, prompts, attention mode, runtime, hardware, and measurement method constant while evaluating dense donor, K32, K24, K16, K8, and adaptive K. Every rung reports benchmark quality and token-level divergence. Only quality-qualified rungs proceed to replicated latency, throughput, resident memory, and calibrated energy measurement.
Reduced selected work, pair count, active parameters, or planner time is not commercial savings until end-to-end wall-clock and resource measurements confirm it.
Failure and uncertainty
Failed, interrupted, warning-bearing, and manually reviewed runs remain in the ledger. Post-result threshold changes create a new exploratory protocol and cannot rescue the original claim. Inconclusive remains inconclusive.