Navigate

Quick find

Navigate NeuroForge

Esc

Start typing to search public pages and evidence.

    Measured evidence

    Extreme-context decode ladder

    A replicated current-host context ladder across three CPU-only synthetic candidates, with cache-representation counts, semantic diagnostics, and admission boundaries.

    Recorded finding: Tiny and Small completed the declared ladder through 524,288 tokens, while Medium completed every safely admitted point through 131,072; 4,991 metric rows closed with zero observatory errors.

    On this page

    Currency and next evidence gate

    Decision signals

    Key metrics and controls

    The recorded values, with their controls and scope.

    524,288 tokens

    Largest admitted context

    Reached by the Tiny and Small synthetic candidates; Medium stopped at its declared admission boundary.

    3

    Replicated candidate scales

    20 decode trials per context and mode; synthetic CPU-only protocol.

    4,991

    Ordered metric rows

    87,451 scalar values; 0 observatory errors.

    0/11

    Tiny shortlist exact contexts

    Retained negative result: the faster fixed 16-token shortlist did not preserve exact sequences.

    Source-linked visual evidence

    What the evidence shows

    Open any chart’s data table to inspect the exact values, or download the CSV.

    Source-linked chart

    Safely admitted context ceiling by candidate scale

    Largest completed power-of-two context under the declared current-host memory admission policy.

    Finding: Tiny and Small completed through 524,288 tokens; Medium completed through 131,072 and refused the two larger points before allocation.

    Safely admitted context ceiling by candidate scaleLargest completed power-of-two context under the declared current-host memory admission policy. Tiny and Small completed through 524,288 tokens; Medium completed through 131,072 and refused the two larger points before allocation.0200,000400,000600,000tokensTinySmallMediumTinyTiny524,288SmallSmall524,288MediumMedium131,072
    View source values

    Scroll sideways if any columns are out of view.

    Safely admitted context ceiling by candidate scale; values in tokens
    GroupMaximum completed context
    Tiny524,288 tokens
    Small524,288 tokens
    Medium131,072 tokens

    Source-linked chart

    Recorded cache tensor-representation ratio

    Sequence-cache tensor storage divided by prepared recurrent-representation storage at each scale's largest completed context.

    Finding: The recorded tensor-only ratio is about 15,887.5× for Tiny and Small and 1,985.9× for Medium at their respective admitted ceilings.

    Recorded cache tensor-representation ratioSequence-cache tensor storage divided by prepared recurrent-representation storage at each scale's largest completed context. The recorded tensor-only ratio is about 15,887.5× for Tiny and Small and 1,985.9× for Medium at their respective admitted ceilings.05,00010,00015,00020,000×Tensor representation ratioTinyTiny15,887.5SmallSmall15,887.5MediumMedium1,985.9
    View source values

    Scroll sideways if any columns are out of view.

    Recorded cache tensor-representation ratio; values in ×
    GroupTensor representation ratio
    Tiny15,887.5 ×
    Small15,887.5 ×
    Medium1,985.9 ×

    Method

    How to read this result

    1. Three CPU-only candidate scales were measured on one disclosed host using a synthetic all-zero K/V history.
    2. Tiny and Small covered powers of two from 512 through 524,288 tokens; Medium covered each safely admitted point through 131,072 tokens.
    3. Each context and decode mode used 20 decode trials; each context and cache profile used 10 transition trials, with 10,000 bootstrap resamples for mean confidence intervals.
    4. The Medium 262,144- and 524,288-token points were not attempted because the declared allocation increment, physical-memory floor, and headroom exceeded the observed admission envelope.
    5. Turbo and fixed-shortlist paths are retained as semantic-change and negative evidence; neither is treated as an equivalent-generation performance result.
    Read the complete evidence method

    What this result establishes

    The history is synthetic all-zero K/V. These results do not establish language quality, generalisation, energy, whole-process memory, cross-host performance, end-to-end throughput, or superiority.

    • Language quality or generalisation
    • Energy savings
    • Cross-host performance
    • End-to-end training or generation throughput
    • Whole-process memory savings
    • Commercial or architectural superiority

    Inspection and reuse

    Downloads and provenance

    Source data

    Machine-readable summary

    The recorded JSON artifact is staged from the governed public evidence pack during the build. Its currency notice determines whether it may be read as current.

    Open source JSON

    Provenance

    Build and evidence receipt

    Use the public build receipt to compare the deployed site source, evidence digest, claim posture, and inference release state.

    Open build receipt

    Continue reviewing

    Related evidence

    measuredbounded
    Historical
    Observed
    Across 20 measurements per arm, direct block-region counting retained the tested plan identity and reduced mean 65,536-token plan-construction time from 784.49 ms to 7.99 ms on the disclosed host.
    Supports
    Exact plan-identity checks and host-local construction timing for one sparse-attention planning component.
    Does not support
    This is component-level engineering evidence. It does not establish end-to-end speed, model quality, memory or energy savings, cross-host performance, or superiority.
    Next decisive evidence
    Integrate the component into a matched quality-bearing end-to-end evaluation before expanding its performance meaning.

    Decision relevance: research · engineering

    Review evidence
    projectionbounded
    Historical
    Observed
    The outlook turns architecture assumptions into inspectable resource scenarios for planning.
    Supports
    Projected active share, stored parameters, decode bandwidth, optimizer state, and weight-stream requirements.
    Does not support
    All values are projections, not measured deployment results, and depend on the published assumptions.
    Next decisive evidence
    Refresh assumptions from the approved current architecture and keep projections separate from measured results.

    Decision relevance: investment · planning

    Review evidence