Navigate

Quick find

Navigate NeuroForge

Esc

Start typing to search public pages and evidence.

    Measured evidence

    Training probe

    Step counts, throughput ranges, parameter footprint, and pair-work reduction from bounded probes.

    Recorded finding: Small controlled probes provide measurement signals suitable for engineering decisions and follow-up experiments.

    On this page

    Currency and next evidence gate

    Decision signals

    Key metrics and controls

    Each value is mapped directly from the recorded summary artifact and retains its experimental, historical, or control context.

    388.6 tok/s

    Overall median throughput

    Across the approved bounded probe rows.

    304.6 tok/s

    Minimum observed throughput

    Minimum in the approved probe set.

    1.58×

    Median pair-work reduction

    Within the bounded probe configuration.

    24.1%

    Trainable parameter share

    125M of 520M parameters.

    Source-linked visual evidence

    Readable at a glance. Inspectable in detail.

    Charts are generated from normalized source fields. Every visualization includes its exact values in an accessible table.

    Ordered source series

    Probe throughput by sequence length

    Minimum and median tokens per second across the approved sequence-length groups.

    Finding: The 1024-token probe group reports higher minimum and median throughput than the 512-token group in this bounded run set.

    Probe throughput by sequence lengthMinimum and median tokens per second across the approved sequence-length groups. Connecting lines show ordered observations and do not imply unmeasured intermediate results. The 1024-token probe group reports higher minimum and median throughput than the 512-token group in this bounded run set.0.0121.0241.9362.9483.9tokens/sMinimumMedian512 tokens: Minimum, 304.6 tokens/s1024 tokens: Minimum, 417.0 tokens/s512 tokens: Median, 340.1 tokens/s1024 tokens: Median, 448.0 tokens/s512 tokens512 tokens1024 tokens1024 tokens

    Connecting strokes aid comparison between published ordered points; they do not represent measurements between those points.

    View source values
    Probe throughput by sequence length; values in tokens/s
    PointMinimumMedian
    512 tokens304.6 tokens/s340.1 tokens/s
    1024 tokens417.0 tokens/s448.0 tokens/s

    Ordered source series

    Pair-work reduction by sequence length

    Median within-probe pair-work reduction for each approved sequence length.

    Finding: The reported reduction is larger in the 1024-token probe group, within the current two-point scope.

    Pair-work reduction by sequence lengthMedian within-probe pair-work reduction for each approved sequence length. Connecting lines show ordered observations and do not imply unmeasured intermediate results. The reported reduction is larger in the 1024-token probe group, within the current two-point scope.0.000.511.021.522.03×Median reduction512 tokens: Median reduction, 1.28 ×1024 tokens: Median reduction, 1.88 ×512 tokens512 tokens1024 tokens1024 tokens

    Connecting strokes aid comparison between published ordered points; they do not represent measurements between those points.

    View source values
    Pair-work reduction by sequence length; values in ×
    PointMedian reduction
    512 tokens1.28 ×
    1024 tokens1.88 ×

    Method

    How to interpret this pack

    1. The public summary groups 30 finite-loss optimizer-step rows across two sequence lengths.
    2. Probe throughput is an engineering signal and must not be generalized to production serving throughput.
    Read the complete evidence method

    Claim boundaries

    Probe-scale behavior is not a substitute for matched end-to-end production benchmarks.

    • Local bounded training-throughput and trainability probe only.
    • Parameter counts are normalized to M.
    • No step-level rows, raw losses, config payloads, probe args, layer names, parameter names, block plans, memory snapshots, source paths, command names, script names, source internals, model weights, checkpoints, raw prompts, or raw datasets are bundled.
    • No dense-baseline superiority, downstream quality, production-readiness, external benchmark reproduction, long-run training stability, or service-level claim.
    • Pair-work reduction is an aggregate local probe signal, not an end-user capability or quality claim.

    Inspection and reuse

    Downloads and provenance

    Source data

    Machine-readable summary

    The recorded JSON artifact is staged from the governed public evidence pack during the build. Its currency notice determines whether it may be read as current.

    Open source JSON

    Provenance

    Build and evidence receipt

    Use the public build receipt to compare the deployed site source, evidence digest, claim posture, and inference release state.

    Open build receipt

    Continue reviewing

    Related evidence

    measuredreviewed
    Historical
    Observed
    The approved runs expose measurable resource behavior with source rows and bounded interpretation.
    Supports
    Public aggregate memory and smoke-loss guardrail curves across the approved experimental scales.
    Does not support
    These curves do not establish dense-baseline superiority or production performance.
    Next decisive evidence
    Repeat under a frozen current stack with matched controls before any broader comparison or performance use.

    Decision relevance: research · diligence

    Review evidence
    measuredpartial
    Historical
    Observed
    The training continuation path is instrumented and produces governed stability evidence.
    Supports
    Loss signals, stability controls, generation gates, reload behavior, and token-scale receipts.
    Does not support
    Generation quality and broad model capability remain under review and are not mature public claims.
    Next decisive evidence
    Run and seal a current exact-resume continuation with frozen evaluation and generation review.

    Decision relevance: research

    Review evidence