Navigate

Quick find

Navigate NeuroForge

Esc

Start typing to search public pages and evidence.

    Measured evidence

    100M scale matrix

    Active parameters, evaluation loss, memory peak, and throughput-ratio results at the approved scale point.

    Recorded finding: The scale matrix exposes a concrete engineering checkpoint with explicit dimensions and receipts.

    On this page

    Currency and next evidence gate

    Decision signals

    Key metrics and controls

    Each value is mapped directly from the recorded summary artifact and retains its experimental, historical, or control context.

    5/5

    Passing variants

    100 steps; sequence length 768.

    25.4%

    ERAIS full active share

    Active share of non-embedding parameters in this variant.

    2.62×

    ERAIS full throughput ratio

    Relative to the matched dense control in this matrix.

    31.6%

    ERAIS full RSS reduction

    Training peak RSS relative to the matched dense control.

    Source-linked visual evidence

    Readable at a glance. Inspectable in detail.

    Charts are generated from normalized source fields. Every visualization includes its exact values in an accessible table.

    Source-linked chart

    Training peak RSS by variant

    Peak resident memory for the five passing 100M-scale matrix variants.

    Finding: ERAIS full reports the lowest peak RSS in this bounded matrix; the comparison remains tied to its disclosed geometry.

    Training peak RSS by variantPeak resident memory for the five passing 100M-scale matrix variants. ERAIS full reports the lowest peak RSS in this bounded matrix; the comparison remains tied to its disclosed geometry.01,3822,7654,1475,529MiBPeak RSSDense controlDense control4,874Dense optimizer controlDense optimizer control3,559Sparse controlSparse control5,120Sparse optimizer controlSparse optimizer control3,508ERAIS fullERAIS full3,331
    View source values
    Training peak RSS by variant; values in MiB
    GroupPeak RSS
    Dense control4,874 MiB
    Dense optimizer control3,559 MiB
    Sparse control5,120 MiB
    Sparse optimizer control3,508 MiB
    ERAIS full3,331 MiB

    Source-linked comparison

    Throughput ratio by variant

    Each variant's training throughput divided by the dense-control throughput in the same matrix.

    Finding: ERAIS full reports 2.62× the dense-control throughput in this specific five-variant matrix.

    Throughput ratio by variantEach variant's training throughput divided by the dense-control throughput in the same matrix. ERAIS full reports 2.62× the dense-control throughput in this specific five-variant matrix.0.000.711.412.122.83× denseThroughput ratioDense controlDense control1.00Dense optimizer controlDense optimizer control1.01Sparse controlSparse control0.92Sparse optimizer controlSparse optimizer control0.92ERAIS fullERAIS full2.62
    View source values
    Throughput ratio by variant; values in × dense
    GroupThroughput ratio
    Dense control1.00 × dense
    Dense optimizer control1.01 × dense
    Sparse control0.92 × dense
    Sparse optimizer control0.92 × dense
    ERAIS full2.62 × dense

    Source-linked chart

    Active non-embedding share

    Minimum active share for each variant's disclosed non-embedding parameter footprint.

    Finding: Only ERAIS full reports a materially reduced active share in this matrix; the other variants remain at 100%.

    Active non-embedding shareMinimum active share for each variant's disclosed non-embedding parameter footprint. Only ERAIS full reports a materially reduced active share in this matrix; the other variants remain at 100%.0.027.054.081.0108.0%Active shareDense controlDense control100.0Dense optimizer controlDense optimizer control100.0Sparse controlSparse control100.0Sparse optimizer controlSparse optimizer control100.0ERAIS fullERAIS full25.4
    View source values
    Active non-embedding share; values in %
    GroupActive share
    Dense control100.0 %
    Dense optimizer control100.0 %
    Sparse control100.0 %
    Sparse optimizer control100.0 %
    ERAIS full25.4 %

    Method

    How to interpret this pack

    1. The matrix compares five variants at scale_100m, 100 steps, and sequence length 768.
    2. Every variant reports pass status and at least 1,700 full-evaluation examples in the approved source rows.
    Read the complete evidence method

    Claim boundaries

    A single bounded scale matrix does not prove performance superiority across hardware or workloads.

    • Bounded 100M local scale-matrix aggregate evidence only.
    • Parameter counts are normalized to M across JSON, charts and tables.
    • Full-eval loss movement is a local likelihood-eval signal, not a general model-quality claim.
    • Throughput ratio is from the local gate metric and is not a service-level or benchmark-leadership claim.
    • No production-readiness, mature generation-quality, dense-superiority, unrestricted reproduction or proven-breakthrough claim is made.
    • Raw paths, invocation details, private configuration, source commits, hashes, checkpoints, prompts, outputs, model weights and mechanism internals are not published.

    Inspection and reuse

    Downloads and provenance

    Source data

    Machine-readable summary

    The recorded JSON artifact is staged from the governed public evidence pack during the build. Its currency notice determines whether it may be read as current.

    Open source JSON

    Provenance

    Build and evidence receipt

    Use the public build receipt to compare the deployed site source, evidence digest, claim posture, and inference release state.

    Open build receipt

    Continue reviewing

    Related evidence

    measuredreviewed
    Historical
    Observed
    The approved runs expose measurable resource behavior with source rows and bounded interpretation.
    Supports
    Public aggregate memory and smoke-loss guardrail curves across the approved experimental scales.
    Does not support
    These curves do not establish dense-baseline superiority or production performance.
    Next decisive evidence
    Repeat under a frozen current stack with matched controls before any broader comparison or performance use.

    Decision relevance: research · diligence

    Review evidence
    projectionbounded
    Historical
    Observed
    The outlook turns architecture assumptions into inspectable resource scenarios for planning.
    Supports
    Projected active share, stored parameters, decode bandwidth, optimizer state, and weight-stream requirements.
    Does not support
    All values are projections, not measured deployment results, and depend on the published assumptions.
    Next decisive evidence
    Refresh assumptions from the approved current architecture and keep projections separate from measured results.

    Decision relevance: investment · planning

    Review evidence