Navigate

Quick find

Navigate NeuroForge

Esc

Start typing to search public pages and evidence.

    Engineering controls

    Terminal evidence handoff

    An exact-source, sanitized public bundle joining current phase and claim reconciliation with a 2M synthetic context candidate and two retained sparse-geometry negatives.

    Recorded finding: The package records 38/41 raw phase discovery, 44 reconciled claim rows, a guarded 2,048K synthetic CPU candidate, and two component-geometry failures that narrow the next refracture and recovery experiments.

    On this page

    Currency and next evidence gate

    Decision signals

    Key metrics and controls

    The recorded values, with their controls and scope.

    Hash verified

    Sanitized package integrity

    Exact signed Git object, strict file allowlist, outer hashes, canonical self-manifests, and disclosure scan.

    38/41

    Raw phase discovery

    Three failures remain. The separate 41/41 result checks canonical registry and declared-artifact integrity only.

    44 / 99

    Claims / evidence links

    49 unique hash-registered paths with 0 reconciliation issues; this does not prove the claims.

    2,097,152 tokens

    Largest completed context

    50 replicates per mode on the Tiny CPU synthetic all-zero K/V candidate; 4,096K was refused before allocation.

    82%

    Turbo sequence agreement

    At 2,048K. The semantic-changing turbo path is retained as a diagnostic and is not promoted.

    2,880 / 3,072

    Fine geometry first pass

    Unrecovered donor-derived component geometry; both registered geometries failed at equal active width 1,152.

    Source-linked visual evidence

    What the evidence shows

    Open any chart’s data table to inspect the exact values, or download the CSV.

    Status matrix

    Registry integrity and raw discovery are separate

    The sanitized phase derivative retains both views rather than converting canonical integrity into an all-pass experimental claim.

    Finding: Canonical registry integrity is 41/41, while raw discovery remains 38 pass and three fail.

    Registry integrity and raw discovery are separateThe sanitized phase derivative retains both views rather than converting canonical integrity into an all-pass experimental claim. Canonical registry integrity is 41/41, while raw discovery remains 38 pass and three fail.0204060phasesPassFailCanonical registry integrityCanonical registry int…410Raw discoveryRaw discovery383
    View source values

    Scroll sideways if any columns are out of view.

    Registry integrity and raw discovery are separate; values in phases
    GroupPassFail
    Canonical registry integrity41 phases0 phases
    Raw discovery38 phases3 phases

    Source-linked comparison

    2,048K synthetic decode throughput

    Mean current-host CPU decode throughput across 50 replicates per mode for one 2.473M-parameter Tiny candidate.

    Finding: The semantic-changing turbo mode measured faster, but exact sequence agreement was only 82% and first-step logits were not bitwise equal.

    2,048K synthetic decode throughputMean current-host CPU decode throughput across 50 replicates per mode for one 2.473M-parameter Tiny candidate. The semantic-changing turbo mode measured faster, but exact sequence agreement was only 82% and first-step logits were not bitwise equal.0200400600tok/sFull vocabularyTurbo (not promoted)2,048K context2,048K context357.8441.9
    View source values

    Scroll sideways if any columns are out of view.

    2,048K synthetic decode throughput; values in tok/s
    GroupFull vocabularyTurbo (not promoted)
    2,048K context357.8 tok/s441.9 tok/s

    Source-linked comparison

    Unrecovered component approximation error

    Calibration mean and p95 per-row relative FFN-output L2-squared error against the frozen 0.02 mean / 0.05 p95 gate.

    Finding: Learned and oracle K8 remained above the gate; record-shared oracle K16 improved the result but still failed. Geometry error dominated router error.

    Unrecovered component approximation errorCalibration mean and p95 per-row relative FFN-output L2-squared error against the frozen 0.02 mean / 0.05 p95 gate. Learned and oracle K8 remained above the gate; record-shared oracle K16 improved the result but still failed. Geometry error dominated router error.00.20.40.6relative L2²MeanP95Learned K8Learned K80.2460.459Oracle K8Oracle K80.2190.393Shared oracle K16Shared oracle K160.1450.270Frozen maximumFrozen maximum0.0200.050
    View source values

    Scroll sideways if any columns are out of view.

    Unrecovered component approximation error; values in relative L2²
    GroupMeanP95
    Learned K80.246 relative L2²0.459 relative L2²
    Oracle K80.219 relative L2²0.393 relative L2²
    Shared oracle K160.145 relative L2²0.270 relative L2²
    Frozen maximum0.020 relative L2²0.050 relative L2²

    Method

    How to read this result

    1. The importer reads only eight allowlisted blobs from the exact signed ERAIS commit and tree; it does not read the ERAIS working tree.
    2. Outer file hashes, the package hash list, six canonical JSON self-manifests, contracts, source identities, and a disclosure scan are verified before an atomic site-side promotion.
    3. Successful CI receipts remain attached to the three source evidence families. The final publication commit is explicitly not labelled CI-verified because its jobs did not start under an external billing limit.
    4. Canonical phase-registry integrity and raw phase discovery remain separate signals. Negative component results are preserved as decision evidence, not hidden or softened.
    5. Only aggregate public derivatives are staged. Raw rows, prompts, token identifiers, weights, protected identities, and moving experiments remain outside the public site.
    Read the complete evidence method

    What this result establishes

    Source-family CI passed, but final assembly CI did not start because of an external billing limit. The package does not promote language quality, sparse serving, end-to-end performance, production readiness, or commercial superiority.

    • The final assembly commit is not CI-verified; successful source-family CI does not silently transfer to the package assembly.
    • The 2,048K result uses synthetic all-zero K/V on one CPU host and a 2.473M-parameter Tiny candidate. It does not establish language quality, generalisation, production context support, cross-host speed, whole-process memory, energy, training throughput, or commercial superiority.
    • Turbo changed semantics: 82% exact sequence agreement and non-bitwise-equal first-step logits prohibit equivalence wording.
    • The route-ceiling and oracle-ladder results are unrecovered donor-derived component diagnostics, not language rollouts. They do not establish recovery impossibility, router readiness, serving performance, or production readiness.
    • Claim-to-evidence hash reconciliation establishes registry consistency, not claim truth or promotion authority.
    • Canonical 41/41 registry integrity does not mean all experiments were rerun and does not erase the three retained raw-discovery failures.

    Inspection and reuse

    Downloads and provenance

    Source data

    Machine-readable summary

    The recorded JSON artifact is staged from the governed public evidence pack during the build. Its currency notice determines whether it may be read as current.

    Open source JSON

    Provenance

    Build and evidence receipt

    Use the public build receipt to compare the deployed site source, evidence digest, claim posture, and inference release state.

    Open build receipt

    Continue reviewing

    Related evidence

    measuredbounded
    Historical
    Observed
    All 28 donor layers produced balanced expert assignments, K32 matched five frozen engineering cases exactly, and all 112 protected trace units completed across four isolated roles.
    Supports
    Real Qwen3 0.6B fracture, exact K32 engineering reconstruction, and a completed protected four-way trace for material-sparsity alignment.
    Does not support
    The bounded alignment path is executable, but no complete aligned checkpoint is promoted; K8/K16/K24 quality, savings, broader generalisation, and production readiness remain unevaluated.
    Next decisive evidence
    Use the retained negative geometry to test refracture and recovery, then rerun the material-sparsity ladder under quality-bearing controls.

    Decision relevance: research · investment · diligence

    Review evidence
    measuredbounded
    Historical
    Observed
    Tiny and Small completed the declared ladder through 524,288 tokens, while Medium completed every safely admitted point through 131,072; 4,991 metric rows closed with zero observatory errors.
    Supports
    A replicated current-host context ladder across three CPU-only synthetic candidates, with cache-representation counts, semantic diagnostics, and admission boundaries.
    Does not support
    The history is synthetic all-zero K/V. These results do not establish language quality, generalisation, energy, whole-process memory, cross-host performance, end-to-end throughput, or superiority.
    Next decisive evidence
    Clear final assembly CI, then advance from synthetic all-zero K/V to representative histories and quality-bearing workloads.

    Decision relevance: research · investment

    Review evidence
    controlreviewed
    Superseded
    Observed
    Critical system surfaces are governed by explicit pass/fail controls rather than narrative assurance alone.
    Supports
    Configuration validation, critical gate coverage, phase-control checks, and API observability posture.
    Does not support
    Control coverage describes engineering discipline; it does not independently prove model capability.
    Next decisive evidence
    Rerun and close—or retain—the three failing raw-discovery validations, then publish an exact-SHA quality-gate pack with executed assembly CI.

    Decision relevance: governance · diligence

    Review evidence