Navigate

Quick find

Navigate NeuroForge

Esc

Start typing to search public pages and evidence.

    Public proof pack

    Evidence organised for decisions, not volume.

    Start with what was observed, then inspect what it supports, what it cannot support, its source boundary, and the next evidence that could change the decision. Historical and superseded records remain visible rather than being mistaken for current proof.

    Evidence maturity: 2 validatedEvidence maturity: 8 provisionalEvidence maturity: 15 candidateEvidence maturity: 7 unprovenEvidence maturity: 12 blocked

    Current aggregate matrix; claim promotion remains held and the detailed wording cards remain the earlier approved snapshot.

    Inspect all 44 technical claim records
    0current
    7historical
    4refresh required
    3superseded or held

    The exact sanitized package is published under a release hold. Its source families have successful CI, but the final assembly commit's jobs did not start because GitHub reported an external billing limit. Raw phase discovery remains 38 pass and 3 fail.

    Find the right evidence

    Filter the public proof pack

    Showing all 14 evidence packs

    Measured evidence

    8 evidence packs
    measuredreviewed
    Historical
    Observed
    The approved runs expose measurable resource behavior with source rows and bounded interpretation.
    Supports
    Public aggregate memory and smoke-loss guardrail curves across the approved experimental scales.
    Does not support
    These curves do not establish dense-baseline superiority or production performance.
    Next decisive evidence
    Repeat under a frozen current stack with matched controls before any broader comparison or performance use.

    Evidence revision

    Decision relevance: research · diligence

    Review evidence
    measuredbounded
    Historical
    Observed
    Across 20 measurements per arm, direct block-region counting retained the tested plan identity and reduced mean 65,536-token plan-construction time from 784.49 ms to 7.99 ms on the disclosed host.
    Supports
    Exact plan-identity checks and host-local construction timing for one sparse-attention planning component.
    Does not support
    This is component-level engineering evidence. It does not establish end-to-end speed, model quality, memory or energy savings, cross-host performance, or superiority.
    Next decisive evidence
    Integrate the component into a matched quality-bearing end-to-end evaluation before expanding its performance meaning.

    Evidence revision

    Decision relevance: research · engineering

    Review evidence
    measuredbounded
    Historical
    Observed
    Tiny and Small completed the declared ladder through 524,288 tokens, while Medium completed every safely admitted point through 131,072; 4,991 metric rows closed with zero observatory errors.
    Supports
    A replicated current-host context ladder across three CPU-only synthetic candidates, with cache-representation counts, semantic diagnostics, and admission boundaries.
    Does not support
    The history is synthetic all-zero K/V. These results do not establish language quality, generalisation, energy, whole-process memory, cross-host performance, end-to-end throughput, or superiority.
    Next decisive evidence
    Clear final assembly CI, then advance from synthetic all-zero K/V to representative histories and quality-bearing workloads.

    Evidence revision

    Decision relevance: research · investment

    Review evidence
    measuredbounded
    Historical
    Observed
    All 28 donor layers produced balanced expert assignments, K32 matched five frozen engineering cases exactly, and all 112 protected trace units completed across four isolated roles.
    Supports
    Real Qwen3 0.6B fracture, exact K32 engineering reconstruction, and a completed protected four-way trace for material-sparsity alignment.
    Does not support
    The bounded alignment path is executable, but no complete aligned checkpoint is promoted; K8/K16/K24 quality, savings, broader generalisation, and production readiness remain unevaluated.
    Next decisive evidence
    Use the retained negative geometry to test refracture and recovery, then rerun the material-sparsity ladder under quality-bearing controls.

    Evidence revision

    Decision relevance: research · investment · diligence

    Review evidence
    measuredreviewed
    Refresh required
    Observed
    Small controlled probes provide measurement signals suitable for engineering decisions and follow-up experiments.
    Supports
    Step counts, throughput ranges, parameter footprint, and pair-work reduction from bounded probes.
    Does not support
    Probe-scale behavior is not a substitute for matched end-to-end production benchmarks.
    Next decisive evidence
    Regenerate from the repaired exact artifact identity and close or retain the phase-41 contract-discovery failure.

    Evidence revision

    Decision relevance: research · engineering

    Review evidence
    measuredpartial
    Historical
    Observed
    The training continuation path is instrumented and produces governed stability evidence.
    Supports
    Loss signals, stability controls, generation gates, reload behavior, and token-scale receipts.
    Does not support
    Generation quality and broad model capability remain under review and are not mature public claims.
    Next decisive evidence
    Run and seal a current exact-resume continuation with frozen evaluation and generation review.

    Evidence revision

    Decision relevance: research

    Review evidence
    measuredpartial
    Held
    Observed
    The integrated training path runs with observable safety and quality signals.
    Supports
    Bounded throughput, loss, safety-counter, warning, and claim-boundary evidence from the full training path.
    Does not support
    Current evidence does not approve commercial product readiness or unrestricted generation claims.
    Next decisive evidence
    Select one exact run and bind its metrics, receipts, model, data, configuration, and review before refresh.

    Evidence revision

    Decision relevance: research · diligence

    Review evidence
    measuredpartial
    Historical
    Observed
    The scale matrix exposes a concrete engineering checkpoint with explicit dimensions and receipts.
    Supports
    Active parameters, evaluation loss, memory peak, and throughput-ratio results at the approved scale point.
    Does not support
    A single bounded scale matrix does not prove performance superiority across hardware or workloads.
    Next decisive evidence
    Repeat the matrix under one frozen current profile and matched control before current-stack use.

    Evidence revision

    Decision relevance: research · engineering

    Review evidence

    Engineering controls

    4 evidence packs
    controlreviewed
    Refresh required
    Observed
    At the recorded boundary, the engineering inventory and exact-reconstruction anchor were visible while sparse alignment remained the active proof gate.
    Supports
    A timestamped repository-scale, evidence-inventory, and nine-claim presentation snapshot retained for historical traceability.
    Does not support
    This snapshot is not the live claim register. Inventory scale and completed engineering gates do not establish material-sparse quality, performance, or production readiness.
    Next decisive evidence
    Regenerate state of play after the claim, phase, fracture, and evidence releases close at one approved exact SHA.

    Evidence revision

    Decision relevance: diligence · research

    Review evidence
    controlreviewed
    Refresh required
    Observed
    Review controls are visible and machine-auditable, including unresolved and held items.
    Supports
    Proof inventory, public claim posture, release checklist, and risk-register status.
    Does not support
    A review-ready evidence pack is not equivalent to approval for public production inference.
    Next decisive evidence
    Rebuild the review packet and release checklist from the next approved evidence release.

    Evidence revision

    Decision relevance: diligence · governance

    Review evidence
    controlreviewed
    Superseded
    Observed
    Critical system surfaces are governed by explicit pass/fail controls rather than narrative assurance alone.
    Supports
    Configuration validation, critical gate coverage, phase-control checks, and API observability posture.
    Does not support
    Control coverage describes engineering discipline; it does not independently prove model capability.
    Next decisive evidence
    Rerun and close—or retain—the three failing raw-discovery validations, then publish an exact-SHA quality-gate pack with executed assembly CI.

    Evidence revision

    Decision relevance: governance · diligence

    Review evidence
    controlbounded
    Held
    Observed
    The package records 38/41 raw phase discovery, 44 reconciled claim rows, a guarded 2,048K synthetic CPU candidate, and two component-geometry failures that narrow the next refracture and recovery experiments.
    Supports
    An exact-source, sanitized public bundle joining current phase and claim reconciliation with a 2M synthetic context candidate and two retained sparse-geometry negatives.
    Does not support
    Source-family CI passed, but final assembly CI did not start because of an external billing limit. The package does not promote language quality, sparse serving, end-to-end performance, production readiness, or commercial superiority.
    Next decisive evidence
    Restore CI availability and execute the exact final package commit, while retaining the three raw phase-discovery failures and all negative scientific results unless they genuinely close.

    Evidence revision

    Decision relevance: research · investment · diligence · governance

    Review evidence

    Projections and diligence

    2 evidence packs
    projectionbounded
    Historical
    Observed
    The outlook turns architecture assumptions into inspectable resource scenarios for planning.
    Supports
    Projected active share, stored parameters, decode bandwidth, optimizer state, and weight-stream requirements.
    Does not support
    All values are projections, not measured deployment results, and depend on the published assumptions.
    Next decisive evidence
    Refresh assumptions from the approved current architecture and keep projections separate from measured results.

    Evidence revision

    Decision relevance: investment · planning

    Review evidence
    controlreviewed
    Refresh required
    Observed
    Diligence artifacts and boundary controls are explicitly inventoried rather than implied.
    Supports
    Artifact coverage, claim posture, clean-room controls, security review, and controlled gates.
    Does not support
    Readiness for diligence is not a declaration that all product, legal, or release risks are closed.
    Next decisive evidence
    Rebuild diligence coverage from the next clean approved evidence release and retain every unresolved gate.

    Evidence revision

    Decision relevance: diligence · investment · governance

    Review evidence

    How to read this material

    Status is part of the result.

    Validated, provisional, and candidate are evidence-maturity states with different review boundaries. Blocked means the required evidence does not currently permit an affirmative capability statement.

    Current blocked boundary

    Dense-baseline superiority, production readiness, mature generation quality, material sparse quality, energy benefit, and external score claims remain blocked.

    Data quality and lineage

    One normalized catalog. Currency travels with every result.

    The catalog records each pack's classification, currency, decision relevance, source artifact, metric and chart lineage, method, boundaries, and next decisive evidence. Presentation completeness never implies scientific confidence or current-stack validity.