THE METHOD

Trust the result.
Inspect the path.

A single score cannot tell you whether a system is ready. Cyberelf separates the questions a responsible release needs to answer.

01 / PROVE

SalienceLean

A content-addressed proof receipt tied to an allowlisted Lean template.

02 / EVALUATE

Whetstone

Paired outcomes and an explicit PASS, HOLD, or BLOCK decision.

03 / ADOPT

Foundry

An admission decision and, when permitted, an explicit baseline adoption record.

04 / AUTHORIZE

ScopeWrit

A deterministic DENY, PAUSE, or ALLOW, with a short-lived action ticket and audit trail.

05 / EXECUTE

Backplane

An inspectable execution receipt and explicit capability lifecycle.

PUBLISHED LOCAL EVIDENCE

The gate must
be able to say no.

These are bounded internal experiments and integration checks. They are not customer results or a benchmark of frontier models.

BLOCK

A better aggregate score, three regressions.

On a fresh 20-item local code cohort, the baseline scored 7/20 and the candidate 10/20: six gains, three regressions, eleven ties. Three candidate replays produced identical item-level outcomes.

Qwen2.5 1.5B → Qwen3 8B · 2,048-token completion budget · exact McNemar p = .5078125.

Inspect the published comparison ↗
PASS

A separate cohort earned promotion.

On a fresh 16-item local code cohort, the candidate moved from 6/16 to 13/16: seven gains and zero regressions, with exact p = .015625.

A distinct local comparison. The result does not generalize to every model, task, or deployment.

Inspect both results ↗
INTEGRATION

The release chain closes end to end.

A separate 38-probe cycle moved from 20/38 to 38/38, with 18 gains and zero regressions. This establishes the reference integration path, not general intelligence improvement.

Inspect the reference cycle ↗

The homepage’s HOLD example is illustrative: identical paired answers supply no evidence of improvement. Published evidence above is retained with its original scope; a new website does not make an old experiment a new result.

PRIVATE EVALUATION

Keep the exam in your environment.

The public workbench is for disposable or sanitized inputs. Bring the core into your own environment for sensitive cohorts, answers, and artifacts. Public benchmark publication is opt-in.

Read the evaluation protocol →