← All tools

EVALUATE / Public workbench

Whetstone

Did the change earn its release?

Compare your current and candidate AI system on the same checks. Keep every gain and regression visible.

AVAILABLE IN THIS RELEASE

YOU BRING

A baseline, a candidate, and the same verifier-backed tasks.

YOU GET

Paired outcomes and an explicit PASS, HOLD, or BLOCK decision.

YOUR FIRST RUN

Get from idea
to an inspectable result.

MCP & API guide →
  1. 01

    Start a shared cohort in the Open Promotion Bench.

  2. 02

    Run each task with your baseline and candidate, keeping their configurations fixed.

  3. 03

    Submit both answer sets. Inspect gains, regressions, scope changes, and the resulting receipt.

WHAT THIS DOES — AND WHAT IT DOESN’T

A useful tool has a clear boundary.

The public benchmark uses six procedural tasks and self-attested system identities. A public run is not a private evaluation or proof of general model quality.