Quality Ledger
Machine-readable: /quality-ledger.json · eval suite: /evals.json · doc: QUALITY_LEDGER.md.
Contract boundary: current signed integrations use POST /receipt/v1, /receipt-v1-test-vector.json, and local receipt-v1 verification. The dated blackbox results and verdict metrics below describe the legacy v2 POST /verify evaluation archive, not a current live run. Current site gates are reported by CI rather than a hard-coded test count.
Engine: raven-solana-live-scan@0.1.0 (vocabulary v4) · legacy v2 keyId rvk_c2997e90215279c2 · archived public blackbox run: 2026-06-05 — 9 pass, 0 fail.
Rules: a failing eval becomes a build order · new capabilities must ship with evals · no undocumented trust expansion. Sanctioned engine-work sources: measured feedback, beta users, explicit coverage gaps, eval failures.
Known limitations: holder beta key-gated · deployer history not yet live · liquidity gap without pool evidence · Raydium v4/CEX custody unadjusted · no price prediction or trading advice, by design.
Visible trust machinery
Current receipt-v1 trust machinery is inspectable: independently authenticated signer pin (/pubkey is cross-check only) · rulesVersion + slot · findings + coverage gaps · payloadHash + receiptId + signerPublicKey + signature · receipt-v1 vector · quality ledger · evidence source registry · consumer decision policy · abuse runbook. ReplayHash/officialAttestationHash/keyId belong to the legacy v2 archive.
Quality is not token consumption
Raven does not measure quality by model used, tokens consumed, agent steps, model calls, lines of generated code, PR count, prompt length, or UI badges. Current quality is: external receipt-v1 artifacts used, local verification rate across distinct axes, finding/gap stability across re-checks, gap resolution, tamper rejection, fail-closed behavior, and contract-gate pass rate. Legacy v2 verdict stability remains a version-labelled historical metric. Full metric lists: /quality-ledger.json. Deterministic receipt verification never depends on an LLM.
Evidence per action, not tokens consumed
The verifier fails closed rather than hanging; agents apply their own timeout and retry policy. Deterministic checks need no LLM calls — do not spend a frontier reasoning model on deciding whether a signed receipt is valid; ed25519 verification is a deterministic function. Pre-spend actions reverify immediately; delayed execution reverifies at execution time. Machine-readable: /quality-ledger.json costLatency.