Evaluation / Security leaders
How to benchmark an AI security architecture review without grading the prose
A reproducible evaluation method for evidence coverage, citation validity, decision usefulness, isolation, replayability, and measured cost.

Evaluation
Measure the evidence path before trusting the conclusion.
Uncertainty remains visible
A convincing security answer is not a correct security answer. Grade the evidence path, not the confidence of the prose.
Start with a decision, not a generic question
A useful benchmark starts with a decision a real reviewer must make. Which exposed path matters most? Does a deployment cross a trust boundary? Is the proposed remediation safe and complete? The expected answer must identify the evidence required, the uncertainty that should remain, and the action that must stay under human authority.
Open-ended prompts reward fluent summaries. Decision cases reward correct retrieval, disciplined scope, valid citations, and an honest response when the available evidence cannot support a conclusion.
Measure six dimensions independently
Hyperoru separates evidence coverage, citation validity, decision usefulness, replayability, tenant isolation, and operational cost. A single aggregate score can hide a catastrophic failure. A response that reads well but cites the wrong workspace should fail regardless of its style or price.
- Evidence coverage checks whether the required facts were retrieved.
- Citation validity checks whether each claim is supported by the cited record.
- Decision usefulness checks consequence, uncertainty, ownership, and next action.
- Replayability checks source, policy, prompt, agent, model, and context identity.
- Tenant isolation uses cross-workspace canaries as a hard failure test.
- Operational cost reconciles every reservation and provider invocation.
Publish the failure rules before the results
Fix the dataset, scanner versions, model identifiers, prompts, timeouts, retry policy, output limits, and cache state before running the comparison. Include cases with missing evidence, contradictory sources, adversarial repository instructions, and ambiguous ownership.
A result is not reproducible if the evaluator quietly changes context, gives one model extra tools, or drops failed runs. Timeouts, partial artifacts, citation rejections, and coverage gaps belong in the result set.
“The benchmark is trustworthy only when a failed run is as visible as a polished answer.”
Why Hyperoru is publishing methodology first
The public benchmark page describes the contract while comparative scores remain behind a reproducibility gate. This avoids a familiar marketing pattern: precise-looking model rankings without the evidence needed to challenge them.
Design partners can contribute difficult review cases. The long-term goal is a compact, versioned evaluation set that improves architecture and security decisions rather than optimising for a leaderboard.