Evidence coverage
How much of a material claim is supported by accessible, immutable evidence?
Benchmark / Methodology first
Hyperoru evaluates whether an AI-assisted review is supported, useful, isolated, repeatable, and honestly costed.
Evaluation contract
A model cannot compensate for missing evidence, invalid citations, tenant leakage, or an unreplayable workflow. Each dimension has its own hard failures.
How much of a material claim is supported by accessible, immutable evidence?
Does each factual statement point to the exact evidence that supports it?
Can a reviewer understand consequence, uncertainty, and the safest next action?
Can the same source, policy, context, agent, and model inputs reproduce the review?
Can any retrieval, cache, stream, or export cross a workspace boundary?
Do model, scanner, remediation, and assistant costs reconcile to measured invocations?
Protocol
The public benchmark will include source fixtures, expected evidence, adversarial instructions, model and prompt versions, timeouts, retries, and raw structured outputs.
Methodology published. Internal evaluation is active. Comparative scores stay unpublished until the fixtures, model versions, and failure conditions can ship with them.
Results pending reproducibility gateThe best evaluation cases contain missing evidence, conflicting sources, prompt injection, ambiguous ownership, or a decision that is expensive to get wrong.