Exploring the validity of AI evaluations and benchmark comparisons.
Project details
Independent verification of AI benchmarks, model comparisons, evaluator results and experimental evidence. Viridian determines what the evidence actually supports, and where the claim goes beyond it.
Viridian Assurance is an independent experimental-assurance service for AI systems. We examine whether the evidence behind a benchmark result, model comparison, evaluator score, regression claim, or release decision actually supports the conclusion being made.
Modern AI teams can produce precise-looking numbers from pipelines whose comparability, provenance, evaluator behaviour, configuration identity, or experimental controls are weaker than the final claim suggests. Viridian audits that evidence chain: what was run, what changed, whether the comparison is valid, whether the evaluator behaved consistently, what confounders remain, and what the maximum defensible claim is.
Our work covers areas including LLM-as-a-judge reliability, benchmark comparability, experimental reproducibility, evaluator integrity, provenance, regression attribution, control validity, configuration drift, and held-out confirmation. Viridian also publishes independent technical research on AI evaluation methodology.
Our public research repository contains Technical Notes on response-order sensitivity in pairwise LLM judging, evaluator-contract stability, and presentation/serialization sensitivity, together with frozen rubrics, analysis plans, integrity records, reproducibility material, and explicit methodological limitations.
Commercially, Viridian provides bounded, asynchronous assurance work rather than generic AI consulting. A customer can submit a benchmark result, model-comparison claim, or experimental evidence package and receive a written determination of what the evidence supports, partially supports, does not support, or leaves unestablished.
Current services include Benchmark Verification, Diagnostic Review, and broader Experimental Audits for higher-stakes evidence packages.
Viridian is built around a simple principle:
We do not sell certainty. We show you exactly what the evidence can justify.
Website: viridianassurance.com
Research: github.com/ViridianAIQualityAssurance/viridian-assurance-research
Comments
0Start the conversation
Share the first comment.