← Back to the work Independent R&D · Cairn · Evaluation infrastructure

Cairn: a provider-neutral evaluation harness

Comparing AI work across providers is difficult when evidence, quality, latency, retries, usage, and credits arrive in different shapes or disappear into a transcript.

Cairn's evaluation harness replays the same versioned task fixture against normalized receipts from Codex, Claude Code, and OpenClaw-style runtimes. It compares verified outcome rate, evidence coverage, observed latency, retries, and fully priced credit equivalents without making paid model calls.

The harness is intentionally not a model-quality oracle. It reports the evidence present in each receipt; missing evidence, pricing, timing, or retry counts remain unknown or null. A result is verified only when the receipt states that outcome and includes passing evidence. Provider adapters emit a common result shape through the shared evaluation harness, with deterministic fixtures, tests, documentation, and a frge eval CLI.

25 tests passed. The durable local foundation is ready for redacted live receipts from each runtime and governed comparison reports in the hosted dashboard.

Working on something like this?

I'm interested in full-time remote roles where I can own the connective tissue between demand, product, sales, data, and revenue.