open measurement for agent skills

Opinions are plentiful. Evidence is scarce.

Batuta is an open measurement layer for Agent Skills: local routing, an explicit control group, and public results tied to the raw evidence.

Measured baseline

0skills with publishable sample size
91 msindexing 506 skills
136 ms50 routes, startup included
15 / 15conformance tests passing

Local benchmark recorded at audited baseline 9fa7471 on August 24, 2026. These are provenance-bound baseline measurements, not ecosystem results or current-device guarantees.

Batuta Zero

A funnel, not a star rating

01

The router is the reference, not the product

A dependency-free Rust binary indexes installed skills and ranks matches locally with BM25. The measurement layer is designed to work with other routers too. Silence is a valid result when the match is weak.

02

Causality needs a control arm

A declared holdout keeps the router silent on a configurable share of turns. Comparing outcomes with and without a suggestion is more useful than confusing correlation with lift.

03

No result without its denominator

Published rankings require multiple installations and carry counts, model context, protocol version, and stated sample bias. An empty ranking is more credible than a polished anecdote.

open source · reproducible

Start with the evidence

The repository contains the protocol, frozen battery, schemas, implementation, and the commands used to reproduce the measured baseline.

LAB · phase 0

What the LAB is measuring now

Phase zero publishes bounded aggregates from trusted LAB runs. Prompts, secrets, signatures, receipts and project content never appear here.

Loading verified LAB aggregates…