Batuta Manifesto
One sentence
Batuta is the open measurement layer for Agent Skills: it measures whether a skill actually works, at what cost, on which model — and publishes everything, immutable, zero-profit.
It's not the 27th router on the market. It's the judge — and it works with any router.
Why this exists
Dozens of skill routers exist today. None of them publish data. Each measures its own goal with its own ruler, so none can prove it's better than the next one — or that it's good for anything at all.
Meanwhile, the people who actually solve problems in the world — teachers, health workers, agricultural technicians, public servants, bakery owners — almost never have the money for an expensive model, and have no way to know what works.
Batuta delivers what's missing: the proven recipe and the real cost.
The north star
Free public infrastructure so that any person, NGO, or government can use cheap AI with a proven result. Education, health, agriculture, public administration. A global development engine funded by donation. Institutional target: Digital Public Good certification.
The sequencing rule
"If nothing works at point one, the rest doesn't matter."
- Batuta measures code skills and publishes the data — prove it first
- Cheap recipes perform as well as expensive ones — prove it with a number
- Global public infrastructure — the mission, which 1 and 2 buy the right to announce
The manifesto announces the north star. The launch announces only what's proven.
The seven principles
1. Zero profit. Zero. Always. Nobody earns anything — not founders, not contributors. Visible credit is the salary. A contributor's name goes on the portal and in the dataset, never money.
2. Total transparency. Every donation and every expense is public. Every published result is immutable and verifiable by anyone.
3. Privacy. The prompt never leaves your machine — only a hash with a local salt, length, and identifiers. Sending data is explicit opt-in. The local report works 100% offline: the value Batuta gives you isn't held hostage by the upload.
4. No conflict of interest. We don't sell skills, we don't sell models, we don't sell SaaS. The number has no reason to lie. No competitor can say that.
5. Declared bias. Whoever installs Batuta already cares about skills. The sample is voluntary and not representative. The dataset states this plainly, always.
6. A vote isn't a grade. The vote decides what gets tested. The test decides what enters the recipe.
7. Civic scope, never partisan. Systems for education, sport, and communication between people and government — independent. We build them, everyone uses them however they like. Electoral disputes are out of scope.
What guarantees the number isn't just talk
The judge is blind. It doesn't know whether the skill fired. If it knew, it would confirm whatever we want to hear, and the number would die right there.
The judge is not the defendant. A model never judges its own output. Cross judgment, always.
The judge is versioned. Model, version, and the full prompt are recorded alongside every verdict. A judge that changes without a changelog invalidates the historical series.
A control group exists. In 5% of turns the router deliberately stays silent. Without this, you're measuring correlation and calling it causation. The holdout is declared plainly to the user, is configurable, and can be turned off — a hidden experiment destroys the project.
The chain cannot be edited. Every published result carries the previous one's hash. Altering an old record breaks the chain in front of everyone. The top is stamped outside our control, periodically, with OpenTimestamps.
What Batuta admits about itself
Credibility is the only product. One hallucinated number or one edited record kills the whole project, and no good intention fixes that afterward.
So: when we decline or limit something, the reason gets written down. When we publish, the raw data comes with it.
Batuta · zero profit · https://batuta.space · code under MIT license