Solution · bots and AI features
A reference set for AI features: no blind changes to models or prompts
An AI feature is easy to break with a model swap or a prompt edit — and to find out from users. The solution is a reference set of labelled examples and a run that measures precision, recall, latency and cost before every change. If the metrics get worse, the change does not ship.
What it looks like
Mock-ups with demo data, no real data
🧪Model comparison
| Model | Phrases | Precision / Recall | Latency | Cost |
|---|---|---|---|---|
| gpt-4o-mini | 60 | 1.0 / 1.0 | 764 ms | $0.06 |
| gpt-5.2 | 63 | 1.0 / 1.0 | 1,203 ms | $0.063 |
At a glance
- In production
- In my own product, Bruno
- Fits
- Any AI feature: classification, extraction, search
- Needs
- A few dozen labelled examples from your real work
Result at the client
63
labelled phrases in Bruno’s set: 27 promises and 36 non-promises
2
models compared on one set — by quality, latency and cost
Estimate for your company
≈1–2 h
of manual checking saved on every model or prompt change
Assumption: 60+ examples × 1–2 minutes of manual checking each (our assumption)
This is a calculation, not a promise: plug in your own volumes and the figure changes.
Before
- After a model swap it “seems to work” — until a user hits an error.
- It is unclear whether a new prompt version is better or just different.
- An expensive model is picked “just in case”, although a cheaper one does as well.
What was built
A set from real work
Examples are collected from real data and labelled with the correct answer. The set includes negative examples — cases where the feature must not fire.
Run and metrics
A script runs the whole set and computes precision, recall, F1, average latency and the cost of the run.
A release gate
A change that makes the metrics worse does not ship. Run results are kept, so you can see how quality changed over time.
Stack
- TypeScript
- OpenAI
- JSON fixtures
Where it runs
Questions
Isn’t 1.0 on the set perfect accuracy?
No: it is the result on a small set, and we use it as a gate, not a promise. Its job is to catch regressions before users do.
How many examples are needed?
A few dozen from your real work to start, including hard and borderline cases. The set grows with every error found.