SERGEY REVIN
← All work

Solution · bots and AI features

A reference set for AI features: no blind changes to models or prompts

An AI feature is easy to break with a model swap or a prompt edit — and to find out from users. The solution is a reference set of labelled examples and a run that measures precision, recall, latency and cost before every change. If the metrics get worse, the change does not ship.

What it looks like

Mock-ups with demo data, no real data

🧪Model comparison

ModelPhrasesPrecision / RecallLatencyCost
gpt-4o-mini601.0 / 1.0764 ms$0.06
gpt-5.2631.0 / 1.01,203 ms$0.063
Bruno runs on one set

At a glance

In production
In my own product, Bruno
Fits
Any AI feature: classification, extraction, search
Needs
A few dozen labelled examples from your real work

Result at the client

  • 63

    labelled phrases in Bruno’s set: 27 promises and 36 non-promises

  • 2

    models compared on one set — by quality, latency and cost

Estimate for your company

  • ≈1–2 h

    of manual checking saved on every model or prompt change

    Assumption: 60+ examples × 1–2 minutes of manual checking each (our assumption)

This is a calculation, not a promise: plug in your own volumes and the figure changes.

Before

  • After a model swap it “seems to work” — until a user hits an error.
  • It is unclear whether a new prompt version is better or just different.
  • An expensive model is picked “just in case”, although a cheaper one does as well.

What was built

A set from real work

Examples are collected from real data and labelled with the correct answer. The set includes negative examples — cases where the feature must not fire.

Run and metrics

A script runs the whole set and computes precision, recall, F1, average latency and the cost of the run.

A release gate

A change that makes the metrics worse does not ship. Run results are kept, so you can see how quality changed over time.

Stack

  • TypeScript
  • OpenAI
  • JSON fixtures

Where it runs

Questions

Isn’t 1.0 on the set perfect accuracy?

No: it is the result on a small set, and we use it as a gate, not a promise. Its job is to catch regressions before users do.

How many examples are needed?

A few dozen from your real work to start, including hard and borderline cases. The set grows with every error found.

Have a similar problem? Get in touch — we start by mapping your processes.

Message me on Telegram