Tell us where it breaksContact

AI-Based

Where the demo worked, and then the model changed underneath it.


Where it breaks

The parts that never made it into the process document.

None of this is a technology problem, which is why buying technology has not fixed it.

  1. 01

    The demo is the product. It was never tested against a set of cases anyone wrote down, so there is no way to tell whether it is still as good as it was in March.

  2. 02

    Nobody agreed what correct means for this system, so every disagreement about quality is an argument about taste.

  3. 03

    The cost per call was never modelled. The bill arrives monthly and is discussed after it arrives.

  4. 04

    The prompt is the product and it lives in one person’s editor, unversioned, with the working version lost twice already.


What we would build

Narrow, boring, and in your environment.

Scoped against a boundary agreed in writing before anything is built, and built in increments you can evaluate one at a time.

  • An evaluation set that says what correct means for your case, run on every model change and every prompt change.
  • Cost modelled per action before launch, with caps and alerts, and a monthly figure you can check against the provider’s own console.
  • Prompts, tools and boundaries in version control, reviewed like the rest of the code, because they are the rest of the code.

Not negotiable

What the agent is not allowed to do here.

Every piece of this work has a line the automation does not cross, and it is cheaper to agree it now than to discover it in an incident. These are ours in ai-based.

  • A measured drop in quality is reported to you whether or not you asked, and whether or not it is convenient.
  • The model provider is a decision with a cost and a switching cost, and both are written down before it is made.
  • It runs on your accounts. There is no layer of ours between you and the provider.

The six that apply to every engagement, whatever it is, are on the home page, and what an engagement costs is on pricing.


Next

Tell us which of these is your week.

Three weeks of mapping tells you whether any of it is worth automating, and you keep the map whether or not you build with us. If it is not worth doing, we will say so.

Tell us where it breaks