The model isn't the product. The harness is.
Every business can rent the same models. What separates AI that ships from AI that embarrasses you is the code around it: what it's fed, how its work is graded, what it's allowed to touch, and who signs off. We call ours the Harness. Here's what it does and why we built it.
Every business now has access to the same models. The best one on the market is a login away, priced per token, available to your competitor at exactly the price it's available to you.
So the model isn't the advantage. The harness is.
A harness is the ordinary software wrapped around a model. It decides what the model sees, what it's asked, how its answer gets checked, what it's allowed to change, and who has to say yes before anything reaches a customer. It's the least glamorous part of any AI system, and it's the part that decides whether the thing works on a Tuesday afternoon with nobody watching.
We've been building ours for a year. We just renamed it to what it always was: the Harness.
Four jobs, in order
1. Feed it the right question. A model asked one clever question does worse than the same model asked ten narrow ones. That isn't our opinion. In one published test, a model scored 62.6% at spotting phishing emails when asked "is this phishing?" and 95.0% when the same judgement was broken into five specific questions. The harness is where that decomposition lives, written down once, versioned, and reviewed like any other code.
2. Grade the work before a person sees it. Every piece of work a model drafts gets checked by something that didn't write it. Sometimes that's a stronger model acting as a strict editor. Sometimes it's a small, fast decision model that answers a narrow yes-or-no in a fraction of a second for a fraction of a cent. The point is the same: nothing arrives in front of a human ungraded, and the riskiest work arrives first.
3. Gate what ships. Models draft. People decide. In the Harness, nothing reaches a customer's inbox or a client's live website without a person approving it, and the model can only ever touch what it was explicitly allowed to. A drafted change to a website isn't a rewrite of the file; it's a set of exact edits, each of which has to match the original precisely or the whole draft is refused.
4. Log every decision. Every answer a model gives is stored with the exact version of the questions that produced it. Change a question and the system knows precisely which answers are now stale. Approve or reject a draft and that outcome feeds a running track record for that model on that kind of work. Over time the harness knows which model to trust with what, based on what actually shipped.
What it looks like on real work
The newest thing running on the Harness is a keyword engine: it finds the Google searches a site should be winning and isn't, then drafts the fix.
We pointed it first at a family-run appliance repair company outside Chicago, one of our clients. The harness generated 942 searches that business could plausibly win, from its own list of services, towns and brands. A small decision model judged every one of them (is this a customer, which service, which town, which page on the site answers it, how likely is this person to pay) in 13.6 seconds, for about three and a half cents. Search volume for all 942 came back in one call for nine cents.
Then it checked Google, as a searcher in the business's home town would see it. The site already ranks eighth for the search that matters most to it. Page one, bottom half, which is where the cheapest wins live.
And it caught its own mistake on the first run. Asked whether "hvac technician jobs" was relevant to a heating and cooling company, the decision model said yes, which is true of the topic and wrong for the business: a job seeker is never a customer. We changed one question to ask about the person searching rather than the subject of the search. Job seekers now read as not-a-customer with better than 90% confidence. That fix took minutes because the question lives in one reviewed file, and every affected answer re-ran on its own.
When the engine finds a gap, it doesn't hand over a report. It drafts the page change, has it graded, opens it for review with a preview of the page as it will look, and waits. Publish is one click. Send back takes a note, and the next draft addresses it.
Why this is the goal
The cost of a model call has fallen so far that the expensive part of AI is no longer the model. It's the mistakes. A support reply that promises the wrong thing. A web page that invents a certification the business doesn't hold. A pricing claim nobody approved.
The Harness exists to make those mistakes cheap to catch and impossible to ship quietly. Cheap review means you can check everything instead of a sample. Narrow questions mean you can see exactly where a judgement went wrong. A human gate means the business keeps the final word. A decision log means you can prove afterwards what happened and why.
That's also what makes it portable. Our clients don't care which model drafted their page. They care that it's right, that it's theirs, and that someone accountable said yes. Models will keep changing every few months. The harness is the part that stays.
What we'd tell you if you're putting AI to work
- Stop shopping for a better model first. Write down the questions your business actually needs answered, narrowly, and see how far a cheap model gets.
- Never let a model grade its own work. Separate the writer from the reviewer, even if the reviewer is just a faster, smaller model.
- Make approval a feature, not a bottleneck. One screen, one button, the riskiest item first. If signing off is painful, people stop doing it.
- Log the question with the answer. When something goes wrong, and it will, you want to know which version of your thinking produced it.
The Harness runs our support drafting and our search work today, with more on the way. If you want AI doing real work in your business, and a person still holding the pen, come talk to us.