A customer found an error nobody had caught
The assistant answered well, but about the wrong product. It sounded correct, and nobody audits what sounds correct. We fix it with structured metadata at ingestion and entity filtering before ranking.
It invents data, answers about the wrong entity, costs more than planned, or the team simply stopped trusting it. We measure it, find why it fails and take it to production quality.
28% get built and never deliver the expected value — RAND · 95% of pilots die on the way, according to MIT · 1 week to know whether to rescue it or rebuild
Built on
Ninety seconds: who we are, how we work and what you get at the end.
A set of real questions with their correct answers, the system as it is, and the failures ranked by impact. That is how you know what to fix first and whether the fix worked.
Example: first you measure how well the system answers today, fix what hurts most, and measure again with the same questions.
100 questions · 100 correct answers
same set, same metric
Swipe to follow the flow →
01
100 real questions from your users with the correct answer validated by your team. It is the part almost no vendor does, because it is days of work with your people and it does not show off in a demo.
02
We distinguish whether the problem is that it does not find the information or that it finds it and uses it badly. Confusing the two is why so many fixes fix nothing.
03
Ingestion, chunking, embeddings, vector store, retrieval, system prompt, model routing and cost per query. Each stage with its finding.
04
Failures ranked by impact, the fix for each with an effort estimate, and a clear recommendation: rescue or rebuild, with the numbers behind it.
01
The assistant answered well, but about the wrong product. It sounded correct, and nobody audits what sounds correct. We fix it with structured metadata at ingestion and entity filtering before ranking.
You can check this in ten minutes: ask it something that is not in your documentation. If it answers as confidently as when it does know, the system has no way to recognize its own limit.
There are opinions about whether the last change made it better or worse, and none of them can be checked against another. That is why the diagnostic starts by building the measurement that never existed.
It started as a cheap experiment and now someone asks every month how much it is running. Almost always the same cause: one configuration, the most expensive one, for every question.
The final symptom and the most expensive, because nobody files a ticket: they just go back to asking the same colleague. Behind it there is usually information loaded by hand once that went stale without anyone noticing.
Formats and guarantee
Reference set, full measurement and a report with the failures ranked by impact and their fix.
1 week
The prioritised fixes, executed, with before-and-after measurement.
3–5 weeks
When rescuing costs more than rebuilding. The diagnostic is deducted.
4–8 weeks
If the diagnostic does not reach at least three actionable findings — each with its fix and its estimated effort — you do not pay. It exists so you do not have to trust us before seeing us work, which is exactly what went wrong last time.
Free if unmet
The reference set and the report stay with you, in open formats: if the fix is executed by your current vendor or your own team, you have something to hold them to and something to verify it with.
Yours, for good
We built and operated the assistant for a loyalty platform serving more than two million active users across ten countries.
We built the full pipeline: document ingestion and normalization, chunking, embedding generation, vector store on Azure AI Search and Pinecone, and retrieval with grounded generation on LangChain.
We held it above 99% availability for three years.
RAG systems →In most cases it can be rescued, because the language model is almost never the problem: how the information was prepared and retrieved is, and that part can be rebuilt without throwing away the rest. The diagnostic exists precisely to answer that with data instead of intuition. If the honest recommendation is to rebuild, we say so and the diagnostic is credited against the new project.
For the diagnostic we need to see the information the system was fed and a sample of real conversations or runs. Production access is not required. An NDA is signed before you send anything, and if your data contains personal information we work on an anonymized sample.
Yes, and it is the most common case. We have worked on implementations built in n8n, Make, subscription chatbot platforms and custom development. The diagnostic is the same because the failures are the same: the technology changes, the design mistakes do not.
It happens often: around 60% of companies wanting to implement AI have undocumented processes and scattered data. AI does not fix disorder, it automates it faster. If that is your case we say so in the report, with what would need organizing first and how much work it represents. We would rather say it than charge you for a rescue that was going to fail.
The diagnostic takes a week. Remediation, three to five weeks depending on scope. But from the first week of remediation there is measurement running, so improvement shows up as a number rather than a feeling. That is the entire point of the process.
Yes. In several cases the role is to diagnose and hand over findings for the team that built it to execute, and then verify they were implemented correctly. You do not need to change vendors to fix the system.
That is an increasingly common reason for engagement. We address it with task routing, context trimming and prompt caching, measuring cost per query before and after. In systems we have optimized, the reduction is usually substantial without touching perceived quality.
Then it is not billed. The guarantee is concrete: if the report does not reach at least three actionable findings — each with its fix and its estimated effort — there is no invoice. In practice we are running little risk, because the five typical symptoms show up almost every time.
By booking a 15-minute call where you describe the symptoms. From that alone we can usually tell you whether it sounds like one of the five typical failures and how serious it looks. No cost and no follow-up pressure.
It gets fixed, it gets built, or it is not worth it. And if the diagnostic does not reach three actionable findings, it is not charged.
We use cookies to improve the user experience. Privacy