Medellín, Colombia · UTC−5 · Remote operation across Latin America and the United States hello@quarl.co EN ES
Project rescue

Your AI answers badly. It is almost never the model’s fault.

It invents data, answers about the wrong entity, costs more than planned, or the team simply stopped trusting it. We measure it, find why it fails and take it to production quality.

Technical capabilities
Measurable diagnosticReference setRAGASCost analysisActionable report
73%of enterprise AI projects never reach production
95%of pilots die on the way, according to MIT
1 weekto know whether to rescue it or rebuild

Key takeaways

  • If an AI implementation is underperforming, it is in the majority: 73% of enterprise projects never reach production.
  • The language model is almost never the problem. How the information was prepared and retrieved is — and that part can be rebuilt without discarding the rest.
  • The five failures that show up almost every time: answering about the wrong entity, inventing when it does not know, nobody measuring quality, cost growing faster than usage, and stale information.
  • The diagnostic takes one week and delivers a reference set of real questions, retrieval and generation measured separately, and failures ranked by impact — all in open formats and owned by the client.

Operating track record

FEMSA loyalty platform

Three years operating the system in production — not one delivery and an exit.

Cross-entity contamination

The most frequent failure in RAG architectures, solved with structured metadata at ingestion.

Fixed scope and date

The proposal arrives within 48 hours with fixed scope, price and date. If scope changes, it is quoted separately and approved first.

It is not you

You are in the majority, not the exception

If you commissioned an AI implementation and it is not delivering, you are with 73% of the market. The documented causes are almost always the same: implementation without technical judgment, expectations set wrong at the sale, and vendors delivering demos dressed as products.

The good news is that the language model is almost never the problem. How the information was prepared and retrieved is — and that part can be rebuilt without throwing away the rest.

The five failures

What we almost always find

  • Right answer, wrong thing. The answer sounds correct and nobody audits it. It happens because the system searches by semantic similarity between nearly identical entities. Fixed with structured metadata at ingestion and entity filtering before ranking.
  • It invents when it does not know. Corrected with grounding constraints, mandatory citation and explicit refusal when search comes back empty.
  • Nobody knows whether it answers well. If your vendor cannot show a quality number, it is not that the number is bad: it was never measured.
  • The bill grows faster than usage. The most expensive model gets sent with all available context on every question.
  • The information is stale and nobody noticed. Loading was manual, done once, and nobody defined how it gets refreshed.
The diagnostic

One week. You leave with numbers, not opinions.

  1. We build your reference set50 real questions from your users with the correct answer validated by your team. Without it no measurement is possible, and it is the part no vendor does because it is work.
  2. We measure retrieval and generation separatelyWe distinguish whether the problem is that it does not find the information or that it finds it and uses it badly. Two different diseases with two different treatments — confusing them is why so many fixes fix nothing.
  3. We review the full pipelineIngestion, chunking, embeddings, vector store, retrieval, system prompt, model routing and cost per query. Each stage with its finding.
  4. We deliver the reportFailures ranked by impact, the fix for each with an effort estimate, and a clear recommendation: rescue or rebuild, with the numbers behind it.
What you keep even if you do not continue with us

The reference set and the report are yours, in open formats. If you decide your current vendor or your internal team should fix it, you have something to hold them to and something to verify with.

It is the only way not to end up in the same position again in six months.

Pricing

Diagnostic and remediation

ScopeWhat it includes
DiagnosticReference set, full measurement and a report with prioritized failures and fixes.1 week
RemediationExecution of the prioritized fixes, with before-and-after measurement.3–5 weeks
RebuildWhen rescuing costs more than rebuilding. The diagnostic is credited.4–8 weeks

Frequently asked questions

Is it worth rescuing or better to start over?

In most cases it can be rescued, because the language model is almost never the problem: how the information was prepared and retrieved is, and that part can be rebuilt without throwing away the rest. The diagnostic exists precisely to answer that with data instead of intuition. If the honest recommendation is to rebuild, we say so and the diagnostic is credited against the new project.

Do you need access to our systems?

For the diagnostic we need to see the information the system was fed and a sample of real conversations or runs. Production access is not required. An NDA is signed before you send anything, and if your data contains personal information we work on an anonymized sample.

Does this work if another vendor or a no-code tool built it?

Yes, and it is the most common case. We have worked on implementations built in n8n, Make, subscription chatbot platforms and custom development. The diagnostic is the same because the failures are the same: the technology changes, the design mistakes do not.

What if the problem is that our information is disorganized?

It happens often: around 60% of companies wanting to implement AI have undocumented processes and scattered data. AI does not fix disorder, it automates it faster. If that is your case we say so in the report, with what would need organizing first and how much work it represents. We would rather say it than charge you for a rescue that was going to fail.

How long until we see improvement?

The diagnostic takes a week. Remediation, three to five weeks depending on scope. But from the first week of remediation there is measurement running, so improvement shows up as a number rather than a feeling. That is the entire point of the process.

Can you work alongside our current vendor?

Yes. In several cases the role is to diagnose and hand over findings for the team that built it to execute, and then verify they were implemented correctly. You do not need to change vendors to fix the system.

What if the problem is cost, not quality?

That is an increasingly common reason for engagement. We address it with task routing, context trimming and prompt caching, measuring cost per query before and after. In systems we have optimized, the reduction is usually substantial without touching perceived quality.

How do we start?

By booking a 15-minute call where you describe the symptoms. From that alone we can usually tell you whether it sounds like one of the five typical failures and how serious it looks. No cost and no follow-up pressure.

Book 15 minutes

Tell us what you are building, or what stopped working. You leave the call with a concrete answer: it can be fixed, it can be built, or it isn't worth it.