Medellín, Colombia · UTC−5 · Remote operation across Latin America and the United States hello@quarl.co EN · ES
quarl ES

Your AI answers badly. It is almost never the model’s fault.

It invents data, answers about the wrong entity, costs more than planned, or the team simply stopped trusting it. We measure it, find why it fails and take it to production quality.

28% get built and never deliver the expected value — RAND · 95% of pilots die on the way, according to MIT · 1 week to know whether to rescue it or rebuild

  • Measurable diagnostic
  • Reference set
  • RAGAS
  • Cost analysis
  • Actionable report
Contact us How we work Free, no pitch · the proposal lands in 48 hours

Built on

  • Anthropic
  • Claude
  • Google Gemini
  • Google Cloud
  • LangChain
  • LangGraph
  • PostgreSQL
  • Qdrant
  • Datadog

An engineering team that has already been on the other side.

Ninety seconds: who we are, how we work and what you get at the end.

Measure before fixing, and measure again after

A set of real questions with their correct answers, the system as it is, and the failures ranked by impact. That is how you know what to fix first and whether the fix worked.

Example: first you measure how well the system answers today, fix what hurts most, and measure again with the same questions.

01Reference set

Real_questions.xlsx100 rows
Expected_answers100
Production_logs30 days
Support_complaints214

100 questions · 100 correct answers

02Current system

finds the chunk61
answers with grounding68
without making it up74

03Failures by impact

  • Answers about the wrong entity19
  • Makes it up when it finds nothing11
  • Chunk cut in half6
  • Old source beats the new one3

04Fix

Tag by entity at ingestion
Explicit refusal without backing
Chunk by section, not by size
Rank by document date

05Re-measure

6191finds the chunk
6894answers with grounding

same set, same metric

Swipe to follow the flow →

One week. You leave with numbers, not opinions.

01

We build your reference set

100 real questions from your users with the correct answer validated by your team. It is the part almost no vendor does, because it is days of work with your people and it does not show off in a demo.

02

We measure retrieval and generation separately

We distinguish whether the problem is that it does not find the information or that it finds it and uses it badly. Confusing the two is why so many fixes fix nothing.

03

We review the full pipeline

Ingestion, chunking, embeddings, vector store, retrieval, system prompt, model routing and cost per query. Each stage with its finding.

04

We deliver the report

Failures ranked by impact, the fix for each with an effort estimate, and a clear recommendation: rescue or rebuild, with the numbers behind it.

01

The five failures

Five symptoms you already recognize

01

A customer found an error nobody had caught

The assistant answered well, but about the wrong product. It sounded correct, and nobody audits what sounds correct. We fix it with structured metadata at ingestion and entity filtering before ranking.

02

It will say anything before saying "I do not know"

You can check this in ten minutes: ask it something that is not in your documentation. If it answers as confidently as when it does know, the system has no way to recognize its own limit.

03

Nobody in your company can show a number

There are opinions about whether the last change made it better or worse, and none of them can be checked against another. That is why the diagnostic starts by building the measurement that never existed.

04

Cost became a meeting topic

It started as a cheap experiment and now someone asks every month how much it is running. Almost always the same cause: one configuration, the most expensive one, for every question.

05

The team stopped using it

The final symptom and the most expensive, because nobody files a ticket: they just go back to asking the same colleague. Behind it there is usually information loaded by hand once that went stale without anyone noticing.

Formats and guarantee

What you hire, and what you keep

  • Diagnostic

    Reference set, full measurement and a report with the failures ranked by impact and their fix.

    1 week

  • Remediation

    The prioritised fixes, executed, with before-and-after measurement.

    3–5 weeks

  • Rebuild

    When rescuing costs more than rebuilding. The diagnostic is deducted.

    4–8 weeks

  • The guarantee

    If the diagnostic does not reach at least three actionable findings — each with its fix and its estimated effort — you do not pay. It exists so you do not have to trust us before seeing us work, which is exactly what went wrong last time.

    Free if unmet

  • What you keep

    The reference set and the report stay with you, in open formats: if the fix is executed by your current vendor or your own team, you have something to hold them to and something to verify it with.

    Yours, for good

A RAG system in production: three years, two million users

We built and operated the assistant for a loyalty platform serving more than two million active users across ten countries.

We built the full pipeline: document ingestion and normalization, chunking, embedding generation, vector store on Azure AI Search and Pinecone, and retrieval with grounded generation on LangChain.

We held it above 99% availability for three years.

RAG systems →

Frequently asked questions

01

Is it worth rescuing or better to start over?

In most cases it can be rescued, because the language model is almost never the problem: how the information was prepared and retrieved is, and that part can be rebuilt without throwing away the rest. The diagnostic exists precisely to answer that with data instead of intuition. If the honest recommendation is to rebuild, we say so and the diagnostic is credited against the new project.

02

Do you need access to our systems?

For the diagnostic we need to see the information the system was fed and a sample of real conversations or runs. Production access is not required. An NDA is signed before you send anything, and if your data contains personal information we work on an anonymized sample.

03

Does this work if another vendor or a no-code tool built it?

Yes, and it is the most common case. We have worked on implementations built in n8n, Make, subscription chatbot platforms and custom development. The diagnostic is the same because the failures are the same: the technology changes, the design mistakes do not.

04

What if the problem is that our information is disorganized?

It happens often: around 60% of companies wanting to implement AI have undocumented processes and scattered data. AI does not fix disorder, it automates it faster. If that is your case we say so in the report, with what would need organizing first and how much work it represents. We would rather say it than charge you for a rescue that was going to fail.

05

How long until we see improvement?

The diagnostic takes a week. Remediation, three to five weeks depending on scope. But from the first week of remediation there is measurement running, so improvement shows up as a number rather than a feeling. That is the entire point of the process.

06

Can you work alongside our current vendor?

Yes. In several cases the role is to diagnose and hand over findings for the team that built it to execute, and then verify they were implemented correctly. You do not need to change vendors to fix the system.

07

What if the problem is cost, not quality?

That is an increasingly common reason for engagement. We address it with task routing, context trimming and prompt caching, measuring cost per query before and after. In systems we have optimized, the reduction is usually substantial without touching perceived quality.

08

What if the diagnostic finds nothing?

Then it is not billed. The guarantee is concrete: if the report does not reach at least three actionable findings — each with its fix and its estimated effort — there is no invoice. In practice we are running little risk, because the five typical symptoms show up almost every time.

09

How do we start?

By booking a 15-minute call where you describe the symptoms. From that alone we can usually tell you whether it sounds like one of the five typical failures and how serious it looks. No cost and no follow-up pressure.

Fifteen minutes. A concrete answer.

It gets fixed, it gets built, or it is not worth it. And if the diagnostic does not reach three actionable findings, it is not charged.

Message on WhatsApp