Medellín, Colombia · UTC−5 · Remote operation across Latin America and the United States hello@quarl.co EN ES
All articles

Twelve questions to ask an AI vendor

Every proposal reads the same. Agents, RAG, no hallucinations. These twelve questions separate the vendor who has run something in production from the one who assembled a demo last weekend.

Quarl team 7 min read

Key takeaways

Vocabulary stopped differentiating anyone: "RAG", "agents" and "no hallucinations" are said by the entire market, including agencies reselling no-code tools. Asking for the metric still differentiates.

The question that returns the most information is the simplest: what number will prove the answers are correct? No number means no acceptance criterion.

Always get scope, price, date and acceptance criterion in writing. Without an acceptance criterion there is no way to argue whether something is finished, and that argument always arrives.

System ownership is the most expensive blind spot: several vendors deliver a configuration inside their own platform, not something you can move elsewhere.

What to ask and what you should hear

QuestionGood signBad sign
What metric proves it answers correctly?Names recall@k, groundedness or similar, with a target"We test it and tune it"
Who writes the test questions?Someone from your business, and they are real questionsThe vendor writes them, or there are none
What does the system do when it does not know?Says so, cites a source, escalates with context"It always finds something to answer"
What does one query cost?Gives the figure and projects it to expected volumeHas not calculated it
What is it built on?Names the pieces and explains the choiceAvoids the question, or says "proprietary technology"
Do I keep the code and the prompts?Yes, with repository and deployment docs"It lives in our platform"
What if the model provider raises prices?Spend is measured and a model swap is planned forHas not considered it
How does the information get updated?Scheduled reindexing and review of new contentManually, when somebody flags it
Who answers if it breaks on a Sunday?Written coverage window and response times"We keep an eye on it"
What warranty, and for how long?A stated period and what it coversNo stated period
Can I speak to two of your clients?Gives name, role and phone numberAnonymous testimonials only
What document do I keep at the end?Architecture, decisions, measurement, operating planA slide deck

Ask for the number, not the demo

Every demo works. It is built from the questions the system answers well and run in a controlled environment. What the demo leaves out is the twenty percent of odd questions that arrive in production and decide whether anyone uses the thing twice.

The question that reorders the conversation is this one: out of a hundred real questions from our customers, how many will it answer correctly, and how will you verify that before go-live? A vendor who has shipped to production has a structured answer. One who has not changes the subject to technology.

In a sweep of more than eighty providers in Colombia in August 2026, none sold on evaluation metrics. That does not mean none use them; it means you have to ask explicitly, because it will not be in the proposal.

Five phrases worth taking seriously

  • Almost nobody trains their own model. It usually means a prompt over somebody else API, which is fine, but calling it something else says something about the rest of the proposal.
  • Every language model can invent. What exists is reducing the rate and measuring it; promising zero is promising what cannot be delivered.
  • A self-serve product, maybe. An implementation with integrations, no.
  • Ask whether they mean training or retrieval. They are different things with different costs and risks, and they get conflated in most proposals.
  • True, which is why scope gets closed before signing. If two meetings in there is still no figure, the problem is not the scope.

Frequently asked questions

What should an AI project proposal include?

Closed scope, price, date and acceptance criterion. The acceptance criterion is the part that almost never appears and the only one that lets you argue objectively about whether the delivery is good: which questions the system must answer correctly and how often. It should also state who owns the code and the prompts at the end.

Is it a bad sign when a vendor does not publish prices?

Not on its own, but it changes what you should demand. With no public price, the first meeting should end with a range, not with another meeting. In the Colombian market as of August 2026, fourteen providers publish exact figures and the rest quote after a call. Both work as long as the figure arrives quickly.

Should I hire an agency that builds on n8n or Make?

It depends what you are paying. Those tools are legitimate and for many cases they are the right call: fast delivery, low cost. The problem starts when custom development rates get charged for a two-day configuration, or when the process needs strict validation and cost control the platform does not provide.

How do I verify an AI vendor experience?

Ask for two references with name, role and phone number, and call them. Ask about what went wrong rather than what went well: the answer to that question says more about how the vendor works than any case study. If all you get is anonymous testimonials or screenshots, treat it as marketing material.

How many proposals should I request?

Three, from the same tier. Comparing a USD 1,000 agency proposal against a USD 50,000 consultancy proposal tells you nothing, because they are not selling the same thing. Before requesting quotes, count the integrations you need and decide whether the system will only read or also write: that is what makes proposals comparable.

Book 15 minutes

Tell us what you are building, or what stopped working. You leave the call with a concrete answer: it can be fixed, it can be built, or it isn't worth it.