Medellín, Colombia · UTC−5 · Remote operation across Latin America and the United States hello@quarl.co EN · ES
quarl ES

LLM applications that survive real users

Integrating a model takes an afternoon. The work is making the output reliable, the cost predictable and the latency acceptable once there is real volume behind it.

5 providers in production: OpenAI, Anthropic, Google, Azure and Bedrock · 3 years running LLM applications at millions-of-users scale · 2 models per request: a cheap one classifies, a capable one answers

  • Structured output
  • Tool use
  • Task routing
  • Prompt caching
  • Streaming
  • Evaluation
  • Fine-tuning
Contact us How we work Free, no pitch · the proposal lands in 48 hours

Built on

  • Anthropic
  • Claude
  • Google Gemini
  • Google Cloud
  • LangChain
  • LangGraph
  • PostgreSQL
  • Qdrant
  • Datadog

An engineering team that has already been on the other side.

Ninety seconds: who we are, how we work and what you get at the end.

Measure before fixing, and measure again after

A set of real questions with their correct answers, the system as it is, and the failures ranked by impact. That is how you know what to fix first and whether the fix worked.

Example: first you measure how well the system answers today, fix what hurts most, and measure again with the same questions.

01Reference set

Real_questions.xlsx100 rows
Expected_answers100
Production_logs30 days
Support_complaints214

100 questions · 100 correct answers

02Current system

finds the chunk61
answers with grounding68
without making it up74

03Failures by impact

  • Answers about the wrong entity19
  • Makes it up when it finds nothing11
  • Chunk cut in half6
  • Old source beats the new one3

04Fix

Tag by entity at ingestion
Explicit refusal without backing
Chunk by section, not by size
Rank by document date

05Re-measure

6191finds the chunk
6894answers with grounding

same set, same metric

Swipe to follow the flow →

The real problem

The hard part is not calling the API

Anyone can wire a model to a form in an afternoon. What decides whether the product survives is everything else: that the output always has the shape your system expects, that cost per request does not explode as usage grows, that latency stays tolerable.

And that you can switch providers without rewriting the application, and measure whether a change improved or degraded the answers. None of that shows up in the demo, and all of it shows up the day there are real users.

The pieces that make the difference

01

Guaranteed structured output

The model returns JSON that validates against a schema, with retries and repair when it does not comply. Your system never receives something it cannot process.

02

Tool use and function calling

The model queries your APIs and databases instead of answering from memory, with limits and validation on every call.

03

Task-based routing

A cheap model classifies, summarizes and filters; a capable one generates the final answer. It is the single biggest lever on cost without touching perceived quality.

04

Prompt caching and context trimming

Reduces cost and latency in applications with long, repeated instructions.

05

Streaming

The answer starts appearing immediately. It does not reduce real latency but it completely changes the user’s perception.

06

Provider abstraction

Moving from OpenAI to Claude or to a self-hosted model should not cost a quarter of engineering time.

07

Prompt injection defenses

Instruction hierarchy, strict separation between system and user content, and input and output filtering.

01

Decisions

Prompting, RAG or fine-tuning

The most frequent question, and the one that wastes the most money when answered badly.

  • The model does not know your information

    RAG

  • The model does not answer in the right format

    Structured output

  • The tone or style is not yours

    Prompting first, fine-tuning later

  • A very specific domain with its own vocabulary

    Fine-tuning

  • Cost is unsustainable

    Routing and caching

  • Nobody knows whether it works

    Evaluation layer

Formats

Scopes and timelines

01 3–5 weeks

Integration into your product

One LLM-powered feature inside an existing application, with evaluation and cost control.

02 6–12 weeks

LLM-native product

Full application: backend, orchestration, interface, evaluation and observability.

03 2–4 weeks

Cost and latency optimization

Routing, caching, context trimming and before-and-after measurement.

04 3–6 weeks

Fine-tuning

Dataset preparation, training, evaluation against the baseline.

Related services

RAG systems

RAG that answers well — and knows when to stay quiet

Chunking, reranking and hybrid search, evaluated with recall@k and NDCG.

AI agents

Agents that do the work, not just talk about it

LangGraph orchestration, durable state and human approval on steps with consequences.

LangChain and LangGraph

LangChain and LangGraph, from proof of concept to production

Implementation and observability instrumented with LangSmith.

AI consulting

AI consulting for enterprises

Architecture, evaluation criteria and cost per query before writing code.

Key takeaways

Four things before the call

  1. 01

    Integrating a language model takes an afternoon; what decides whether the product survives is structured output, cost control and latency at real volume.

  2. 02

    Task-based routing — a cheap model classifies and summarizes, a capable one writes the final answer — is the single biggest lever on cost without touching perceived quality.

  3. 03

    Choosing between prompting, RAG and fine-tuning moves the budget by an order of magnitude: RAG when the model does not know the information, fine-tuning only for tone or domain vocabulary.

  4. 04

    Quarl works with OpenAI, Anthropic, Google, Azure OpenAI and Amazon Bedrock behind an abstraction layer, so switching providers is configuration rather than a quarter of work.

This is where the data comes in

An AI system is worth what its sources are worth. These are the standard connectors; anything with an API or a database connects the same way, and what has no API is handled by file.

SAP

Enterprise ERP

ERP

Oracle

ERP and database

ERP

NetSuite

Cloud ERP

ERP

Salesforce

CRM and service

CRM

HubSpot

CRM and marketing

CRM

PostgreSQL

Database and pgvector

Databases

A RAG system in production: three years, two million users

We built and operated the assistant for a loyalty platform serving more than two million active users across ten countries.

We built the full pipeline: document ingestion and normalization, chunking, embedding generation, vector store on Azure AI Search and Pinecone, and retrieval with grounded generation on LangChain.

We held it above 99% availability for three years.

RAG systems →

Frequently asked questions

01

Which model should we use?

It depends on the task, and it is almost always several within the same system. For classifying, extracting and summarizing, a cheap model is enough and costs a fraction. For complex reasoning or the final user-facing answer, a capable one. We work with OpenAI, Anthropic, Google, Azure OpenAI and Amazon Bedrock, and the choice is justified with numbers in the proposal, not by fashion.

02

How do you control cost?

Four levers: task-based routing, context trimming so you are not sending everything just in case, prompt caching for long repeated instructions, and hard limits per request and per user. In the systems we have operated, those four levers are what keeps the bill predictable as usage multiplies. We also instrument spend so you can see it without asking.

03

Can we switch providers later?

Yes, and we design for it from the start. The application talks to an abstraction layer rather than directly to one provider’s API, so switching models is configuration plus an evaluation run to confirm quality holds. Without that layer, migrating costs a quarter.

04

What is structured output and why does it matter?

It is forcing the model to return data in an exact shape — JSON that satisfies a schema — instead of free text. It matters because your system has to process the response: if it occasionally returns a field under a different name or a number as a string, the integration breaks in production in ways that are hard to reproduce. We validate against a schema and retry with repair when it does not comply.

05

What is prompt injection and how do you handle it?

It is when someone embeds instructions inside content the model processes — a document, an email, a message — to make it ignore its rules. Mitigated with instruction hierarchy in the system prompt, strict separation between system and user content, input and output filtering, and limiting which tools the model can invoke. In agent systems this stops being optional.

06

When does fine-tuning make sense?

When you need a tone, format or vocabulary that prompting cannot achieve consistently, and you have at least a few hundred high-quality examples. It is not for injecting updatable knowledge — that is what RAG is for. In practice, nine out of ten cases that arrive asking for fine-tuning are better solved with RAG and better prompting.

07

How do you know an improvement worked?

With a set of real cases and their expected outcomes, run before and after each change. We measure answer quality, cost per request and latency. Without that, "we improved the prompt" is an opinion, and that is exactly how systems degrade without anyone noticing.

08

Is this useful for processing documents at volume?

Yes, and it is one of the clearest-return use cases: extracting data from invoices, contracts or forms in variable formats where rigid templates fail. It combines with schema validation and human review on low-confidence cases. We also cover this from the automation side.

Fifteen minutes. A concrete answer.

It gets fixed, it gets built, or it is not worth it. And if the diagnostic does not reach three actionable findings, it is not charged.

Message on WhatsApp