Integration into your product
One LLM-powered feature inside an existing application, with evaluation and cost control.
Integrating a model takes an afternoon. The work is making the output reliable, the cost predictable and the latency acceptable once there is real volume behind it.
5 providers in production: OpenAI, Anthropic, Google, Azure and Bedrock · 3 years running LLM applications at millions-of-users scale · 2 models per request: a cheap one classifies, a capable one answers
Built on
Ninety seconds: who we are, how we work and what you get at the end.
A set of real questions with their correct answers, the system as it is, and the failures ranked by impact. That is how you know what to fix first and whether the fix worked.
Example: first you measure how well the system answers today, fix what hurts most, and measure again with the same questions.
100 questions · 100 correct answers
same set, same metric
Swipe to follow the flow →
The real problem
Anyone can wire a model to a form in an afternoon. What decides whether the product survives is everything else: that the output always has the shape your system expects, that cost per request does not explode as usage grows, that latency stays tolerable.
And that you can switch providers without rewriting the application, and measure whether a change improved or degraded the answers. None of that shows up in the demo, and all of it shows up the day there are real users.
01
The model returns JSON that validates against a schema, with retries and repair when it does not comply. Your system never receives something it cannot process.
02
The model queries your APIs and databases instead of answering from memory, with limits and validation on every call.
03
A cheap model classifies, summarizes and filters; a capable one generates the final answer. It is the single biggest lever on cost without touching perceived quality.
04
Reduces cost and latency in applications with long, repeated instructions.
05
The answer starts appearing immediately. It does not reduce real latency but it completely changes the user’s perception.
06
Moving from OpenAI to Claude or to a self-hosted model should not cost a quarter of engineering time.
07
Instruction hierarchy, strict separation between system and user content, and input and output filtering.
01
Decisions
The most frequent question, and the one that wastes the most money when answered badly.
RAG
Structured output
Prompting first, fine-tuning later
Fine-tuning
Routing and caching
Evaluation layer
One LLM-powered feature inside an existing application, with evaluation and cost control.
Full application: backend, orchestration, interface, evaluation and observability.
Routing, caching, context trimming and before-and-after measurement.
Dataset preparation, training, evaluation against the baseline.
Chunking, reranking and hybrid search, evaluated with recall@k and NDCG.
LangGraph orchestration, durable state and human approval on steps with consequences.
Implementation and observability instrumented with LangSmith.
Architecture, evaluation criteria and cost per query before writing code.
Key takeaways
Integrating a language model takes an afternoon; what decides whether the product survives is structured output, cost control and latency at real volume.
Task-based routing — a cheap model classifies and summarizes, a capable one writes the final answer — is the single biggest lever on cost without touching perceived quality.
Choosing between prompting, RAG and fine-tuning moves the budget by an order of magnitude: RAG when the model does not know the information, fine-tuning only for tone or domain vocabulary.
Quarl works with OpenAI, Anthropic, Google, Azure OpenAI and Amazon Bedrock behind an abstraction layer, so switching providers is configuration rather than a quarter of work.
An AI system is worth what its sources are worth. These are the standard connectors; anything with an API or a database connects the same way, and what has no API is handled by file.
SAP
Enterprise ERP
Oracle
ERP and database
NetSuite
Cloud ERP
Salesforce
CRM and service
HubSpot
CRM and marketing
PostgreSQL
Database and pgvector
Microsoft SQL
Database
Snowflake
Data warehouse
BigQuery
Google data warehouse
Databricks
Data platform
Redshift
AWS data warehouse
Synapse
Azure data warehouse
Supabase
Managed Postgres
Workday
Payroll and HR
QuickBooks
Accounting
Sage
Accounting and ERP
Xero
Cloud accounting
Shopify
Catalogue and orders
WooCommerce
Catalogue and orders
Magento
Catalogue and orders
Stripe
Payments and subscriptions
Google Drive
Documents and folders
CSV y Excel
Flat files
Nothing in this category
We built and operated the assistant for a loyalty platform serving more than two million active users across ten countries.
We built the full pipeline: document ingestion and normalization, chunking, embedding generation, vector store on Azure AI Search and Pinecone, and retrieval with grounded generation on LangChain.
We held it above 99% availability for three years.
RAG systems →It depends on the task, and it is almost always several within the same system. For classifying, extracting and summarizing, a cheap model is enough and costs a fraction. For complex reasoning or the final user-facing answer, a capable one. We work with OpenAI, Anthropic, Google, Azure OpenAI and Amazon Bedrock, and the choice is justified with numbers in the proposal, not by fashion.
Four levers: task-based routing, context trimming so you are not sending everything just in case, prompt caching for long repeated instructions, and hard limits per request and per user. In the systems we have operated, those four levers are what keeps the bill predictable as usage multiplies. We also instrument spend so you can see it without asking.
Yes, and we design for it from the start. The application talks to an abstraction layer rather than directly to one provider’s API, so switching models is configuration plus an evaluation run to confirm quality holds. Without that layer, migrating costs a quarter.
It is forcing the model to return data in an exact shape — JSON that satisfies a schema — instead of free text. It matters because your system has to process the response: if it occasionally returns a field under a different name or a number as a string, the integration breaks in production in ways that are hard to reproduce. We validate against a schema and retry with repair when it does not comply.
It is when someone embeds instructions inside content the model processes — a document, an email, a message — to make it ignore its rules. Mitigated with instruction hierarchy in the system prompt, strict separation between system and user content, input and output filtering, and limiting which tools the model can invoke. In agent systems this stops being optional.
When you need a tone, format or vocabulary that prompting cannot achieve consistently, and you have at least a few hundred high-quality examples. It is not for injecting updatable knowledge — that is what RAG is for. In practice, nine out of ten cases that arrive asking for fine-tuning are better solved with RAG and better prompting.
With a set of real cases and their expected outcomes, run before and after each change. We measure answer quality, cost per request and latency. Without that, "we improved the prompt" is an opinion, and that is exactly how systems degrade without anyone noticing.
Yes, and it is one of the clearest-return use cases: extracting data from invoices, contracts or forms in variable formats where rigid templates fail. It combines with schema validation and human review on low-confidence cases. We also cover this from the automation side.
It gets fixed, it gets built, or it is not worth it. And if the diagnostic does not reach three actionable findings, it is not charged.
We use cookies to improve the user experience. Privacy