RAG system from scratch
Full pipeline: ingestion and normalization, chunking strategy chosen from the corpus’s real structure, embeddings, vector store, hybrid search with reranking, grounded generation and evaluation.
Retrieval-Augmented Generation connects a language model to your company’s real information. Done right, it answers with verifiable sources. Done wrong, it answers confidently about what it does not know.
2M+ users served by the RAG system we ran in production · 3 years sustaining it — which is where the real failures appear · 10 countries with different catalogs and languages
Built on
Ninety seconds: who we are, how we work and what you get at the end.
Contracts, price lists, email, the CRM: what you already have becomes sourced answers and agents that do the work. Pick a case and follow the path.
Example: a customer asks what their policy covers. The system answers with the exact page of the contract and opens the case.
Connected by API or uploaded. Nothing leaves your infrastructure when the case demands it.
84 pp → 612 chunks
Searches by meaning and by exact word, and filters by the tag before choosing.
answers with source · 91 of 100
Swipe to follow the flow →
What RAG is
A language model knows everything except your company: not your prices, your inventory, your policies or an order’s status. RAG solves that by inverting the order: it searches your information first, then writes the answer using only what it found, and cites where it came from.
The idea is simple; the execution is where it gets decided. A mediocre RAG and a good one look identical in the demo, and separate by month three, when the questions nobody rehearsed start arriving.
01
The most common failure and the most dangerous, because the answer sounds correct and nobody audits it. It happens because "Premium plan pricing" and "Basic plan pricing" are nearly identical in vector space. We fix it by extracting structured metadata at ingestion and filtering by entity before ranking.
02
A model prefers a plausible answer to admitting ignorance. We correct it with grounding constraints, mandatory citation of the source passage, and explicit refusal when retrieval comes back empty.
03
If your vendor cannot show a quality number, it is not that the number is bad: it was never measured. That is where we start: a reference set, and measurement running on every change.
04
Because every question gets the most expensive model with all available context, just in case.
05
Loading was manual, done once, and nobody defined how it gets refreshed.
01
Full pipeline: ingestion and normalization, chunking strategy chosen from the corpus’s real structure, embeddings, vector store, hybrid search with reranking, grounded generation and evaluation.
Reference set built from real questions, retrieval and generation measured separately, and failures ranked by impact with their fix.
When the system exists but retrieves badly: query rewriting, metadata filtering, reranking, chunking strategy adjustment.
Reference set, retrieval and generation metrics, A/B testing between configurations and tracing. So your team can change things without breaking them.
How it is measured
We work with RAGAS and LLM-as-judge alongside the classic retrieval metrics.
recall@k · MRR · NDCG
Groundedness · Faithfulness · Answer relevance
Cost per query · p95 latency
This is not theory. We operated a RAG system for three years over a knowledge base of product, points, promotions and policies, with high similarity between entities, serving more than two million users across ten countries.
The cross-entity contamination failure — number one on the list above — was solved there, with structured metadata at ingestion and pre-ranking filtering. It is the kind of problem that only appears once the system has been in production for months and the catalog has grown.
Retrieval over policy wordings, measured clause by clause.
Clinical and administrative documentation, refusing before guessing.
Every answer with the exact article or clause it came from.
The system is already in production and answers badly. We measure it and determine what to fix.
Key takeaways
RAG — retrieval-augmented generation — inverts the order: it searches the company's own information first and only then writes, citing the source of each claim.
The most frequent RAG failure is answering correctly about the wrong entity, because similar entities sit almost on top of each other in vector space. It is fixed with structured metadata at ingestion and pre-ranking filtering.
Quarl measures retrieval and generation separately: recall@k, MRR and NDCG on one side; groundedness, faithfulness and relevance on the other, with RAGAS and LLM-as-judge.
The track record is a RAG system operated for three years over a catalogue with high entity similarity, serving two million users across ten countries.
An AI system is worth what its sources are worth. These are the standard connectors; anything with an API or a database connects the same way, and what has no API is handled by file.
SAP
Enterprise ERP
Oracle
ERP and database
NetSuite
Cloud ERP
Salesforce
CRM and service
HubSpot
CRM and marketing
PostgreSQL
Database and pgvector
Microsoft SQL
Database
Snowflake
Data warehouse
BigQuery
Google data warehouse
Databricks
Data platform
Redshift
AWS data warehouse
Synapse
Azure data warehouse
Supabase
Managed Postgres
Workday
Payroll and HR
QuickBooks
Accounting
Sage
Accounting and ERP
Xero
Cloud accounting
Shopify
Catalogue and orders
WooCommerce
Catalogue and orders
Magento
Catalogue and orders
Stripe
Payments and subscriptions
Google Drive
Documents and folders
CSV y Excel
Flat files
Nothing in this category
We built and operated the assistant for a loyalty platform serving more than two million active users across ten countries.
We built the full pipeline: document ingestion and normalization, chunking, embedding generation, vector store on Azure AI Search and Pinecone, and retrieval with grounded generation on LangChain.
We held it above 99% availability for three years.
RAG systems →RAG stands for Retrieval-Augmented Generation. Instead of asking the model to answer from memory, the system searches your information first, selects the relevant passages and asks the model to write using only those. ChatGPT on its own does not know your prices or policies: ask it and it answers generically or invents. With RAG the answer comes from your documentation and can be verified because it cites the source.
Almost always RAG, for a practical reason: your information changes. Fine-tuning teaches a model to speak a certain way, not to memorize updatable facts — when prices change you would have to retrain. RAG separates knowledge from the model, so updating means reindexing. Fine-tuning makes sense for tone, format or very specific domains, and sometimes both are combined. We tell you which applies on the first call.
There is no strict minimum, but below roughly twenty FAQs with information that barely changes, RAG is over-engineering: a subscription platform solves that more cheaply. RAG starts paying off when the corpus is large, changes often, has many similar entities, or when a wrong answer carries real cost.
PDF, Word, Excel, HTML, Markdown, plain text, web pages and databases. Also scanned documents through OCR. What takes work is not the format but the structure: tables, documents with many sections and catalogs with variants need different chunking strategies, and that gets decided by looking at your actual corpus, not by default.
Four things combined: grounding constraints in the system prompt, mandatory citation of the passage backing each claim, explicit refusal when search returns nothing relevant, and groundedness measurement over a reference set to catch when the system starts drifting from sources. None of the four is sufficient alone.
It depends on how fast your information changes. For catalogs and pricing we usually schedule daily or event-driven reindexing; for policies and documentation, weekly. What matters is that ingestion is reproducible and has alerts when a source has gone too long without refreshing, because the most common failure mode is not breaking: it is going stale without anyone noticing.
Only if you want it to. We can deploy on Azure OpenAI or Amazon Bedrock inside your own subscription with the vector store in your cloud, so information never leaves your perimeter. In any configuration we use enterprise plans where API data is not used for training, and we filter personal data before sending.
Model consumption plus the vector store. For a mid-market company it usually lands between USD 150 and 900 per month depending on query volume and corpus size. Model routing — a cheap one to classify and summarize, a capable one for the final answer — and context trimming exist precisely to keep that figure predictable as usage grows.
It gets fixed, it gets built, or it is not worth it. And if the diagnostic does not reach three actionable findings, it is not charged.
We use cookies to improve the user experience. Privacy