Medellín, Colombia · UTC−5 · Remote operation across Latin America and the United States hello@quarl.co EN · ES
quarl ES

LangChain and LangGraph, from proof of concept to production

The frameworks are not the problem. The problem is what happens when the flow holds state, fails halfway through, and someone has to understand why. That is what LangGraph solves and what we know how to operate.

3 years building on LangChain in production · 2M+ users served by systems built with these tools · 3–5 days for an architecture review

  • LangGraph
  • LangChain
  • LangSmith
  • Durable state
  • Human-in-the-loop
  • Replay
  • Checkpointing
Contact us How we work Free, no pitch · the proposal lands in 48 hours

Built on

  • Anthropic
  • Claude
  • Google Gemini
  • Google Cloud
  • LangChain
  • LangGraph
  • PostgreSQL
  • Qdrant
  • Datadog

An engineering team that has already been on the other side.

Ninety seconds: who we are, how we work and what you get at the end.

Measure before fixing, and measure again after

A set of real questions with their correct answers, the system as it is, and the failures ranked by impact. That is how you know what to fix first and whether the fix worked.

Example: first you measure how well the system answers today, fix what hurts most, and measure again with the same questions.

01Reference set

Real_questions.xlsx100 rows
Expected_answers100
Production_logs30 days
Support_complaints214

100 questions · 100 correct answers

02Current system

finds the chunk61
answers with grounding68
without making it up74

03Failures by impact

  • Answers about the wrong entity19
  • Makes it up when it finds nothing11
  • Chunk cut in half6
  • Old source beats the new one3

04Fix

Tag by entity at ingestion
Explicit refusal without backing
Chunk by section, not by size
Rank by document date

05Re-measure

6191finds the chunk
6894answers with grounding

same set, same metric

Swipe to follow the flow →

What they are

One composes, the other sustains

LangChain is the framework for composing applications on language models: connecting the model to data sources, tools, memory and reasoning chains. It solves building very well.

LangGraph solves what comes after. It models the flow as a graph with durable state: the process can pause waiting for a person’s approval, retry a single step without repeating the previous ones, survive a restart, and replay in full for debugging. A flow without that works right up until the first failure halfway through.

That is why most of what we put into production with agents runs on LangGraph, with LangChain covering the composition pieces.

Services

How we can come in

01 3–5 days

Architecture review

For when a design is already on the table. A second opinion on technical grounds, before committing quarters of work. Report with findings, risks and concrete recommendations.

02

Custom development

We build the system: graph, tools, guardrails, state persistence, evaluation and observability with LangSmith.

03 1–2 weeks

Audit of an existing project

When the system is built but fails, costs too much, or nobody understands why it makes the decisions it makes.

04

Embedded team

We join your team, your process and your repository for a few months, and leave capability behind in your people.

Where these systems break

01

State and checkpointing

If the process does not persist, any failure forces a full redo — and in flows with already-executed actions, that duplicates work or charges.

02

Human interrupt points

Which actions should require approval and do not.

03

Step, time and cost limits

A graph without a ceiling loops. It is the error that burns the most budget.

04

Tool failure handling

What happens when the API the agent calls returns an error or takes too long.

05

Observability

Without step-level traces, debugging is guesswork. LangSmith configured properly changes this completely.

06

Evaluation

A set of cases with expected outcomes running in continuous integration alongside the rest of the tests.

01

An honest opinion about frameworks

Not everything needs LangGraph. For a linear three-step flow with no state and no approvals, a well-written function with direct model calls is simpler to maintain and easier to debug.

We will tell you when the framework is overkill. Adding a heavy abstraction where none is needed is the most common way to turn a simple project into an expensive one.

Related services

AI agents

Agents that do the work, not just talk about it

LangGraph orchestration, durable state and human approval on steps with consequences.

Language models

LLM applications that survive real users

LLM applications with structured output and tool calling.

RAG systems

RAG that answers well — and knows when to stay quiet

Chunking, reranking and hybrid search, evaluated with recall@k and NDCG.

AI consulting

AI consulting for enterprises

Architecture, evaluation criteria and cost per query before writing code.

Key takeaways

Four things before the call

  1. 01

    LangChain composes and LangGraph sustains: the first connects models to data and tools; the second models the flow as a graph with durable state.

  2. 02

    Durable state allows pausing for human approval, retrying a single step without repeating earlier ones, surviving restarts, and replaying a full run for debugging.

  3. 03

    In an audit Quarl reviews the graph, checkpointing, human interrupt points, step and cost limits, tool failure handling, observability with LangSmith, and evaluation.

  4. 04

    Not everything needs a framework: for a linear flow of a few steps with no state and no approvals, direct code is simpler to maintain and cheaper to run.

This is where the data comes in

An AI system is worth what its sources are worth. These are the standard connectors; anything with an API or a database connects the same way, and what has no API is handled by file.

SAP

Enterprise ERP

ERP

Oracle

ERP and database

ERP

NetSuite

Cloud ERP

ERP

Salesforce

CRM and service

CRM

HubSpot

CRM and marketing

CRM

PostgreSQL

Database and pgvector

Databases

A RAG system in production: three years, two million users

We built and operated the assistant for a loyalty platform serving more than two million active users across ten countries.

We built the full pipeline: document ingestion and normalization, chunking, embedding generation, vector store on Azure AI Search and Pinecone, and retrieval with grounded generation on LangChain.

We held it above 99% availability for three years.

RAG systems →

Frequently asked questions

01

What is the difference between LangChain and LangGraph?

LangChain is for composing: it connects models to data, tools and memory. LangGraph is for operating: it models the flow as a graph with durable state, with the ability to pause for human approval, retry a step without repeating earlier ones, survive restarts and replay a full run for debugging. LangChain helps you build it; LangGraph lets you sustain it.

02

Do we need a framework at all?

Not always. For a linear flow of a few steps, with no state and no approvals, direct code is simpler to maintain. The framework starts paying off with branching, state that survives restarts, human approvals, partial retries, or the need to replay runs. Adding the abstraction before you need it makes the project more expensive for nothing.

03

Will you work on a project we already started?

Yes, and it is a significant part of what we do. The audit reviews the graph, state handling, guardrails, cost limits, failure handling and observability, and delivers findings ranked by impact with the fix for each. We can execute those fixes or leave them documented for your team.

04

Can you train our team?

Yes. In the embedded team format we work inside your repository and your process, pairing with your people, with the explicit goal of leaving capability behind. It is slower than doing it externally and handing over, and over the medium term it is far cheaper for you.

05

What is durable state and why does it matter so much?

It means the flow records where it is, so a restart, a network failure or a wait for approval does not force starting over. It matters because in a flow that already executed actions — sent an email, charged a card, created a record — restarting is not just slow: it duplicates real-world effects. It is the main reason LangGraph exists.

06

Do you use LangSmith?

Yes, for tracing and evaluation. It lets you see every step of a run with its inputs and outputs, compare configurations and run evaluation sets systematically. When a client prefers not to depend on an external service, we build equivalent tracing on Datadog or in-house tooling.

07

What about alternative frameworks?

We know CrewAI, AutoGen and the lighter alternatives, and in some cases they are the right call. Our criterion is which one operates better in production: state, debugging, cost control and ecosystem maturity. Today LangGraph wins in most enterprise cases, but it is not an ideological position.

08

Which language do you work in?

Python and TypeScript, depending on what your team uses. LangChain and LangGraph have implementations in both, and we pick the one that lets your people maintain the system, not the one we find more comfortable.

Fifteen minutes. A concrete answer.

It gets fixed, it gets built, or it is not worth it. And if the diagnostic does not reach three actionable findings, it is not charged.

Message on WhatsApp