# LangChain, LlamaIndex, Haystack or direct API calls: choose a framework for the client's problem, not for its popularity

> Even Anthropic advises starting without a framework. So the question at a client site is no longer which library to use, but which part of the work you are willing to hand to someone else's library.

Bản gốc: https://fdetimes.net/en/analysis/choosing-agent-framework-langchain-llamaindex-haystack/

In its guide to building agents, published on 19 December 2024, Anthropic offered advice that runs against what many developers do by reflex: start by calling the LLM API directly. The reason is practical. Many patterns need only a few lines of code.

Every framework, meanwhile, makes its own promise. Haystack describes itself as an open-source framework for building production-ready agents. LangChain promises to get agents running quickly with any model provider. LlamaIndex invites you to build agents on your organisation's own data.

For an FDE, choosing a framework is not a matter of taste. The code you write stays in the client's codebase after you leave, and it is the client's team that has to debug it at two in the morning.

The right question, then, is not "which framework is best" but "where is the client stuck, and should that part be handed to someone else's library?"

## Each tool solves a different part of the job

Read closely how each vendor describes itself and it becomes clear they are not competing on the same pitch. LangChain focuses on getting started quickly and working with any model provider. When you need a production agent with a degree of determinism, it points you to a separate product, LangGraph, which offers low-level control.

LangChain also has LangSmith, a platform for observing, evaluating and deploying agents. This detail matters to an FDE: part of this ecosystem is a product the client may have to pay for, not just an open-source library. When recommending LangChain, be clear about which parts the client is using and which could bring costs later.

LlamaIndex starts from data: a Python toolkit for building LLM-powered agents that run on the client's own data. Its workflows are event-driven, multi-step processes combining agents, data connectors and tools. RAG is just one tool among them, not the whole system.

Haystack does not hide logic behind automatic layers. According to deepset's documentation, a pipeline is a graph you wire yourself, so it runs the same way every time. It stresses explicit control over retrieval, routing, memory and generation, and promises that you can swap models or infrastructure without rewriting the system.

RAGFlow is a different kind of thing: not a library for assembling code but a ready-built RAG engine. Its strength is deep understanding of unstructured documents with complex formatting. It also lets users see visually how documents are chunked, so that people can intervene.

The table below maps each tool to situations common at client sites:

| Tool | Strong at | Fits when the client needs | Question to ask before bringing it in |
|---|---|---|---|
| LangChain + LangGraph | Fast start, works with many providers; LangGraph for production needing determinism | A quick demo to win over decision-makers, then a move to a controlled flow | Will the client use LangSmith, and who pays? |
| LlamaIndex | Agents on private data, event-driven workflows, many data connectors | Data scattered across many sources | Which connectors are really needed, and which are not yet? |
| Haystack | Developer-wired pipelines, consistent runs, swappable models and infrastructure | Auditability, reproducible results, fear of lock-in | Does the client's team have the people to maintain that graph? |
| RAGFlow | Ready-built engine, understands complex documents, chunking can be inspected and edited | A messy PDF archive, with business users checking retrieval quality | Will they accept running a separate system rather than a library? |
| Direct API calls | Fewest intermediate layers, full view of the prompt | Simple flows, fast debugging | Which parts are you rewriting that a framework already solves well? |

## Where is the client stuck?

Take a hypothetical case: a logistics company wants an internal agent that answers questions about shipping contracts. The data is several thousand scanned PDFs containing tables, drawn up from different templates over the years. You have two weeks.

In the first week, you call the API directly, split documents into fixed-length chunks and print the full prompt before each request. A user asks for the warehousing fee on the Hai Phong–Hanoi route, and the model gets it wrong. The printed prompt (illustrative figures, in the original Vietnamese) looks like this:

```
Ngữ cảnh:
[Đoạn 15] Hải Phòng – Hà Nội | 4.200.000 | 350.000
[Đoạn 31] Phí lưu kho được tính theo tháng, thanh toán cuối kỳ...
Câu hỏi: Phí lưu kho tuyến Hải Phòng – Hà Nội là bao nhiêu?
```

Roughly: "Context: [Chunk 15] Hai Phong – Hanoi | 4,200,000 | 350,000. [Chunk 31] Warehousing fees are charged monthly, payable at the end of the period... Question: What is the warehousing fee on the Hai Phong–Hanoi route?"

The fault is obvious at once. The header row "Route | Shipping fee | Warehousing fee" fell into chunk 14, which was not retrieved, so the model saw two numbers with no column names and picked the first one. The problem lies in reading and chunking documents, not in orchestrating the agent.

That is the moment to bring in a tool strong at document understanding, such as RAGFlow. Its chunking view lets someone on the operations side see that the fee table has been separated from its header and fix it so the whole table sits in one chunk. Print the prompt again and the context now carries the column names; the model picks the correct figure, 350,000.

The lesson for day one at a client site: before discussing frameworks, print the prompt and ask whether the retrieved chunks contain enough information for a human reader to answer correctly.

Other situations lead to other tools. If the documents are clean but scattered across SharePoint, databases and email, LlamaIndex's data layer is what saves you effort.

Suppose instead the client is an audited business that needs every answer to be reproducible. Then Haystack's hand-wired graph, which runs the same way each time, is easier to defend before the compliance department.

If the client's boss just needs a demo in three days to approve a budget, a quick start with LangChain is reasonable, provided you say up front that the production version will need LangGraph or another approach.

**Điểm mấu chốt:** Bring in a framework only after you have called the API directly and know exactly which part of the work it will carry.

## You leave; the client keeps the abstraction

Anthropic identifies the biggest risk: frameworks often add layers of abstraction that obscure the actual prompts underneath, making debugging harder. For a developer building an internal product, that is a minor annoyance. For an FDE, it is a debt left to the client's team.

Anthropic's second piece of advice is blunter: if you use a framework, understand the code underneath. It notes that incorrect assumptions about what is happening inside are a common source of customer errors. In other words, failures often come not from the model but from users believing the framework does one thing when it actually does another.

So every framework needs to pass a test before handover. Can you print the exact final prompt sent to the model, as in the shipping-contract example above? Can the client's team trace a wrong answer back to the exact chunk and routing step that caused it?

If the answer is no, the framework is saving you time by making the client pay later.

## Writing it yourself does not mean rewriting everything

Reading Anthropic's advice as "never use a framework" is a misreading. It says many patterns need only a few lines of code, not all patterns. A tool-calling loop, a summarisation step or a simple router all fit in a few dozen lines that anyone on the client's team can read.

But writing your own parser for PDFs with tables, or your own connectors for ten data sources, means redoing work in which RAGFlow or LlamaIndex have invested heavily. Haystack's promise of "swap models without rewriting" is also only worth something if the client genuinely might swap models.

If not, a thin adapter layer you write yourself is enough to reduce lock-in risk.

The sensible approach at a client site is therefore usually a mix. Keep the orchestration hand-written and thin enough to print every prompt. Bring a framework in at exactly one heavy layer, where it carries real work. At handover, include a page explaining why that layer uses a library and the others do not.

## What to practise

The valuable skill here is not memorising the APIs of all four tools. It is being able to explain, in front of the client's engineers, why you chose this tool for this part of the job. Practise by solving the same problem two ways, calling the API directly and using a framework, then note where the framework helped and where it hid the prompt.

Framework names in job descriptions are also a clue: a role that mentions LangGraph and observability is probably looking for someone to put agents into production under control. When describing your experience, show employers that you weigh why to use or drop a tool, rather than listing library names.

Frameworks will be renamed, merge products and change how they present themselves. The habit of calling the API directly first, finding exactly where the client is stuck and only then choosing a tool will outlast any of them.

**Thử ngay tuần này:**

- Build a question-answering flow over 5 documents by calling the LLM API directly, then rewrite it with Haystack or LlamaIndex, and count how many steps it takes to print the final prompt sent.
- Open an FDE or AI engineer job description, underline the framework names, and write one sentence for each: which layer of work it solves.
- Add a line to your CV describing a time you chose NOT to use a framework (or chose to use one), with a specific technical reason.

## Nguồn

- [Haystack Overview (Haystack Documentation, deepset)](https://docs.haystack.deepset.ai/docs/intro)

- [deepset-ai/haystack (GitHub)](https://github.com/deepset-ai/haystack)

- [LlamaIndex Python Framework documentation (docs.llamaindex.ai/en/stable redirects here)](https://developers.llamaindex.ai/python/framework/)

- [LangChain](https://www.langchain.com/)

- [infiniflow/ragflow (GitHub)](https://github.com/infiniflow/ragflow)

- [Building effective agents (Anthropic)](https://www.anthropic.com/research/building-effective-agents)
