# Sub-agent, skill or MCP server: where a client's workflow belongs

> An agent can fail at a client site even when the model is perfectly capable. Often the cause is simply that business process knowledge was put in the wrong place: the system prompt, a tool description, or a sub-agent nobody needed.

Bản gốc: https://fdetimes.net/en/guides/sub-agent-skill-or-mcp-server/

Imagine you are in your second week at a health insurer. The head of claims has just spent three hours explaining how her team assesses a claim.

First, check whether the policy is still in force. Then reconcile the hospital invoices, look up the exclusion code table, and, if the claim exceeds a certain threshold, escalate it.

You now have a notebook full of notes and an agent that needs to learn all of it. Before writing the first line of code, you have to answer one question: where should each piece of this knowledge live?

It sounds simple, but put things in the wrong place and the agent burns context, becomes hard to change whenever the client alters its process, and may be able to read data it should never touch. A good FDE distinguishes three containers: the MCP server, the agent skill and the sub-agent.

## Three containers answer three different questions

MCP (Model Context Protocol) is an open standard for connecting AI applications to external systems such as data sources, tools and workflows. The official documentation likens MCP to a USB-C port for AI applications.

You write a connection once, and any application that supports the standard can plug into it. MCP answers the question: which systems can the agent touch?

As Anthropic defines them, Agent Skills are organised folders of instructions, scripts and resources that an agent discovers and loads when needed. A skill answers the question: how should the agent do this work? Anthropic sees skills as complementary to MCP: MCP opens up access, while skills teach the agent more complex workflows.

A sub-agent is a specialised assistant. It runs in its own context window, with its own system prompt, its own set of tools and independent permissions. A sub-agent answers the third question: where should this work happen, and with what permissions?

**Điểm mấu chốt:** MCP gives the agent access to systems; a skill teaches it to work the way the client does.

## Taking apart a claim

Go back to the notebook from that hypothetical insurer and sort it piece by piece.

The policy management system and the invoice store are the client's infrastructure, so they belong in an MCP server, with tools such as looking up a policy by number or fetching invoices by claim ID. If the client later wants to switch to a different agent or IDE, this connection can be reused.

The seven assessment steps, the exclusion code table and the payout formula are process knowledge, so they belong in a skill. The folder might look like this (the names are Vietnamese: the skill is "claims assessment", with an exclusion-code table and a payout script):

```
tham-dinh-boi-thuong/
├── SKILL.md            # các bước, ngưỡng chuyển cấp
├── ma-loai-tru.md      # bảng tra cứu dài, chỉ đọc khi cần
└── tinh_chi_tra.py     # script tính toán, không để model tự nhẩm
```

The comments read: the steps and escalation thresholds; a long lookup table, read only when needed; and a calculation script, so the model is not left to do the arithmetic in its head. The top of SKILL.md only needs to be short:

```
---
name: tham-dinh-boi-thuong
description: Dùng khi thẩm định hồ sơ bồi thường y tế, kiểm tra
hiệu lực hợp đồng, đối chiếu hóa đơn và quyết định chuyển cấp duyệt.
---
1. Gọi tool tra cứu hợp đồng, dừng nếu hợp đồng hết hiệu lực.
2. ...
```

The description says, in effect: use this when assessing medical claims, checking policy validity, reconciling invoices and deciding on escalation. Step one: call the policy lookup tool and stop if the policy has lapsed.

The description is the line that matters most. In Claude Code, only a skill's description sits in context so the agent knows which skills it has; the full content loads only when the skill is invoked. Anthropic calls this progressive disclosure (load only what is needed, when it is needed) and treats it as the core design principle of skills.

The last task is scanning a few hundred scanned invoices for duplicates. This work is noisy: it pours out large amounts of file content the main agent will never read again, so it suits a sub-agent:

```
---
name: doi-chieu-hoa-don
description: Rà kho hóa đơn của một hồ sơ, trả về danh sách hóa đơn trùng hoặc bất thường.
tools: Read, Grep, Glob
---
Chỉ đọc. Không sửa, không xóa file. Kết quả trả về dạng bảng ngắn.
```

This "invoice reconciliation" sub-agent scans a claim's invoices and returns duplicates or anomalies. Its instructions: read only, do not edit or delete files, return results as a short table.

The `tools` field works as an allowlist, while `disallowedTools` serves as a denylist. With a client's medical data, this is where you draw the permission boundary in configuration, rather than relying on the model to discipline itself.

The two layers can be connected. When a skill declares `context: fork`, Claude Code starts a sub-agent of the type named in the `agent` field and uses the skill's content as its prompt:

```
---
name: ra-hoa-don-trung
description: Dùng khi cần rà toàn bộ hóa đơn của một hồ sơ để tìm hóa đơn trùng.
context: fork
agent: doi-chieu-hoa-don
---
Lấy danh sách hóa đơn theo mã hồ sơ, so số tiền và ngày khám,
trả về bảng các cặp nghi trùng.
```

Here the skill tells the sub-agent to fetch the invoice list by claim ID, compare amounts and visit dates, and return a table of suspected duplicate pairs.

How to find duplicate invoices is written in the skill, but execution happens inside a sub-agent that has only Read, Grep and Glob. When the client changes its definition of a duplicate, you edit the skill. When the security team wants tighter permissions, you edit the sub-agent. The two changes never collide.

## Context is a budget; spend it well

This split follows from one fact: context costs money. The Claude Code documentation states plainly that, unlike content in CLAUDE.md, a skill's body loads only when it is used, so long reference material costs almost nothing until it is actually needed.

A ten-page exclusion code table in CLAUDE.md eats context on every turn of the conversation. Put in a skill, it takes up space only while the agent is assessing a claim.

MCP has its own cost: preloading definitions for too many tools consumes context. Anthropic has experimented with having the agent write code to call tools instead of loading every definition up front, and in one example token usage fell from 150,000 to 2,000. An MCP server with fifty thin tools, each carrying a long description, is a liability.

The sub-agent is the most expensive option. Anthropic reports that its multi-agent research system uses about 15 times as many tokens as ordinary chat, and that this kind of architecture is a poorer fit for coding tasks with little room for parallelism.

That figure measures a whole multi-agent system rather than a single sub-agent call, but the direction is clear: every separate context is another token bill.

So use a sub-agent only when a side task would flood the main conversation with search results, logs or file contents nobody will read again, or when you need a permission boundary.

## In what order should you work at the client site?

After the discovery session, do not open the editor straight away. Write down everything the client said as separate lines, then ask three questions of each: which system does it need access to, how should it be done, and does it need its own space and permissions? Any line that does not get a "yes" to the third question is, by default, not a sub-agent.

Next, build the MCP server with few tools, each with a clear meaning, and keep process out of the tool descriptions. Only then write the skills. Put the most effort into the description line, because it is the only thing the agent sees when deciding whether to invoke the skill. Write it in the words the client's users actually say, not in your own jargon.

Finally, use real questions from the client's users to test whether the agent calls the right skill and the right tool. When the client changes an approval threshold and you only have to edit one markdown file, that is a sign you packaged things correctly.

## Four common mistakes

The most common mistake is stuffing the entire process into CLAUDE.md or the system prompt. The demo still works, but every new procedure takes up more context on every turn, and by the third month nobody dares touch that file.

The second is writing business logic into the MCP server, for example wrapping a tool such as "assess the whole claim" around dozens of if-else branches. The process is then locked inside the server's code, where the business team can neither read nor change it. Keep MCP thin and let skills carry the process.

The third is splitting everything into sub-agents to look architectural. Every separate context is more tokens to pay for, and the client will see the bill before it sees the value. The fourth is the opposite: creating a sub-agent but letting it inherit every tool, which throws away its biggest benefit, the permission boundary.

## How to put it on a CV so recruiters understand

The claims example above can become a concrete CV line.

Instead of writing "built an agent for an insurance company", write something like: "Split the claims assessment process into a skill made up of SKILL.md, an exclusion code table and a payout script; an MCP server with read-only access to policies and invoices; and a duplicate-invoice sub-agent granted only Read, Grep and Glob."

That line shows the reader you know how to put process, connections and permissions in three different places, and why. If asked about it in an interview, go on to describe how the client changed the approval threshold and you only had to edit one file.

Models will change many times in the coming years. What stays with the client is how you organised their knowledge: processes that can be read, connections that can be reused, and permissions written down explicitly.

**Thử ngay tuần này:**

- Pick a process you know well at work, such as release review or ticket handling. Write out each step, then label each one MCP, skill or sub-agent.
- Write a real SKILL.md for that process, with a description of no more than two sentences. Ask questions the way real users would and see whether the agent picks the right skill on its own.
- Open the CLAUDE.md or system prompt of an agent project in production. Move every block of process instructions longer than ten lines into a skill.

## Nguồn

- [Equipping agents for the real world with Agent Skills](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills)

- [Claude Code Docs: Skills](https://code.claude.com/docs/en/skills)

- [Claude Code Docs: Create custom subagents](https://code.claude.com/docs/en/sub-agents)

- [What is the Model Context Protocol (MCP)?](https://modelcontextprotocol.io/docs/getting-started/intro)

- [Code execution with MCP: building more efficient agents](https://www.anthropic.com/engineering/code-execution-with-mcp)

- [How we built our multi-agent research system](https://www.anthropic.com/engineering/multi-agent-research-system)
