LiteLLM Proxy: one shared LLM gateway, a virtual key per client, separate budgets and bills
When three clients share one OpenAI account, someone will ask who spent what sooner than you expect. A self-hosted gateway is the cleanest answer.
In brief
- LiteLLM Proxy is a self-hosted server with one OpenAI-compatible endpoint for more than 100 LLM providers, under the MIT licence.
- Each virtual key has its own models, budget and rate limits; set a team_id so several keys draw on one client's shared budget.
- The gateway records the cost of every request by key, user and team. A budget without a reset period never resets.
- 1Application sends a requestCalls the OpenAI-compatible endpoint with a virtual key and a user parameter
- 2Check the virtual keyIs this key allowed to use this model, and is it within its rate limit?
- 3Check the budgetThe budget of the key or the client's team; requests are blocked once the cap is reached
- 4Forward to the providerThe gateway holds the provider key and calls one of more than 100 LLM providers
- 5Record the costEach request's cost is attributed by key, user and team for client billing
The application holds only a virtual key. The gateway checks permissions and budget, then records the cost against each client.
Graphic: FDE Times
The monthly LLM bill usually arrives as a single number. For a Forward Deployed Engineer (FDE) working with three clients, or three departments of one client, that number is close to useless. Nobody knows which part belongs to whom, and nothing stops an agent stuck in a loop at midnight.
LiteLLM Proxy solves exactly that problem. The official documentation describes it as a self-hosted server that gives applications one OpenAI-compatible endpoint for more than 100 LLM providers, as well as MCP tools and A2A agents. The BerriAI/litellm repository on GitHub calls it an open-source AI Gateway, and the package on PyPI is MIT-licensed.
For an FDE, the value is not that it can call many models. It is that the gateway turns cost and access into something you configure per client, before anything goes wrong.
Why not hand out the raw API key?
Giving each application the provider’s API key means losing control from day one. The raw key does not know which application is using it, it has no spending cap of its own, and revoking it takes everything down at once.
LiteLLM’s virtual keys separate that layer. The documentation states that each key has its own list of models, budget and rate limits, and that spend is tracked automatically per key.
The README lists virtual keys, spend tracking, guardrails, load balancing and an admin dashboard as built-in features. Applications hold only virtual keys; the provider’s key stays inside the gateway.
An example: three clients, one gateway
Imagine you work for a services company running chatbots for three clients, A, B and C. Each client has two applications: a chatbot and an overnight job that summarises documents.
A sensible setup is one team per client and one virtual key per application. The LiteLLM documentation lets you attach a team_id to a virtual key so that the key draws on the team’s shared budget; a key without a team_id is a personal key.
Client A’s two applications then spend from a single budget, while costs are still split by key, so you can see how much the overnight job consumes.
The application side needs almost no changes, because the endpoint is OpenAI-compatible. The example below uses a placeholder address and model name:
from openai import OpenAI
client = OpenAI(
base_url="https://llm-gateway.noi-bo.example", # địa chỉ gateway, giả định
api_key="VIRTUAL_KEY_CUA_CHATBOT_KHACH_A",
)
resp = client.chat.completions.create(
model="ten-model-duoc-phep",
messages=[{"role": "user", "content": "Tóm tắt hợp đồng này"}],
user="nguoi-dung-cuoi-1024", # để gateway tính chi phí theo end-user
)
(The comments read “gateway address, placeholder” and “so the gateway attributes cost per end user”; the prompt asks the model to summarise a contract.)
The line to watch is the user parameter. The spend-tracking documentation says costs can be tracked by key, user and team, and that end-user tracking relies on the user parameter in the request. Its stated purpose is billing other teams, customers or users.
The gateway records the cost of every request, so at the end of the month you have figures at three levels: which client, which application, which end user.
A budget is only safe with a reset period
The LiteLLM homepage describes hard budgets at several levels (key, team, org and model) with daily and monthly resets, and blocking once the cap is reached. This is what spares you an early-morning call about an agent looping forever.
There is a small trap, though. The documentation says budgets reset at the end of the period you specify; if you do not specify one, the budget never resets.
Suppose you set a cap for client B but forget the period. The first month runs smoothly. When the cap is reached, every key in team B is blocked, and stays blocked until someone fixes it by hand.
So the first thing to do when creating a team for a client is to set the reset period explicitly, then let a key hit its cap in a test environment to see what error the application receives and what it shows the user.
Limits to raise with the client up front
Self-hosting is both a strength and a burden. The gateway sits on the path of every LLM request, so it is your infrastructure or the client’s: someone has to run it, monitor it and upgrade it.
The release cadence is also very fast: the latest litellm release on PyPI is 1.104.2, published on 8 October 2026. At a client site, pin a fixed version and upgrade only after rerunning your own test suite.
Finally, per-end-user figures are only accurate if every application passes the user parameter consistently. That is a discipline for the application teams; the gateway cannot guess it.
What to learn first, and what to put on your CV
A sensible order: get the proxy running with one provider, create a virtual key restricted to certain models, then move on to teams, budgets with reset periods and cost reports by user. Load balancing and guardrails can wait until the cost problem is under control.
When reading FDE job descriptions, phrases such as “multi-tenant”, “cost attribution” or “LLM gateway” signal that the work will touch exactly this area. On a CV, a concrete line such as “built a self-hosted LLM gateway, separating budgets and costs for 3 clients with virtual keys and teams” says more than any list of frameworks.
Clients rarely ask which model you use. They ask how much this month cost, and why. With a gateway, you can answer that with data rather than guesswork.