FDE PulseFDE jobs open 448New in 7 days 29Companies hiring 52Remote-friendly 25%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Analysis

OpenAI, Anthropic, Gemini or open-weight for a client: where the data goes matters more than the leaderboard

At a client company, the "best" model is rarely the one that gets chosen. The chosen model is the one compliance will sign off on, deployed so that switching it means editing a config file.

OpenAI, Anthropic, Gemini or open-weight for a client: where the data goes matters more than the leaderboard
Photo: Taylor Vick / Unsplash

In brief

  • According to a mid-2025 Menlo Ventures survey, Anthropic holds 32% of enterprise LLM usage, OpenAI 25% and Google 20%. Market share is still no reason to pick a model for a particular client.
  • For an individual client, the provider is usually decided by data retention, processing region and the cloud contracts already in place.
  • Put the model behind an abstraction layer so the choice stays reversible instead of becoming a long-term commitment.
ShareLinkedInFacebookX
Decision tree. The first question, highlighted in orange: may the data leave the client's infrastructure? If not, use self-hosted open-weight (gpt-oss-120b). If so, ask next: does the client already have a cloud contract? If yes, use Claude or Gemini through that cloud. If not, call the OpenAI or Anthropic API directly and choose the right region. Every branch ends with running evals and then adding a fallback.
Data and contract constraints rule out options before benchmarks come into play. Source: article, based on documentation from OpenAI, Anthropic and Google Cloud.

In mid-2025, Menlo Ventures published a figure that surprised many people: Anthropic accounted for 32% of enterprise LLM usage, ahead of OpenAI and Google (20%). OpenAI had fallen to 25%, half its own share of two years earlier.

Many engineers reading this will draw a simple conclusion: pick whichever provider is in the lead. That is not entirely wrong. But it answers the question “what is the market using?”, while the real question at a client company is which model is allowed to run on their data.

If you want to work as an FDE, learn this early. At a client company, model choice is decided far more by where the data sits, how long it is kept and which contract it passes through than by benchmarks. Whoever understands these constraints runs the meeting.

Whoever only knows the leaderboard usually ends up listening to the legal team explain things.

What market share tells you, and what it doesn’t

Menlo’s data is still useful. It shows that enterprises choose models on performance rather than price, and that builders tend to pick frontier models over cheaper, faster ones. Your default, then, should be a frontier model, not the cheapest one.

Menlo also found that open-source (open-weight) models account for only 13% of enterprise AI workloads, down from 19% six months earlier. On that number alone, you would conclude that open-weight is losing ground.

But a market-wide average says nothing about the client in front of you. The question worth asking is whether that client has a constraint that forces them onto open-weight.

Remember too that the survey dates from July 2025, so by late 2026 the shares may well have shifted. Use it to understand the trend, but do not cite it as the reason when recommending a model to a specific client.

The first question: how long is data retained?

Imagine you are deploying an agent that reads claims files for an insurance company. The first question compliance asks will almost certainly be how long the provider stores their prompts, not which model is smarter.

OpenAI and Anthropic both have a 30-day figure, but it is not the same mechanism. OpenAI creates abuse-monitoring logs for all API features and by default retains them for up to 30 days. With Anthropic’s standard API, the inputs and outputs themselves are automatically deleted within 30 days of being received or generated.

Those two sentences describe different things, so do not copy “30 days” into a compliance document as if the two were identical. State exactly what is retained. Zero data retention at Anthropic is a separately signed agreement; it does not come with simply creating an API key.

This detail directly affects the timeline. If the client requires that no data be stored for even a day, you cannot promise that in the first week, because the agreement still has to be signed. Put this requirement into the plan from the start so it does not block you on the way to production.

Where is the data processed?

The second question is the processing region. OpenAI lets you set a data residency region for each new project, with options including the US, Europe (EEA and Switzerland) and the UAE. Because the setting applies to new projects, create the project in the right region from the outset rather than planning to move it later.

The client’s cloud has often already chosen for you

This is something many new FDEs overlook: the channel through which a model is bought sometimes matters more than the model itself. Claude is sold through Google Cloud, Amazon Bedrock, Microsoft Foundry and several other channels. If the client already has a contract with a hyperscaler, they can use Claude under that contract without onboarding a new vendor.

For procurement, this can shorten the process considerably. The trade-off is that data handling rules change with the channel: when you use Claude on Google Cloud, data processing is governed by Google Cloud. So the documents you hand to compliance must be Google Cloud’s, not Anthropic’s privacy page.

Then there is the question of endpoints. On Google Cloud, Claude’s regional and multi-region endpoints cost 10% more than the global endpoint, and the global endpoint is recommended as the default. If a client’s bill on the global endpoint is $1,000 a month, switching to regional makes it $1,100. That is not a large sum, but the client needs to know in advance so the invoice is not a surprise.

Gemini on Google Cloud works in a similar way. Requests sent to the global endpoint may be processed in any Google Cloud location worldwide, while jurisdictional endpoints keep ML processing within that region.

If you leave the global endpoint in place simply because it is the default, you may have unwittingly breached a data residency commitment the client has made to its own users.

The table below brings these constraints together. These are the questions to answer before running any benchmark; any cell marked “check” is something you must read up on yourself before the meeting.

Client question OpenAI API Claude (direct API or via cloud) Gemini on Google Cloud Self-hosted open-weight (gpt-oss-120b)
How long is data retained? Abuse-monitoring logs kept for up to 30 days Direct API: inputs/outputs deleted within 30 days; zero retention needs a separate agreement Check Google Cloud documentation Client decides
Where is data processed? Region chosen per project: US, Europe, UAE On Google Cloud: global is the default, regional costs 10% more Global may process anywhere; jurisdictional keeps it in region On the client’s infrastructure
Which contract does it go through? Signed with OpenAI Signed with Anthropic or via Google Cloud, Bedrock, Foundry Signed with Google Cloud Apache 2.0 licence
The cost One more vendor A premium if data must stay in region Choosing the right endpoint Operating an 80GB GPU yourself

When is open-weight the right choice?

For some clients, none of the columns on the left of the table will work: for instance, when data may not leave internal infrastructure, or when the region they need is not on the providers’ supported list. For those clients, open-weight is a serious option, even with only 13% of the market.

OpenAI’s gpt-oss-120b is a concrete example. It is released under the Apache 2.0 licence and runs on a single 80GB GPU such as an NVIDIA H100 or AMD MI300X. In other words, the client needs one GPU server, not a cluster.

The licence matters as much as the hardware. The model page describes Apache 2.0 as free of copyleft restrictions and patent risk. That is a specific argument you can take to the legal team when the client plans a commercial deployment, though the approval decision remains theirs.

The hard part that remains is yours: operating, monitoring and upgrading the model, work the API provider used to do for you.

So propose open-weight only when data constraints genuinely require it.

Don’t tie the system to one provider

The constraints above can change mid-project. The client might sign zero data retention in month three, move to a different cloud, or a new model might appear and outperform the one you are using. If the provider’s name is scattered throughout the code, each of these becomes a refactoring exercise.

LiteLLM solves exactly this problem. It lets you call more than 100 LLMs, including OpenAI, Anthropic, Vertex AI, Bedrock and many other providers, through a single OpenAI format. It also supports fallback and per-team cost tracking, so changing models becomes a config edit rather than a code change.

The process at a client company should therefore follow three steps. First, establish the data and contract constraints to eliminate options that cannot be used. Then run evals on the client’s real data with the remaining frontier models, and put the winner behind an abstraction layer, with fallback to the runner-up.

What should developers practise?

With clients operating across several countries, the question of where data lives tends to come up in the very first meeting. Practise answering it before you answer which model is stronger. Whenever a provider ships an update, reread its data controls page and update your own comparison table.

When reading FDE job descriptions, look for terms such as “data residency”, “compliance”, “multi-cloud” or “Bedrock/Vertex”. They signal that the actual work will revolve around exactly these constraints. On your CV, do not just write “integrated GPT-4”.

Write that you designed a model-calling layer with fallback between two providers and selected endpoints according to the client’s data residency requirements.

A good side project for this is an internal chatbot that runs in three configurations: the OpenAI API, Claude via a cloud provider, and self-hosted gpt-oss-120b. Add a README explaining when to use each one.

That README shows you can read a client’s constraints and pick the right configuration, something a demo that calls an API cannot show.

Market-share rankings will keep shifting every six months. The questions about retention, processing region and contracts will come up again with every client.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
7 sources
Read next on the roadmap · Stage 3: Applied AIMulti-agent systems: parallelise the reading, keep the writing in one threadAnthropic and Cognition published opposite conclusions a day apart. Read the two posts together, though, and the real dividing line is fairly clear: don't ask how many agents you need. Ask which agent is allowed to write.