# Bedrock, Azure OpenAI, Vertex AI: getting the model into the customer's cloud

> A prototype that runs on your personal API key isn't finished. The real work starts when it has to run inside the customer's account, in their region and under their access controls.

Original: https://fdetimes.net/en/guides/bedrock-azure-openai-vertex-ai-customer-cloud/

Picture your first week at a customer. The demo runs smoothly on your personal API key. Then the customer's security team sends you one sentence: every model call must go through the company's cloud account, in an approved region, using access rights they grant.

At that point the demo needs a near-total rewrite. Nothing is wrong with the model. The demo just wasn't built inside the customer's environment.

This is the biggest difference between a product developer and an FDE. Developers usually get to choose their own stack. An FDE walks into a cloud the customer chose years ago, and has to find their way around it like someone who works there.

Learn the three services below and you won't spend your whole first week asking where to call the model.

## You don't choose the cloud, but you have to know how to get in

On AWS, the entry point is Amazon Bedrock. AWS describes it as a fully managed service that gives businesses access to foundation models from many AI companies. Bedrock supports more than 100 models, from Amazon, Anthropic, DeepSeek, Moonshot AI, MiniMax, OpenAI and xAI.

For an FDE, the most useful detail is that Bedrock offers several API styles: Messages, Responses, Chat Completions, Converse and Invoke. You can keep the Anthropic or OpenAI SDK you already know, or use boto3. For new applications, AWS recommends the `bedrock-runtime` endpoint.

On Azure, Azure OpenAI now lives inside Microsoft Foundry. The model catalogue has two groups: models sold directly by Azure, and models from partners and the community. Beyond the GPT family, it includes model sets from Cohere, DeepSeek, Meta, Mistral AI and xAI.

Microsoft also says plainly that which models are available depends on the region and the cloud type, and Azure Government has its own documentation page.

On Google Cloud, the name you will see most often is Vertex AI. Google's newer documentation calls the platform Gemini Enterprise Agent Platform, a toolset covering the whole ML lifecycle, from building and training to managing models. Older blog posts say "Vertex AI" and newer docs use the other name, but it is the same platform.

| Aspect | AWS | Azure | Google Cloud |
|---|---|---|---|
| Service to look for | Amazon Bedrock | Azure OpenAI in Microsoft Foundry | Gemini Enterprise Agent Platform (formerly Vertex AI) |
| Scope | Access to more than 100 foundation models from many providers | Models sold directly by Azure, plus partner and community models | Tools to build, train and manage ML models |
| What to ask the customer first | Do you use cross-Region inference? | Is the model available in your region and cloud type? | Which regions have you approved for resources and data? |

## Example: moving a prototype onto Bedrock for a logistics company

Say the customer is a logistics company that runs everything on AWS. They want an agent that reads complaint emails and summarises them for customer service staff. Your prototype currently calls the model provider's API directly.

The first step isn't code. On any cloud, the habit to build is opening the console in the customer's account and checking whether the model you need is available in the region they use. Microsoft states this explicitly for Azure; on AWS or Google Cloud, treat it as the safe assumption.

If the model isn't there, the whole plan changes, and the earlier you find out, the cheaper the change.

The second step is cross-Region inference. Bedrock supports two kinds, Global and Geo, which give higher throughput and a lower price per token. That sounds attractive, but requests may then be processed outside the source region, and complaint emails contain customers' names, phone numbers and addresses.

If the customer's policy requires data to stay in one region, you have to ask before turning this feature on.

The third step is access. Don't put access keys in environment variables. Run the application under an IAM role issued by the customer, allowed to call only the model it needs. Bedrock applies this principle thoroughly: even the built-in Web Search tool only fetches data from the outside web when IAM explicitly allows it.

This approach will convince the customer's security team more than any slide deck.

Only now do you get to the code. Calling the model with boto3 through the Converse API is short:

```python
import boto3

client = boto3.client("bedrock-runtime", region_name=REGION)  # region the customer has approved

def generate(prompt: str) -> str:
    resp = client.converse(
        modelId=MODEL_ID,
        messages=[{"role": "user", "content": [{"text": prompt}]}],
    )
    return resp["output"]["message"]["content"][0]["text"]
```

The detail that matters is the `generate()` function. The rest of the application calls only this function and knows nothing about Bedrock. If the customer later acquires a company that runs on Azure, you just write one more adapter for Foundry, and the complaint-summarising logic stays as it is.

That new adapter still has to answer the region question from the start, because on Azure model availability depends on the region and the cloud type, Azure Government included. Switching clouds costs one extra adapter file, but the checks have to be run again from scratch.

Try one more scenario: one of the customer's subsidiaries runs on Google Cloud. The first job is to search the documentation under both names, Vertex AI and Gemini Enterprise Agent Platform, so you don't miss either the old or the new guidance.

Google Cloud has regions across Asia, Australia, Europe, Africa, the Middle East and the Americas, so you ask the subsidiary which regions it has approved, place resources exactly there, and then write a third adapter for the same `generate()` function.

The last step is infrastructure. Don't click through the console by hand to create roles and resources. AWS has CDK, a framework for declaring infrastructure in code and deploying it through CloudFormation. Google Cloud also lets you manage resources through the console, the CLI, client libraries or Terraform; for the subsidiary, Terraform is the natural choice.

When the infrastructure lives in the repo, the customer's team can review it, rebuild it and take it over after you leave.

**Key point:** Is the model available in the customer's region, may data leave that region, and who grants access? Answer all three before you write code.

## Five questions for the kickoff

The example gives you a process that works across all three clouds. First: which cloud does the customer use, and which regions are approved? Google Cloud has regions in Asia, Australia, Europe, Africa, the Middle East, North America and South America, but customers usually allow only a few.

Second: is the model you need in the catalogue for that exact region? Third: who grants access, and what is the minimum permission you need to request? Fourth: which data is sensitive, and may it be processed outside the region?

Last: which IaC tool does the customer deploy with? Knowing this, you write to their standards rather than bringing your own.

## Mistakes that eat the first week

The most common mistake is assuming a model is available in every region. Microsoft states plainly that availability depends on region and cloud type. A demo that works in a US region guarantees nothing about the customer's region.

The second is turning on cross-Region inference because it is cheaper and faster, without asking about the policy on where data is stored and processed. The third is leaving personal credentials in the code. The code still runs, but it will fail the first security review.

The fourth is small but costly: searching for "Vertex AI" in the new documentation without knowing the platform has been renamed, or the other way round. The last is building everything by hand in the console. On handover day, nobody can rebuild the environment you set up.

## Turning this skill into a line on your CV

When reading FDE job descriptions, look for phrases such as "deploy into customer cloud", "Bedrock", "Azure OpenAI" or "IAM". When you see them, the job will look a lot like the example above.

On your CV, don't just write "AWS experience". Be specific: moved an LLM application onto Bedrock in the customer's account, running under a least-privilege IAM role, with infrastructure written in CDK.

That line shows a hiring manager you can build an application that runs reliably in someone else's environment. Anyone can call a model. In an FDE role, what you're paid for is everything that comes after.

**Try this week:**

- Use a personal AWS account to call a model through Bedrock with boto3 and the Converse API, then note which regions that model is available in
- Rewrite your most recent prototype in adapter style: one shared generate() function, one file per cloud
- Draft a checklist of the five questions (region, model, permissions, data, IaC) and use it at your next kickoff with a customer or your team

## Sources

- [What is Amazon Bedrock? (Amazon Bedrock User Guide)](https://docs.aws.amazon.com/bedrock/latest/userguide/what-is-bedrock.html)

- [Foundry Models sold by Azure - Microsoft Foundry | Microsoft Learn](https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)

- [Introduction to machine learning on Gemini Enterprise Agent Platform | Google Cloud Documentation](https://docs.cloud.google.com/vertex-ai/docs/start/introduction-unified-platform)

- [Google Cloud overview](https://docs.cloud.google.com/docs/overview)

- [AWS Cloud Essentials](https://aws.amazon.com/getting-started/cloud-essentials/)
