FDE PulseFDE jobs open 441New in 7 days 29Companies hiring 47Remote-friendly 24%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Guides

Hands-on: deploying vLLM in a customer VPC without exposing internal ports

Getting a model to run is the easy part. The hard part is deploying it inside the customer's network so that the security team signs off and the system holds up when load rises.

In brief

  • vLLM has an official Docker image, vllm/vllm-openai. The container needs --ipc=host or --shm-size to access the host's shared memory.
  • vLLM's security documentation is explicit: --api-key alone is not enough, internal ports must never be exposed to untrusted networks, and VLLM_SERVER_DEV_MODE=1 must not be enabled in production.
  • When the KV cache runs out, vLLM preempts requests and recomputes them later. Tensor parallelism splits the weights across GPUs, leaving each GPU more room for KV cache.
ShareLinkedInFacebookX
GraphicWhere each port goes when vLLM runs in a customer VPC
  1. Customer internal appsCall the OpenAI-style API from other hosts on the same network, never over the internet
  2. Access control layerGateway, reverse proxy or security group open only to the specific app hosts
  3. vLLM HTTP API portThe only port that accepts calls; --api-key is on, but not relied on alone
  4. Isolated net for internal portsDistributed comms and KV cache transfer ports; never exposed to untrusted networks

Only the API port passes through the customer's protection layer; all internal vLLM ports stay on an isolated network.

Graphic: FDE Times

vLLM’s security documentation contains a sentence every FDE should tape to their monitor: do not rely solely on --api-key to protect access to vLLM. Many people stand up a server, add a key and call it done. In the VPC of a bank or a factory, that is only the starting point.

For customers who will not let data leave their internal network, production means their machines, their subnet and their security review process. A deployment that runs on your laptop but fails the security review is not yet a deployment.

This guide follows exactly that path: deploy vLLM with Docker, test it with a real API call, lock down the network layer, deal with GPU memory, then turn everything into something reproducible.

By the end, you will have a vLLM server in a private subnet serving an OpenAI-style API to internal applications, along with a port diagram clear enough to hand to the customer’s security team.

What do you need before typing any commands?

You need a GPU machine in a subnet with no public address, with Docker installed and configured so containers can use the GPU. You also need a second machine on the same network to make test calls. Pick an open-weight model the customer is permitted to use; in the commands below, the model name is held in the $MODEL variable.

The commands here have been trimmed for readability. GPU flags, how the model cache directory is mounted and the default port all change between versions, so before running anything at the customer’s site, check them against the “Using Docker” and “Online Serving” pages in the vLLM documentation.

Step 1: The official image, and the shared memory flag you must not forget

vLLM has an official Docker image, vllm/vllm-openai. Using it instead of building your own lets you quickly answer the security team’s first question: where does the software come from?

docker pull vllm/vllm-openai

When running in a container, vLLM needs access to the host’s shared memory. The documentation offers two options: the --ipc=host flag or the --shm-size flag. The skeleton of the run command looks like this:

# COMMAND SKELETON, NOT RUNNABLE AS WRITTEN:
# missing the flag that gives the container GPU access, the flag mapping the API port to the host
# and the model cache mount. Take the full set of flags from the "Using Docker" page.
export MODEL=ten-model-open-weight
docker run --ipc=host vllm/vllm-openai --model "$MODEL"

To be blunt: if you copy the command above verbatim, the container will not see the GPU, and even if it did run, other machines could not reach it because the API port has not been mapped to the host. The shared memory flag is what this guide wants you to remember; take the GPU and port flags exactly as specified for the vLLM version you are using.

If the customer’s container isolation policy forbids --ipc=host, switch to --shm-size and record the reason in the handover documentation. Check: the container log gets through model loading without any shared memory errors.

Step 2: Test with a real request

Besides Docker, you can start the server directly with vllm serve followed by the model name. Either way, vLLM opens an HTTP server compatible with OpenAI-style interfaces. That is why the customer’s application team usually does not need to change much code: mostly they just change the server address.

# Illustration: set VLLM_BASE_URL, VLLM_API_KEY, MODEL to match your actual configuration
# e.g. VLLM_BASE_URL=http://10.0.1.15:8000/v1 (internal IP of the GPU machine)
import os
from openai import OpenAI

client = OpenAI(
    base_url=os.environ["VLLM_BASE_URL"],
    api_key=os.environ["VLLM_API_KEY"],
)

resp = client.chat.completions.create(
    model=os.environ["MODEL"],
    messages=[{"role": "user", "content": "Answer in one sentence: what is 2 + 2?"}],
)
print(resp.choices[0].message.content)

Check: run this from the second machine on the same subnet, not from the GPU machine itself, and the terminal should print an answer. If you only test via localhost, you have tested nothing about the network.

Step 3: Why is an API key not enough?

You can certainly start the server with --api-key, and you should. But the vLLM documentation is clear that you should not rely on it alone. At a customer site, the other layer of protection is usually access control infrastructure they already have, such as a gateway, a reverse proxy or a security group that lets only the right application machines call in.

The internal ports matter even more. When running distributed, vLLM uses additional ports for inter-process communication and for transferring KV cache. The documentation requires that these ports are never exposed to the internet or to untrusted networks, so they must sit on an isolated network.

Before handover, there is one more check: the environment variable VLLM_SERVER_DEV_MODE=1 must never be enabled in production. Run env inside the container and look for it, because it can easily slip in from a configuration file used during testing.

Step 4: What happens when the KV cache runs out?

When the KV cache has no room left, vLLM preempts some requests and recomputes them once space frees up. The system does not crash, but users see latency rise for no apparent reason. So if the application team complains that things are “sometimes fast, sometimes slow”, look for signs of preemption in the logs before suspecting the network.

Consider a simplified case: an 80 GB GPU, with the model weights taking 60 GB, leaves only about 20 GB for KV cache and everything else. Enable tensor parallelism across 2 GPUs and the weights are split in half: each GPU holds about 30 GB and has about 50 GB free.

The vLLM documentation describes exactly that mechanism: splitting weights across multiple GPUs so each GPU has more memory for KV cache. The specific parameter names are on the “Optimization and Tuning” page.

But the cost comes with clear numbers too.

IBM also warns that infrastructure and maintenance costs can eat up most of a deployment budget. The FDE’s job is to give the customer both numbers, latency and GPU count, and let them choose.

Step 5: Package it into a script that builds from scratch

Once you have typed all the commands by hand, package them into a script or infrastructure configuration that builds everything from scratch: the image, the shared memory flag, the GPU configuration, the network rules and the dev mode check. AWS argues that automating model creation and deployment shortens time to market and reduces operating costs.

At the customer site, the benefit is more concrete: when their IT team needs to rebuild after a maintenance window, they do not have to call you.

Once the server is running stably, monitoring comes next. IBM regards tracking model performance, especially model drift, as the core of monitoring. Agree with the customer from the outset who looks at which metrics, and how often.

The most common mistakes

Symptom Common cause Fix
Container fails while loading the model No access to shared memory Add --ipc=host or --shm-size
Security team refuses sign-off Only --api-key in place, internal ports not isolated Put the API behind an access control layer, give internal ports their own network
Latency rises and falls KV cache exhausted, requests preempted Check the logs, consider tensor parallelism or reduce load
Security risk in production VLLM_SERVER_DEV_MODE=1 still enabled, against vLLM documentation guidance Remove the variable, add a check to the script

What does this skill look like at the customer site and on a CV?

On a real project, most of the time is not spent on vllm serve. It goes into meetings with the network team to open exactly one port, sessions explaining to the security team why internal ports never leave the isolated network, and conversations about GPU count with whoever holds the budget.

When reading FDE job descriptions, watch for phrases such as “self-hosted LLM”, “on-prem”, “VPC deployment” or “customer environment”. On your CV, do not just write “used vLLM”. Describe the work: deployed vLLM in a private subnet, isolated the network for internal ports, handled preemption with tensor parallelism across 2 GPUs, handed over with a reproducible build script.

Next time someone asks whether you have deployed an LLM, the most valuable answer is not the model name. It is a port diagram you can draw yourself, and the reason behind every arrow on it.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
6 sources
Read next on the roadmap · Stage 5: DeploymentSAML, SCIM and the security questionnaire: what FDEs run into on their first enterprise contractThe demo can work perfectly, but the customer's security team will not let you into production until you can say how long a departing employee keeps access.