# Ollama and LM Studio: demoing open-weight models on a laptop when the client bans data egress

> When a client's legal team rules out every cloud API, the demo depends on a single laptop. Whether it works comes down to how you set that laptop up before you reach the client's office.

Original: https://fdetimes.net/en/tools/ollama-lm-studio-offline-client-demos/

"Data cannot leave the building" often ends a demo before the first slide is up. For banks, hospitals and factories this is a hard requirement, not an opening position. In that situation, all a forward deployed engineer can bring is a laptop running an open-weight model locally.

Ollama and LM Studio are two common ways to do this. They do not compete on model intelligence, because the models come from third parties. The difference is how each tool behaves once the machine is cut off from the network, and that detail decides whether you get through the meeting with the client's security team.

## Ollama suits people at home in the terminal

Ollama is an open-source project under the MIT licence, built on the llama.cpp backend. Its official repository lists many model families, including Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen and Gemma. The line worth quoting to a client is in the FAQ: when run locally, Ollama does not see users' prompts or data.

What forward deployed engineers value most in Ollama is that its defaults are already fairly safe. The server binds only to loopback, on port 11434. To let other machines on the network reach it, you have to change the `OLLAMA_HOST` environment variable yourself, so exposing it is always a deliberate choice. A demo app on the same laptop just calls the local REST API:

```bash
curl http://localhost:11434/api/chat -d '{
  "model": "<downloaded-model-name>",
  "messages": [{"role": "user", "content": "Summarize this confidentiality clause"}]
}'
```

(The placeholder is the name of a downloaded model; the prompt reads "Summarise this confidentiality clause".)

Take a demo for a bank that wants to summarise contracts. Before the meeting, you download the model at home. You then point `OLLAMA_MODELS` at a folder on an encrypted drive approved by the client's IT department, since the FAQ allows the model storage location to be changed with this variable.

You also turn off Ollama's cloud features as the FAQ describes, giving up cloud models and web search. On an engagement that bans egress, that trade-off is the right one.

## LM Studio suits demos for people who do not write code

LM Studio has a graphical interface that you can hand to a business specialist to try for themselves. Its official documentation says that what you type when chatting with a local model does not leave the device. Documents users upload for question answering are also processed entirely on the machine.

The privacy policy in force from June 2026 adds a further assurance: in local use, messages, chat history and documents are never transmitted off the system, and only the company's Cloud Services process content.

LM Studio's runtime is based on MLX and llama.cpp, so on an Apple Silicon MacBook it can run through MLX. To connect an application, you start a local OpenAI-compatible server from the Developer tab or with `lms server start`. This server can run on `localhost` or be exposed to the network. In a strict environment, keep it on localhost.

**Key point:** The promise that "data never leaves the machine" holds for the chat. It does not necessarily hold for the whole tool.

## Both tools still have parts that need the network

This is the part many people skip. LM Studio still sends requests out, for example to huggingface.co, when you search for models in the Discover tab. Without internet access, the in-app updater does not work either.

Models therefore have to be downloaded in advance or sideloaded, and updates have to come through a separate channel, not on the locked-down laptop. Settle the list of models you need and download them the day before. Do not open Discover at the client's site just to grab one more model "to be safe".

Ollama has cloud models and web search, and both can be turned off. Put that step in your machine setup script so you do not have to remember it next time. If the client's security team is watching network traffic, one unexpected request in the middle of a demo will cost you trust faster than any wrong answer from the model.

## A dry run that catches mistakes before the client does

The easiest mistake is leaving the server exposed to the network. LM Studio has this option built in, and switching it on once to test at the office is easy to forget when you carry the laptop to the client. With Ollama, the equivalent risk is changing `OLLAMA_HOST` and not changing it back.

First step of the dry run: turn off Wi-Fi and run the whole demo script from start to finish, including chatting with documents. Any step that fails is a step still quietly relying on the network.

Next, check where the server is listening. On a Mac, the command below should return a loopback address such as `127.0.0.1:11434`. If you see `*:11434` or the IP of a network interface, the server is open to the outside. Do the same for whatever port LM Studio's server is using.

```bash
lsof -nP -iTCP:11434 -sTCP:LISTEN
```

Finally, turn the network back on, run the script again and list open connections with `lsof -nP -iTCP -sTCP:ESTABLISHED`. Every connection not pointing to 127.0.0.1 is a question you need to be able to answer before the client's security team asks it.

## What to learn first, and what to put on your CV

Learn Ollama first. Its API is simple, it is configured through environment variables so it is easy to turn into a setup script, and the loopback default makes it straightforward to explain to a security team. Keep LM Studio ready for when the person trying the demo is a business user, or when you want to use MLX on a Mac.

When reading forward deployed engineer job descriptions, look for phrases such as "on-prem", "air-gapped" or "regulated industries". They signal that you will need this skill. On your CV, do not just write "used Ollama". Describe building a no-egress demo environment: models downloaded in advance, stored on an encrypted drive, the server listening only on localhost, cloud features turned off, and an offline dry run completed.

A client's security team rarely asks how smart your model is. They ask what your laptop will send out, and you should be able to answer that from the terminal on the spot.

**Try this week:**

- Install Ollama, download a model from the Qwen or Gemma family, turn off Wi-Fi and call http://localhost:11434/api/chat with curl
- Point OLLAMA_MODELS at a folder on an encrypted drive, download the model again and confirm Ollama reads it from there
- Start LM Studio's server with `lms server start`, then point a script using the OpenAI SDK at that localhost address

## Sources

- [GitHub - ollama/ollama](https://github.com/ollama/ollama)

- [Ollama FAQ](https://docs.ollama.com/faq)

- [Offline Operation - LM Studio Docs](https://lmstudio.ai/docs/app/offline)

- [LM Studio Privacy Policy](https://lmstudio.ai/privacy)

- [LM Studio Developer Docs - Local LLM API Server](https://lmstudio.ai/docs/developer/core/server)

- [LM Studio Bionic - Your Agent for Work and Code](https://lmstudio.ai/)
