# L4 and L7 load balancers and reverse proxies: running your service behind a customer's network

> When your service has to run behind a customer's load balancer, three faults tend to surface together: users' real IPs disappear, sessions jump between servers and health checks report the wrong thing. All three can be fixed once you know what each network layer can see.

Bản gốc: https://fdetimes.net/en/guides/l4-l7-load-balancer-reverse-proxy/

Imagine it is your third day at a bank. Your internal document Q&A service has been running smoothly on the staging machine. Then the infrastructure team sends a message: all traffic must go through the bank's load balancer, no exceptions.

An hour after the switch, the logs record the same IP address for every request. Your per-IP rate limit blocks an entire department, and users are occasionally logged out mid-conversation. There is nothing wrong with your code. The problem is that you do not yet understand how the customer's network sees your requests.

This is everyday work for an FDE. You rarely get to build infrastructure from scratch; you have to fit your service into an existing system, run by other people, with its own rules. To do that well, you need to tell three things apart clearly: the layer 4 load balancer, the layer 7 load balancer and the reverse proxy.

## What does the customer's load balancer see?

According to F5's glossary, a layer 4 load balancer routes using transport-layer information, meaning addresses and ports, and does not read packet contents. It simply passes a TCP stream to a server, with no idea which URL or cookie is inside.

A layer 7 load balancer makes decisions based on characteristics of the HTTP header, such as the URL or a cookie. Because it reads HTTP, it can send `/api` to one cluster and `/static` to another. It can also insert extra headers before passing the request inside.

This matters because each layer determines what your service receives, for example whether there is an `X-Forwarded-For` header inserted by a proxy. TLS is a separate question: do not guess from the words L4 or L7 alone. Whether decryption happens at the load balancer or at your service depends on how the customer has configured it, so you have to ask.

## A reverse proxy is not a load balancer

The two terms are often used interchangeably. A reverse proxy takes a request from a client and forwards it to a server that can handle it, while a load balancer distributes requests across a group of servers. That is why a reverse proxy is still useful even with only one server behind it.

Even if the customer already has a load balancer, you should still run a reverse proxy of your own, such as NGINX, directly in front of your app. It is the place you control: recovering the real IP, setting timeouts and answering health checks. You do not have to ask the infrastructure team to change their equipment every time you need a configuration tweak.

NGINX suits this role because of its architecture. The NGINX engineering blog explains that the common way to design network applications is to assign a thread or process to each connection, and that this incurs context-switching costs.

NGINX opts for an event-driven architecture: a master process handles privileged work such as reading configuration and binding ports, while worker processes do the main work.

**Điểm mấu chốt:** Bug-free code can still break when you do not know what the customer's network sees.

## A worked example: recovering the user's real IP

Back to the bank. A request now travels: employee's browser → the bank's L7 load balancer → your NGINX → app. When a connection passes through any proxy, the server sees only the IP of the last proxy, so your logs are full of the load balancer's IP.

The fix is the `X-Forwarded-For` header: the proxy inserts the client's original IP into this header before forwarding the request inside. The app reads that header to learn who is really calling.

But MDN warns that if you use this header for security purposes, such as rate limiting or IP-based access control, you must use only the IPs added by trusted proxies. The reason is that a client can send a forged header of its own.

Suppose the infrastructure team tells you their load balancer sits in the `10.20.0.0/24` range. Your NGINX configuration would look like this:

```nginx
# Chỉ tin X-Forwarded-For khi request đến từ load balancer của khách
set_real_ip_from 10.20.0.0/24;
real_ip_header   X-Forwarded-For;
real_ip_recursive on;

upstream app {
server 127.0.0.1:8000;
}

server {
listen 80;

location /healthz {
proxy_pass http://app;
}

location / {
proxy_set_header X-Forwarded-For   $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $http_x_forwarded_proto;
proxy_set_header Host              $host;
proxy_pass http://app;
}
}
```

(The comment on the first line reads: "Only trust X-Forwarded-For when the request comes from the customer's load balancer.")

The `set_real_ip_from` line is the most important in the whole configuration. Without it, anyone who sends `X-Forwarded-For: 1.2.3.4` can impersonate someone else and get past your rate limit.

With it, NGINX trusts the value only when it comes from the load balancer's IP range. The exercise at the end shows how to verify this with `curl`.

## Why sessions jump, and why sticky sessions will not save you

Next, the logouts. Suppose you run 3 instances and configure balancing by source IP. HAProxy's architecture documentation describes this method as follows: the same IP always reaches the same server, as long as the number of servers does not change.

Look at how both conditions break. If you hash on the IP your layer sees, every request carries the load balancer's IP, so your 3 instances collapse into 1 overloaded instance. If next week you scale to 4 instances, the hash redistributes, many users are moved to a different server and lose the session held in RAM.

The more durable approach is to keep no sessions in the app. In a horizontal scaling model, the servers sit behind the load balancer and sessions live in a shared data store that every app server can reach.

At a customer site, that store is usually a Redis instance or a database they already run. In the same bank scenario, the two approaches play out like this:

| Scenario | Sticky by IP, sessions in RAM | Sessions in a shared store |
|---|---|---|
| Every request carries the load balancer's IP | All 3 instances collapse into 1 overloaded instance | Any instance can read the user's session |
| Scaling from 3 to 4 instances | The hash redistributes and many users are logged out | Users stay logged in |

## Five steps before deployment day

1. **Ask the infrastructure team three questions.** Is the load balancer L4 or L7, where does TLS terminate, and what is the proxies' IP range? These three answers determine most of your configuration.
2. **Decide who handles TLS.** The load balancer can take on encryption and decryption so the servers can focus on their main work. If the customer already does this, read `X-Forwarded-Proto` to tell whether the original request was HTTPS, rather than configuring certificates yourself.
3. **Write a `/healthz` that checks for real.** According to AWS documentation, Elastic Load Balancing checks the health of its targets and sends requests only to healthy ones. An endpoint that always returns 200 while the app has lost its database connection will keep sending users to a broken instance.
4. **Move sessions and temporary data to a shared store.** Then the customer can add or remove instances without anyone being logged out.
5. **Agree on timeouts.** Ask what the equipment's timeout is. Requests that call an LLM can run longer than the default, and being cut off midway is a very hard bug to track down.

## Common mistakes

The most common mistake is trusting every `X-Forwarded-For` without restricting its source, then using it for rate limiting. The second is enabling IP-based sticky sessions without realising that every request appears to come from the same address.

The third is assuming that a load balancer removes the need for a reverse proxy, so every configuration change means waiting on another team's ticket. The fourth is harder to spot: a health check that only confirms the process is running, not whether the app can actually serve requests.

For developers moving into FDE roles, this is a skill worth spelling out on a CV. Do not just write "knows NGINX". Describe a time you put a service behind an existing proxy, kept the real client IP and ran the service on multiple instances without losing sessions.

When a job description mentions "customer environment" or "on-prem deployment", you can expect to be asked about exactly these things.

## Exercise: try to fool your own proxy

Set up NGINX with the configuration above in front of a small app that does just one thing: log the client IP. Set `set_real_ip_from` to the IP range of a machine you treat as the load balancer.

From a machine outside that range, run `curl -H "X-Forwarded-For: 1.2.3.4" http:///`. If the log shows `1.2.3.4`, your proxy is being fooled. If it shows the sending machine's real IP, the configuration is correct.

Then remove the `set_real_ip_from` line, run the command again and compare the two log lines. The customer's network is something you cannot change, but a ten-minute test like this tells you whether your service will hold up inside it.

**Thử ngay tuần này:**

- Set up NGINX in front of a small app, log the client IP, then send a request with a forged X-Forwarded-For header to see whether the app is fooled
- Draft four questions for the customer's infrastructure team before your first deployment: is the load balancer L4 or L7, where does TLS terminate, what is the proxies' IP range, and what is the default timeout
- Move the sessions in one of your old projects into Redis, then run 2 instances behind a load balancer to check whether users get logged out

## Nguồn

- [Layer 4 Load Balancing (F5 Glossary)](https://www.f5.com/glossary/layer-4-load-balancing)

- [Reverse Proxy (F5/NGINX Glossary)](https://www.f5.com/glossary/reverse-proxy)

- [Inside NGINX: How We Designed for Performance & Scale](https://blog.nginx.org/?p=5201)

- [HAProxy Architecture Guide](http://www.haproxy.org/download/1.2/doc/architecture.txt)

- [X-Forwarded-For header - HTTP | MDN](https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Headers/X-Forwarded-For)

- [Scalability for Dummies](https://cs.fyi/guide/scalability-for-dummies)

- [What is Elastic Load Balancing? (AWS Docs)](https://docs.aws.amazon.com/elasticloadbalancing/latest/userguide/what-is-load-balancing.html)
