# Latency, throughput, scalability: answering the client's infrastructure team with measurements

> When the client's infrastructure team asks "what's the p99, and how many requests per second can it take?", answering "it runs pretty fast" will cost you their trust in the very first meeting.

Bản gốc: https://fdetimes.net/en/guides/latency-throughput-scalability-with-numbers/

You have just deployed an invoice data extraction service at a client. At the review meeting, the head of their infrastructure team asks two questions: "What's the p99? How many requests per second can one instance handle?" You reply that in testing it seemed pretty fast. The room goes quiet for a moment, and you realise you have just been filed under "vendor team that hasn't measured".

This happens to FDEs more often than you might think. Writing code that works is the part you are used to. Convincing the people who will be on call for that system that it will not fall over at peak is a different skill. To do it, you need to use three words correctly (latency, throughput, scalability), and each one has to come with a number.

## Three words, three different questions

Latency is the time a system takes to respond to a request. Throughput is how many requests the system can handle at once, usually expressed in practice as requests per second. The system design primer on cs.fyi notes that in most systems these two trade off against each other: pile on more work and each piece of work usually waits longer.

Because there is a trade-off, there has to be a goal. Jonas Bonér, in his slide deck on scalability patterns, sets the goal as maximum throughput with acceptable latency. What counts as "acceptable" is for the client to decide, not you. So the first question to ask the infrastructure team is what their latency threshold is, not to show off your own numbers.

The third word is scalability, and this is where it most often gets confused with performance. A post on the Professor Beekums blog separates the two neatly: scalability is handling large numbers of users, data or traffic, while performance is about speed.

The author uses the example of a checkout counter: scalability is keeping the speed of service constant whether customers arrive in crowds or in a trickle.

Werner Vogels, Amazon's CTO, defines it more strictly: a service is scalable if adding resources results in increased performance. For always-on services he adds a condition: resources added for redundancy must not degrade performance.

Vogels also makes clear that "increased performance" can mean two things. Usually it means serving more units of work, but it can also mean handling larger units of work. In load testing, people tend to remember the first meaning and forget the second.

## Slow for one user, or only slow under load?

Bonér offers a test you can use straight away. If the system is slow even for a single user, you have a performance problem. If it is fast for one user but slow under heavy load, you have a scalability problem. The two need different fixes, so a wrong diagnosis wastes the effort you put in.

Imagine your invoice service takes 4 seconds per request when nobody else is using it. Adding servers will not help, because each request still takes 4 seconds. You need to open a trace and see whether the time goes into OCR, the model call or a database query. Only once a single user finds it fast is it time to ask about load.

## A complete measurement, step by step

Suppose the service is already fast for one user. The next step is to fire increasing load at a single replica and record the latency distribution, not just the average. The script below is enough for a first pass:

```python
import asyncio, time, statistics, httpx

async def one(client, url, payload, out):
t = time.perf_counter()
await client.post(url, json=payload, timeout=60)
out.append(time.perf_counter() - t)

async def run(url, payload, concurrency, total=200):
lat, sem = [], asyncio.Semaphore(concurrency)
async with httpx.AsyncClient() as c:
async def task():
async with sem:
await one(c, url, payload, lat)
start = time.perf_counter()
await asyncio.gather(*[task() for _ in range(total)])
wall = time.perf_counter() - start
q = statistics.quantiles(lat, n=100)
print(concurrency, round(q[49], 2), round(q[98], 2), round(total / wall, 1))
```

Run the script at concurrency 1, 4, 8 and 16 against one replica. Then put a second replica behind the load balancer, point the script at the load balancer and run it again with 16 concurrent requests, roughly 8 per replica. The table below shows illustrative results, not real measurements, but its shape is one you will see again and again:

| Replicas | Concurrent requests | p50 (s) | p99 (s) | Requests/s |
|---|---|---|---|---|
| 1 | 1 | 1.2 | 1.5 | 0.8 |
| 1 | 4 | 1.3 | 1.9 | 3.0 |
| 1 | 8 | 1.6 | 3.8 | 5.0 |
| 1 | 16 | 3.2 | 9.5 | 5.0 |
| 2 | 16 | 1.6 | 3.9 | 9.5 |

Here is how to read it. From 1 to 8 concurrent requests, throughput rises while p50 creeps up only slightly, although p99 has already started to stretch, from 1.5 to 3.8 seconds. From 8 to 16, throughput stays flat at 5 requests per second while p50 doubles and p99 jumps to 9.5 seconds. The replica is saturated, and the extra requests simply queue.

The last row is the scalability test exactly as Vogels defines it. Same 16 concurrent requests, but with 2 replicas throughput rises to 9.5 requests per second, nearly double, and p50 and p99 return to roughly the single-replica level at 8 requests. Adding resources increases performance proportionally: this service scales.

If that row had shown 6 requests per second instead of nearly 10, the two replicas would be competing for a shared resource, perhaps the database, perhaps the model API's rate limit. In that case a third replica would be pointless until you find that bottleneck.

Suppose the client sets a p99 threshold of 5 seconds. Your answer in the meeting would be: "One replica serves about 5 requests per second with p99 under 4 seconds, if we keep it at around 8 concurrent requests.

Two replicas measured 9.5 requests per second at the same p99, so to reach 20 requests per second we estimate 4 replicas, and we will measure again at 4 replicas to confirm." That is a sentence an infrastructure team can take away and use for capacity planning.

You should also test a 50-page invoice, because according to Vogels, handling larger units of work is also a dimension of scalability.

## The tail of the distribution is what keeps you up at night

There is a reason the table uses p99 rather than the average. Uwe Friedrichsen, analysing "The Tail at Scale" by Jeffrey Dean and Luiz Barroso, reiterates that at scale, temporary spikes in latency can dominate the performance of an entire service.

He also observes that the more successful a service becomes, the more it is affected by whatever tail latency remains.

A little arithmetic shows why. Suppose one user request has to call 100 sub-services in parallel, and each sub-service is unusually slow only 1% of the time. The probability of at least one slow call is 1 − 0.99^100, about 63%, so each component's tail has become the common case for the system as a whole.

**Điểm mấu chốt:** The average only tells you what a typical user sees, while clients remember the times the system was at its slowest.

Friedrichsen describes one technique for cutting this tail: hedged requests, which means sending the same request to several replicas and using whichever response comes back first.

Sending multiple copies means adding load, so a careful way to apply it is to send the copy only once the first request has passed a threshold, such as the p95 you measured. And when you propose this technique to the infrastructure team, tell them up front what percentage of extra load it costs.

## The mistakes that lose an infrastructure team's trust

The most common mistake is reporting averages. The second is measuring on a laptop with one user and calling the result "production performance", which mixes up performance and scalability. The third is adding replicas for redundancy without re-testing: by Vogels's criterion, if redundancy makes the system slower, the service does not yet meet the bar.

There is a subtler one too: quoting a throughput figure without a latency threshold. "Handles 5 requests per second" without saying what p99 is at that level means almost nothing, because at 16 concurrent requests the system still "handles" the load; each user just waits almost 10 seconds.

If you are aiming for an FDE role, a CV line such as "load-tested and capacity-planned service X: 5 req/s per replica at p99 under 4 seconds, scaling near-linearly with added replicas" is far more convincing than "optimised system performance".

When a job description mentions taking systems to production or working with a client's infrastructure team, bring a measurement table like the one above to the interview.

Next time a client's infrastructure team asks about p99, you will not need to recite definitions. You only need to open a table of a few rows of numbers you measured yourself.

**Thử ngay tuần này:**

- Pick an endpoint you own, measure it at 1, 4, 8 and 16 concurrent requests, and record p50, p99 and requests per second in a table.
- From that table, write exactly one answer of the form "one replica serves X requests/second with p99 under Y seconds" and send it to an infrastructure colleague for feedback.
- Add a replica and measure again to see whether throughput nearly doubles.

## Nguồn

- [A Word on Scalability (All Things Distributed)](https://www.allthingsdistributed.com/2006/03/a_word_on_scalability.html)

- [Performance Vs Scalability](https://blog.professorbeekums.com/performance-vs-scalability/)

- [System Design: Latency vs Throughput (cs.fyi)](https://cs.fyi/guide/latency-vs-throughput/)

- [Scalability, Availability & Stability Patterns (Jonas Bonér)](https://www.slideshare.net/jboner/scalability-availability-stability-patterns/)

- [The tail at scale (Uwe Friedrichsen)](https://ufried.com/blog/tail_at_scale/)
