# Streaming LLM output to the UI: SSE first, WebSocket only when you need it

> Most LLM chat features only need the server to push tokens to the browser. Reaching for WebSocket out of habit can leave you with infrastructure trade-offs the client never needed.

Original: https://fdetimes.net/en/guides/stream-llm-responses-sse-vs-websocket/

OpenAI's streaming guide states plainly that it focuses on HTTP streaming with `stream=true` over server-sent events. The model provider has already chosen how tokens leave its side. What remains is getting those tokens from your backend to the client's screen, and this is where many teams choose badly.

Picture the first week at a banking client. They want an internal assistant that answers questions about procedures, and staff do not want to stare at a blank screen while the model thinks. An engineer used to building realtime apps will set up WebSocket straight away.

But if the interface only needs to receive text as it flows down, that choice brings infrastructure trade-offs the problem does not require.

For an FDE, choosing the streaming channel is the first architectural decision the client will actually see, because it determines whether the demo feels smooth. This guide covers three options in the order you should try them: SSE first, WebSocket when there is a reason, polling when nothing else works.

## What you will build, and what you need

You will build one endpoint that accepts a question and calls the LLM with `stream=true`, plus a second endpoint that pushes each chunk of text to the browser over SSE. The browser uses `EventSource` to render the text as it arrives. You will then rewrite the client side with WebSocket for comparison.

You need any backend you are comfortable with, an API key from an LLM provider that supports streaming, and a browser with DevTools. The code below is a sketch with error handling and authentication stripped out. Treat it as a skeleton to port to your framework, not something to paste into production.

## Step 1: Separate calling the model from reading the stream

According to MDN, the server-side script sending events must respond with the MIME type `text/event-stream`. Without this header `EventSource` will not work, even though the data still crosses the network.

MDN also describes SSE as a one-way connection: the client cannot send events back to the server. So the user's question travels in a separate POST request that creates the conversation turn. The GET endpoint only reads back what the model has already generated and never calls the model itself.

```text
# Simplified sketch, not tied to any framework
POST /chat/turns                 # create the turn, call the model exactly once
  turn_id = new_id()
  run_in_background:
      for chunk in llm.call(messages, stream=true):
          buffer[turn_id].append(chunk)
  return turn_id

GET /chat/stream?id=TURN_ID      # only reads the buffer, never calls the model
  set header Content-Type: text/event-stream
  for chunk in buffer[TURN_ID], wait for new chunks until done:
      write "data: " + chunk + "\n\n"
      flush
```

How to check: open the Network tab in DevTools, call the GET endpoint and look at the `Content-Type` header. If the response appears all at once at the end instead of trickling in, the backend or a proxy layer may be buffering data before sending it. Check the flush call first.

## Step 2: The browser side takes only a few lines

Once the POST returns a `turnId`, the browser opens the stream to receive the answer.

```javascript
// Simplified sketch
const es = new EventSource('/chat/stream?id=' + turnId);
es.onmessage = (e) => { output.textContent += e.data; };
```

The most important check is to switch off the network for a few seconds mid-stream and then switch it back on. According to MDN, the browser reconnects automatically by default when the connection closes, and the wait time is controlled by the `retry` field. Open the backend logs and confirm that the model was not called again.

Note that the sketch above rereads the buffer from the start on reconnection, so text on screen may be printed twice. In a real build you need to clear what was already rendered before rereading, or track the position already read.

Phil Sturgeon, who writes the APIs You Won't Hate blog, observes that SSE works very well inside an HTTP/REST API for sending updates. Because SSE is still HTTP, the stream endpoint can sit alongside your other APIs.

But as Step 1 warned, the client's proxy can still buffer the stream, so test on their actual infrastructure before the demo.

## Step 3: When is it worth moving to WebSocket?

WebSocket opens a two-way session between browser and server. roadmap.sh describes it as a persistent full-duplex channel over a single TCP connection. Ably draws the line in a similar way: SSE suits one-way updates pushed by the server, while WebSocket suits two-way communication such as games or chat.

```javascript
// Simplified sketch
const ws = new WebSocket('wss://example.com/chat');
ws.onmessage = (e) => { output.textContent += e.data; };
ws.send(JSON.stringify({ type: 'question', text: q }));
```

The deciding question is this: while the model is answering, does the interface need to send anything up? An agent that pauses to ask the user for confirmation before taking an action, or a screen where several people watch the same answer stream, are legitimate reasons. A "Stop" button, by contrast, usually needs only a separate request to cancel that turn.

Switching to WebSocket costs you a few things. MDN notes that the standard WebSocket API does not support backpressure, so if the client processes data more slowly than tokens arrive, data piles up. roadmap.sh points out that CDNs and proxies cannot cache WebSocket connections.

For these reasons, roadmap.sh recommends SSE over HTTP rather than WebSocket when you only need the server to push data down.

**Key point:** If the interface only needs to receive text as it flows down, do not pay for a two-way channel.

## Polling: a fallback, not a default

With polling, the client periodically sends a request asking the server whether more text is available. Hookdeck observes that this wastes resources, and that you constantly have to weigh whether the next call will return anything. With an LLM, most requests will come back empty or with only a few extra tokens.

Polling still has a place when the client's environment does not let long-lived connections survive. In that case, use it as a conditional fallback mode and tell the client clearly that the experience will be choppier.

| | SSE | WebSocket | Polling |
|---|---|---|---|
| Direction | Server → client | Two-way | Client asks periodically |
| Weakness to remember | 6-connection limit without HTTP/2 | No backpressure, cannot be cached through a CDN/proxy | Wastes resources, many empty requests |
| Use when | Chat, single-turn answers | Agent needs to ask back, multiple users sharing a view | Long-lived connections are blocked |

## Three failures that wreck a demo

The first comes from automatic reconnection itself. If the stream endpoint both accepts the request and calls the model, every browser reconnection can generate a fresh answer from scratch, wasting tokens and printing duplicate text. The safer approach is to separate creating the conversation turn from reading the stream, as in Step 1.

The second is the connection limit. MDN warns that when not running over HTTP/2, SSE is limited to a very low 6 connections per browser. If a user opens several tabs of the same chat page, the seventh tab may hang waiting for a connection. Ask the client's infrastructure team whether the server runs HTTP/2 before the demo.

The third is more about process than technology. OpenAI warns that streaming model output in production makes content moderation harder, because a partial answer is difficult to evaluate.

With financial or healthcare clients, ask early whether content must pass through a filter before it is displayed, because the answer may force you to change the whole design.

## Three questions before the first line of code

On a client site, the hard part is not writing `EventSource`. Before writing the first line of code, ask all three questions: does the interface need to send anything up mid-stream, does the infrastructure have HTTP/2 and which proxy layers sit in between, and does the content need moderation?

When reading job descriptions for FDE or AI engineer roles, watch for phrases such as "streaming", "realtime UI" or "production LLM app". On your CV, instead of writing "used WebSocket", write one line that gives the reasoning: chose SSE for one-way chat, separated the conversation turn from the read stream so reconnections never call the model again.

A line like that shows you understand the trade-offs, not just the tools.

Choose the simplest option that still meets the need, and have your reasoning ready when someone asks why you did not use WebSocket.

**Try this week:**

- Build the minimal SSE endpoint from Steps 1–2, then cut the network mid-stream to see whether the browser reconnects and whether the model gets called again
- Open the same chat page in 7 tabs on a server not running HTTP/2 and note which tab hangs
- Write a five-sentence portfolio paragraph explaining why you chose SSE or WebSocket for a project, including the trade-offs you accepted

## Sources

- [Using server-sent events - MDN](https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_events)

- [The WebSocket API (WebSockets) - MDN](https://developer.mozilla.org/en-US/docs/Web/API/WebSockets_API)

- [What are realtime APIs and when to use them?](https://ably.com/topic/what-is-a-realtime-api)

- [WebSocket vs. HTTP: Which protocol should you use?](https://roadmap.sh/network-engineer/websocket-vs-http)

- [When to Use Webhooks, WebSocket, Pub/Sub, and Polling](https://hookdeck.com/webhooks/guides/when-to-use-webhooks)

- [Streaming Data with REST APIs](https://apisyouwonthate.com/blog/streaming-data-with-rest-apis/)

- [Streaming API responses - OpenAI](https://developers.openai.com/api/docs/guides/streaming-responses)
