> ## Documentation Index
> Fetch the complete documentation index at: https://docs.pipellm.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Streaming

> Stream chat tokens as they are generated.

Set `stream: true` on Chat Completions. PipeLLM returns [server-sent events](https://html.spec.whatwg.org/multipage/server-sent-events.html) in the same shape as OpenAI.

<CodeGroup>
  ```python Python theme={"dark"}
  import os
  from openai import OpenAI

  client = OpenAI(
      api_key=os.environ["PIPELLM_API_KEY"],
      base_url="https://api.pipellm.ai/v1",
  )

  stream = client.chat.completions.create(
      model="gpt-5",
      stream=True,
      messages=[{"role": "user", "content": "Count to five."}],
  )

  for chunk in stream:
      delta = chunk.choices[0].delta.content
      if delta:
          print(delta, end="", flush=True)
  ```

  ```typescript TypeScript theme={"dark"}
  import OpenAI from "openai";

  const client = new OpenAI({
    apiKey: process.env.PIPELLM_API_KEY,
    baseURL: "https://api.pipellm.ai/v1",
  });

  const stream = await client.chat.completions.create({
    model: "gpt-5",
    stream: true,
    messages: [{ role: "user", content: "Count to five." }],
  });

  for await (const chunk of stream) {
    const delta = chunk.choices[0]?.delta?.content;
    if (delta) process.stdout.write(delta);
  }
  ```

  ```bash cURL theme={"dark"}
  curl https://api.pipellm.ai/v1/chat/completions \
    -H "Content-Type: application/json" \
    -H "Authorization: Bearer $PIPELLM_API_KEY" \
    -N \
    -d '{
      "model": "gpt-5",
      "stream": true,
      "messages": [{"role": "user", "content": "Count to five."}]
    }'
  ```
</CodeGroup>

Anthropic and Gemini native routes also stream when you set their protocol flags (`stream: true` on Messages, or `:streamGenerateContent` on Gemini). Video generation is asynchronous and is **not** streamed — poll [`GET /v2/videos/{id}`](/api-reference/video/get) instead.

See [Chat Completions](/api-reference/openai/chat-completions) for the non-streaming response shape.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.