Streaming
Streaming reduces time to first text by returning Server-Sent Events while the model is generating.
curl
bash
# -N disables curl output buffering so each event appears immediately.
curl -N https://sprelaytoken.com/v1/chat/completions \
-H "Authorization: Bearer $SPRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-sol",
"stream": true,
"messages": [
{"role": "user", "content": "Explain streaming output in three points."}
]
}'stream: true belongs in the JSON body. -N is a curl flag and must stay outside the JSON.
Python SDK
python
# The SDK parses SSE events; os reads the key.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["SPRELAY_API_KEY"],
base_url="https://sprelaytoken.com/v1",
)
# stream=True returns an iterator instead of one completed response.
stream = client.chat.completions.create(
model="gpt-5.6-sol",
stream=True,
messages=[{"role": "user", "content": "Explain streaming output in three points."}],
)
# Process each parsed event and ignore events that do not contain text.
for chunk in stream:
text = chunk.choices[0].delta.content
if text:
# Avoid an automatic newline and flush output immediately.
print(text, end="", flush=True)An event may contain a role, tool fragment, or finish reason instead of text. Use an SDK or proper SSE parser, preserve incremental UTF-8 decoding, handle errors after HTTP 200, and close the connection when the user cancels. Configure connection, first-byte, idle, and total timeouts separately.
Do not blindly replay interrupted requests with tools or side effects. Check Streaming Issues for buffering and parsing problems.
