Streaming
Set stream: true to receive tokens as server-sent events — same semantics as the native APIs.
How it works
- Add
"stream": trueto the request body. - The response is
text/event-streamwith chunks shaped exactly like the OpenAI or Anthropic format of that endpoint. - Billing is recorded once, from the usage data in the final chunks — never per-chunk.
- Failover/retry stops once the first chunk has been sent to you.
Examples
bash · OpenAI-style
curl https://api.laalaa.me/v1/chat/completions \ -H "Authorization: Bearer sk-xxxxxxxxxxxx" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-opus-4-8", "stream": true, "messages": [{"role": "user", "content": "Count to five"}] }'bash · Anthropic-style
curl https://api.laalaa.me/v1/messages \ -H "x-api-key: sk-xxxxxxxxxxxx" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-opus-4-8", "max_tokens": 1024, "stream": true, "messages": [{"role": "user", "content": "Count to five"}] }'In the SDKs
Python · OpenAI SDK
stream = client.chat.completions.create(
model="claude-opus-4-8",
messages=[{"role": "user", "content": "Count to five"}],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="")TypeScript · Anthropic SDK
const stream = client.messages.stream({ model: "claude-opus-4-8", max_tokens: 1024, messages: [{ role: "user", content: "Count to five" }],});stream.on("text", (t) => process.stdout.write(t));const msg = await stream.finalMessage();SDK stream helpers are just SSE parsing — billing is identical to non-streaming.
Errors mid-stream
If the upstream fails before the first chunk, you get a normal JSON error. After the stream has started, an error surfaces as a final event.
Review error codes → ·
POST /v1/chat/completionsទំព័រនេះមានប្រយោជន៍?