Laalaa Docs

Streaming

Set stream: true to receive tokens as server-sent events — same semantics as the native APIs.

How it works

  • Add "stream": true to the request body.
  • The response is text/event-stream with chunks shaped exactly like the OpenAI or Anthropic format of that endpoint.
  • Billing is recorded once, from the usage data in the final chunks — never per-chunk.
  • Failover/retry stops once the first chunk has been sent to you.

Examples

bash · OpenAI-style
curl https://api.laalaa.me/v1/chat/completions \  -H "Authorization: Bearer sk-xxxxxxxxxxxx" \  -H "Content-Type: application/json" \  -d '{    "model": "claude-opus-4-8",    "stream": true,    "messages": [{"role": "user", "content": "Count to five"}]  }'
bash · Anthropic-style
curl https://api.laalaa.me/v1/messages \  -H "x-api-key: sk-xxxxxxxxxxxx" \  -H "anthropic-version: 2023-06-01" \  -H "Content-Type: application/json" \  -d '{    "model": "claude-opus-4-8",    "max_tokens": 1024,    "stream": true,    "messages": [{"role": "user", "content": "Count to five"}]  }'

In the SDKs

Python · OpenAI SDK
stream = client.chat.completions.create(
    model="claude-opus-4-8",
    messages=[{"role": "user", "content": "Count to five"}],
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta.content
    if delta:
        print(delta, end="")
TypeScript · Anthropic SDK
const stream = client.messages.stream({  model: "claude-opus-4-8",  max_tokens: 1024,  messages: [{ role: "user", content: "Count to five" }],});stream.on("text", (t) => process.stdout.write(t));const msg = await stream.finalMessage();
SDK stream helpers are just SSE parsing — billing is identical to non-streaming.

Errors mid-stream

If the upstream fails before the first chunk, you get a normal JSON error. After the stream has started, an error surfaces as a final event.

Review error codes → · POST /v1/chat/completions
ទំព័រនេះមានប្រយោជន៍?