POST /v1/chat/completions is the primary inference endpoint. It is fully compatible with the OpenAI Chat Completions API, so you can point any OpenAI SDK at Darkbloom by changing the base URL — no other code changes required. The endpoint supports both streaming (server-sent events) and non-streaming responses.
Request parameters
object[]
required
The conversation history as an array of message objects. Each message has a
role (system, user, or assistant) and a content string.boolean
default:"false"
When
true, the response is delivered as a stream of server-sent events. Each event contains a delta chunk in the standard OpenAI SSE format, ending with data: [DONE].number
default:"8192"
Maximum number of tokens to generate. If you do not set this, the coordinator injects a default of 8192 to bound the worst-case cost and ensure the pre-flight balance reservation covers the full generation.
number
default:"1.0"
Sampling temperature between 0 and 2. Lower values produce more deterministic output.
Examples
Basic request
Streaming
With a system prompt
Response headers
Every response from the inference endpoint includes trust metadata headers that let you verify which provider handled the request and inspect their attestation status.Related endpoints
POST /v1/responses — OpenAI Responses API alias. The coordinator auto-detects the input vs messages field shape and routes to the same underlying handler.
POST /v1/completions — Legacy text completions endpoint. Accepts a prompt string instead of a messages array. This is the older OpenAI /v1/completions format, not recommended for new integrations.
POST /v1/messages — Anthropic Messages API compatible endpoint. Use this if you are working with the Anthropic Python or TypeScript SDK. See the note below.
To use the Anthropic SDK, set
base_url="https://api.darkbloom.dev" (without /v1) and pass your Darkbloom API key as api_key. The Anthropic SDK appends /v1/messages automatically.Thinking tags
Some models emit extended thinking wrapped in<think>...</think> tags. These tags are stripped from responses by default before the content reaches you. The final content field contains only the model’s visible output.
