Create Chat Completion
POST /v1/chat/completions
Creates a model response for the given chat conversation. Returns a chat completion object, or a streamed sequence of chat completion chunk objects if the request is streamed.
- Uses the OpenAI Chat Completions API request format
- Plain text messages and multi-turn conversations
- Streaming and non-streaming responses
For image, audio, and file analysis, see File Analysis. For tool use, see Tool Calling.
Endpoint
https://api.tokatlas.ai/v1/chat/completionsAuthentication
All endpoints require Bearer Token authentication.
Authorization: Bearer YOUR_API_KEY
Content-Type: application/jsonImportant
Never commit a real API key to a repository or expose it in client-side code.
Request Body
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
model | string | Yes | - | Model ID used to generate the response |
messages | array | Yes | - | List of messages comprising the conversation |
temperature | number | No | 1.0 | Sampling temperature, range 0–2 |
top_p | number | No | 1.0 | Nucleus sampling parameter, range 0–1 |
max_tokens | integer | No | - | Maximum tokens to generate (deprecated for o-series; use max_completion_tokens) |
max_completion_tokens | integer | No | - | Upper bound for generated tokens, including reasoning tokens |
stream | boolean | No | false | Whether to stream the response via SSE |
stream_options | object | No | - | Options for streaming (only when stream: true) |
stop | string or array | No | - | Up to 4 sequences where generation stops |
n | integer | No | 1 | Number of chat completion choices to generate |
frequency_penalty | number | No | 0 | Frequency penalty, range -2.0 to 2.0 |
presence_penalty | number | No | 0 | Presence penalty, range -2.0 to 2.0 |
response_format | object | No | - | Output format: text or JSON schema |
reasoning_effort | string | No | - | Reasoning effort for reasoning models |
metadata | object | No | - | Up to 16 key-value pairs attached to the object |
store | boolean | No | - | Whether to store the output for later retrieval |
model
First list available models and confirm Chat Completions support. Current upstream models are listed by OpenAI, Claude, and Gemini.
For example, GPT-6.1 Sol and GPT-6 Astra support text requests in Chat, but tool calls require Responses. Other providers also require protocol conversion on the Tokatlas route; a model name alone does not establish compatibility.
Current upstream IDs include:
- OpenAI:
gpt-6.1-sol,gpt-6-astra,gpt-6-luna - Anthropic:
claude-opus-5-5,claude-sonnet-5-5,claude-fable-5-1 - Google:
gemini-3.8-flash,gemini-3.5-flash-lite
messages
A list of messages comprising the conversation so far. Each message includes role and content.
| Role | Description |
|---|---|
developer | Developer-provided instructions. With o1 models and newer, developer messages replace system messages |
system | System prompt to set AI behavior (use developer for o1 and newer) |
user | Messages sent by the end user |
assistant | Messages sent by the model in response to user messages |
tool | Tool results returned by your application with the matching tool_call_id |
Basic user message:
[{"role": "user", "content": "Hello!"}]Developer / system prompt:
[
{"role": "developer", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"}
]Multi-turn conversation:
[
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi! How can I help you?"},
{"role": "user", "content": "Tell me about AI"}
]For text messages, content is a plain string. It may also be an array of text content parts:
{"role": "user", "content": [{"type": "text", "text": "Hello!"}]}temperature
What sampling temperature to use, between 0 and 2. Higher values like 0.8 make output more random; lower values like 0.2 make it more focused. Recommend using either temperature or top_p, not both.
top_p
Nucleus sampling parameter, range 0–1.
max_tokens / max_completion_tokens
max_tokens— maximum tokens to generate. Deprecated in favor ofmax_completion_tokens; not compatible with o-series models.max_completion_tokens— upper bound for generated tokens, including visible output and reasoning tokens.
stream
If set to true, the model response is streamed to the client as it is generated using Server-Sent Events. Each chunk has object set to chat.completion.chunk.
true— Streaming responsefalse— Complete response at once (default)
stream_options
Options for streaming response. Only set when stream: true.
| Field | Type | Description |
|---|---|---|
include_usage | boolean | Stream a final chunk with token usage before data: [DONE] |
include_obfuscation | boolean | Add obfuscation fields to normalize payload sizes (default true) |
stop
Up to 4 sequences where the API stops generating further tokens. Not supported with latest reasoning models o3 and o4-mini.
n
How many chat completion choices to generate for each input message. Keep n as 1 to minimize costs.
frequency_penalty / presence_penalty
Number between -2.0 and 2.0.
frequency_penalty— penalizes tokens based on their existing frequency in the textpresence_penalty— penalizes tokens based on whether they appear in the text so far
response_format
An object specifying the format the model must output.
Plain text (default):
{"type": "text"}Structured JSON output (JSON Schema):
{
"type": "json_schema",
"json_schema": {
"name": "math_response",
"schema": {
"type": "object",
"properties": {
"steps": {"type": "array", "items": {"type": "string"}},
"final_answer": {"type": "string"}
},
"required": ["steps", "final_answer"],
"additionalProperties": false
},
"strict": true
}
}reasoning_effort
Constrains effort on reasoning for reasoning models. Supported values: none, minimal, low, medium, high, xhigh, max.
{
"model": "o3-mini",
"messages": [{"role": "user", "content": "Solve this math problem"}],
"reasoning_effort": "high"
}metadata
Up to 16 key-value pairs (keys max 64 chars, values max 512 chars) for storing additional information about the object.
Response
The table describes the protocol response object. Where the non-streaming examples use a { "code": 200, "data": { ... } } envelope, read the object in data; for routes returning the protocol object directly, read the root object. Parse streaming responses as events, not one JSON document. Native SDKs require a route that returns their expected protocol format directly.
| Field | Type | Description |
|---|---|---|
id | string | Unique identifier for the chat completion |
object | string | Object type, always chat.completion |
created | integer | Unix timestamp of when the completion was created |
model | string | Model used for the chat completion |
choices | array | List of chat completion choices |
usage | object | Token usage statistics |
system_fingerprint | string | Backend configuration fingerprint |
service_tier | string | Processing tier used to serve the request |
choices[]
| Field | Type | Description |
|---|---|---|
index | integer | Index of the choice in the list |
message | object | Message generated by the model |
finish_reason | string | Why the model stopped generating tokens |
logprobs | object or null | Log probability information |
message
| Field | Type | Description |
|---|---|---|
role | string | Always assistant |
content | string or null | Generated text content |
refusal | string or null | Refusal message, if any |
finish_reason possible values:
| Value | Description |
|---|---|
stop | Natural stop point or stop sequence reached |
length | Maximum token limit reached |
content_filter | Content omitted due to content filters |
tool_calls | The model requested tool calls; the application must execute them and return results |
usage
| Field | Type | Description |
|---|---|---|
prompt_tokens | integer | Number of tokens in the prompt |
completion_tokens | integer | Number of tokens in the generated completion |
total_tokens | integer | Total tokens used (prompt + completion) |
prompt_tokens_details | object | Breakdown including cached_tokens |
completion_tokens_details | object | Breakdown including reasoning_tokens |
Usage Examples
Basic Conversation
{
"model": "gpt-5",
"messages": [
{"role": "developer", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"}
]
}System Prompt
{
"model": "gpt-5",
"messages": [
{"role": "system", "content": "You are a professional Python programming tutor"},
{"role": "user", "content": "How should I learn programming?"}
]
}Multi-turn Conversation
{
"model": "gpt-5",
"messages": [
{"role": "user", "content": "What is machine learning?"},
{"role": "assistant", "content": "Machine learning is a branch of artificial intelligence..."},
{"role": "user", "content": "Can you give me a practical example?"}
]
}Non-streaming Output
{
"model": "gpt-5",
"messages": [
{"role": "user", "content": "Write a poem about spring"}
],
"stream": false
}Streaming Output
{
"model": "gpt-5",
"messages": [
{"role": "developer", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"}
],
"stream": true
}Request Examples
See language setup. Set API_KEY and replace model, file URL, and ID placeholders first. Each version displays the raw response to the same request.
curl --fail-with-body --silent --show-error --max-time 180 \
--request POST \
--url "https://api.tokatlas.ai/v1/chat/completions" \
--header "Authorization: Bearer $API_KEY" \
--header "Content-Type: application/json" \
--data '{
"model": "gpt-5",
"messages": [
{
"role": "developer",
"content": "You are a helpful assistant."
},
{
"role": "user",
"content": "Hello!"
}
],
"stream": false
}'import os
import requests
headers = {
'Authorization': 'Bearer ' + os.environ["API_KEY"],
'Content-Type': 'application/json',
}
payload = {'model': 'gpt-5',
'messages': [{'role': 'developer', 'content': 'You are a helpful assistant.'},
{'role': 'user', 'content': 'Hello!'}],
'stream': False}
response = requests.request(
'POST', 'https://api.tokatlas.ai/v1/chat/completions', headers=headers,
json=payload,
timeout=180,
)
response.raise_for_status()
print(response.text)if (!process.env.API_KEY) throw new Error("Set API_KEY first.");
const response = await fetch("https://api.tokatlas.ai/v1/chat/completions", {
method: "POST",
headers: {
"Authorization": "Bearer " + process.env.API_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({
"model": "gpt-5",
"messages": [
{
"role": "developer",
"content": "You are a helpful assistant."
},
{
"role": "user",
"content": "Hello!"
}
],
"stream": false
}),
signal: AbortSignal.timeout(180_000),
});
if (!response.ok) {
throw new Error(`HTTP ${response.status}: ${await response.text()}`);
}
console.log(await response.text());import java.net.URI;
import java.net.http.*;
import java.time.Duration;
public class Example {
public static void main(String[] args) throws Exception {
String apiKey = System.getenv("API_KEY");
if (apiKey == null || apiKey.isBlank()) {
throw new IllegalArgumentException("Set API_KEY first.");
}
String payload = String.join("\n",
"{",
" \"model\": \"gpt-5\",",
" \"messages\": [",
" {",
" \"role\": \"developer\",",
" \"content\": \"You are a helpful assistant.\"",
" },",
" {",
" \"role\": \"user\",",
" \"content\": \"Hello!\"",
" }",
" ],",
" \"stream\": false",
"}"
);
HttpClient client = HttpClient.newBuilder()
.connectTimeout(Duration.ofSeconds(30)).build();
HttpRequest request = HttpRequest.newBuilder()
.uri(URI.create("https://api.tokatlas.ai/v1/chat/completions"))
.timeout(Duration.ofSeconds(180))
.header("Authorization", "Bearer " + apiKey)
.header("Content-Type", "application/json")
.method("POST", HttpRequest.BodyPublishers.ofString(payload))
.build();
HttpResponse<String> response = client.send(
request, HttpResponse.BodyHandlers.ofString());
if (response.statusCode() < 200 || response.statusCode() >= 300) {
throw new IllegalStateException("HTTP " + response.statusCode() + ": "
+ response.body());
}
System.out.println(response.body());
}
}package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
)
func main() {
url := "https://api.tokatlas.ai/v1/chat/completions"
payload := map[string]interface{}{
"model": "gpt-5",
"messages": []map[string]string{
{"role": "developer", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"},
},
"stream": false,
}
jsonData, _ := json.Marshal(payload)
req, _ := http.NewRequest("POST", url, bytes.NewBuffer(jsonData))
req.Header.Set("Authorization", "Bearer "+os.Getenv("API_KEY"))
req.Header.Set("Content-Type", "application/json")
resp, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer resp.Body.Close()
body, _ := io.ReadAll(resp.Body)
fmt.Println(string(body))
}Response Examples
Non-streaming (stream: false)
{
"code": 200,
"data": {
"id": "chatcmpl-B9MBs8CjcvOU2jLn4n570S5qMJKcT",
"object": "chat.completion",
"created": 1741569952,
"model": "gpt-5",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I assist you today?",
"refusal": null
},
"logprobs": null,
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 19,
"completion_tokens": 10,
"total_tokens": 29,
"prompt_tokens_details": {
"cached_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 0
}
},
"service_tier": "default"
}
}Streaming (stream: true)
When stream is true, the API returns a Server-Sent Events (SSE) stream with Content-Type: text/event-stream. Each event is a JSON object with object set to chat.completion.chunk. Concatenate the choices[].delta.content fields from each chunk to assemble the full response. The stream ends with data: [DONE].
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-5","choices":[{"index":0,"delta":{"role":"assistant","content":""},"logprobs":null,"finish_reason":null}]}
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-5","choices":[{"index":0,"delta":{"content":"Hello"},"logprobs":null,"finish_reason":null}]}
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-5","choices":[{"index":0,"delta":{"content":"! How can I assist you today?"},"logprobs":null,"finish_reason":null}]}
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-5","choices":[{"index":0,"delta":{},"logprobs":null,"finish_reason":"stop"}]}
data: [DONE]Related Topics
- File Analysis — Image, audio, and file analysis
- Tool Calling — Function calling and tool use
