建立聊天回應
POST /v1/chat/completions
針對指定對話產生模型回應。傳回聊天回應物件;若啟用串流,則傳回一系列聊天回應片段。
- 採用 OpenAI Chat Completions API 請求格式
- 支援純文字訊息與多輪對話
- 支援串流與非串流回應
圖像、音訊與檔案分析請參閱檔案分析;工具使用請參閱工具呼叫。
端點
https://api.tokatlas.ai/v1/chat/completions身分驗證
所有端點均需要 Bearer Token 驗證。
Authorization: Bearer YOUR_API_KEY
Content-Type: application/json重要
請勿將真實 API 金鑰提交至程式碼儲存庫,或公開於用戶端程式碼中。
請求本文
| 參數 | 類型 | 必填 | 預設值 | 說明 |
|---|---|---|---|---|
model | string | 是 | - | 產生回應的模型 ID |
messages | array | 是 | - | 組成對話的訊息清單 |
temperature | number | 否 | 1.0 | 取樣溫度,範圍 0–2 |
top_p | number | 否 | 1.0 | 核心取樣參數,範圍 0–1 |
max_tokens | integer | 否 | - | 生成 Token 上限(o 系列已棄用,請使用 max_completion_tokens) |
max_completion_tokens | integer | 否 | - | 生成 Token 上限,包含推理 Token |
stream | boolean | 否 | false | 是否透過 SSE 傳回串流回應 |
stream_options | object | 否 | - | 串流選項(僅適用於 stream: true) |
stop | string or array | 否 | - | 停止生成的序列,最多 4 組 |
n | integer | 否 | 1 | 要產生的聊天回應選項數量 |
frequency_penalty | number | 否 | 0 | 頻率懲罰,範圍 -2.0 至 2.0 |
presence_penalty | number | 否 | 0 | 出現懲罰,範圍 -2.0 至 2.0 |
response_format | object | 否 | - | 輸出格式:文字或 JSON Schema |
reasoning_effort | string | 否 | - | 推理模型的推理程度 |
metadata | object | 否 | - | 物件附加的鍵值對,最多 16 組 |
store | boolean | 否 | - | 是否儲存輸出以供日後取得 |
model
先取得帳戶可用模型,確認所選模型支援 Chat Completions。最新型號請見 OpenAI、Claude 與 Gemini 官方列表。
例如 GPT-6.1 Sol、GPT-6 Astra 的純文字可使用 Chat,但工具呼叫需使用 Responses。跨供應商模型還需要 Tokatlas 路由提供協定轉換;不能只憑模型名稱判定相容。
目前官方型號包括:
- OpenAI:
gpt-6.1-sol,gpt-6-astra,gpt-6-luna - Anthropic:
claude-opus-5-5,claude-sonnet-5-5,claude-fable-5-1 - Google:
gemini-3.8-flash,gemini-3.5-flash-lite
messages
目前對話的訊息清單,每則訊息均包含 role 和 content。
| 角色 | 說明 |
|---|---|
developer | 開發者提供的指令。o1 及後續模型以 developer 訊息取代 system 訊息 |
system | 設定 AI 行為的系統提示(o1 及後續模型請使用 developer) |
user | 終端使用者傳送的訊息 |
assistant | 模型回覆使用者的訊息 |
tool | 應用程式回傳的工具結果,須包含對應的 tool_call_id |
基本使用者訊息:
[{"role": "user", "content": "Hello!"}]開發者/系統提示:
[
{"role": "developer", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"}
]多輪對話:
[
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi! How can I help you?"},
{"role": "user", "content": "Tell me about AI"}
]文字訊息的 content 可以是純字串,也可以是文字內容部分的陣列:
{"role": "user", "content": [{"type": "text", "text": "Hello!"}]}temperature
取樣溫度介於 0 至 2。較高值(如 0.8)增加隨機性,較低值(如 0.2)讓輸出更集中。建議只調整 temperature 或 top_p 其中一項。
top_p
核心取樣參數,範圍為 0–1。
max_tokens / max_completion_tokens
max_tokens:生成 Token 上限。已棄用,請改用max_completion_tokens;不相容 o 系列模型。max_completion_tokens:生成 Token 上限,包含可見輸出與推理 Token。
stream
設為 true 時,模型透過 SSE 邊生成邊傳回回應。每個片段的 object 均為 chat.completion.chunk。
true:串流回應false:一次傳回完整回應(預設)
stream_options
串流回應選項,僅於 stream: true 時設定。
| 欄位 | 類型 | 說明 |
|---|---|---|
include_usage | boolean | 在 data: [DONE] 之前傳回包含 Token 用量的最後片段 |
include_obfuscation | boolean | 加入混淆欄位以統一負載大小(預設為 true) |
stop
最多 4 組停止序列,符合時 API 停止生成 Token。推理模型 o3 和 o4-mini 不支援此參數。
n
每則輸入訊息要產生的回應選項數量。將 n 保持為 1 可降低成本。
frequency_penalty / presence_penalty
介於 -2.0 至 2.0 的數值。
frequency_penalty:依 Token 在文字中出現的頻率施加懲罰presence_penalty:依 Token 是否已出現在文字中施加懲罰
response_format
指定模型輸出格式的物件。
純文字(預設):
{"type": "text"}結構化 JSON 輸出(JSON Schema):
{
"type": "json_schema",
"json_schema": {
"name": "math_response",
"schema": {
"type": "object",
"properties": {
"steps": {"type": "array", "items": {"type": "string"}},
"final_answer": {"type": "string"}
},
"required": ["steps", "final_answer"],
"additionalProperties": false
},
"strict": true
}
}reasoning_effort
限制推理模型的推理程度。支援值為 none、minimal、low、medium、high、xhigh、max。
{
"model": "o3-mini",
"messages": [{"role": "user", "content": "Solve this math problem"}],
"reasoning_effort": "high"
}metadata
最多 16 組鍵值對,用於儲存物件的額外資訊。鍵最長 64 個字元,值最長 512 個字元。
回應
下表描述協定回應物件。本文非串流範例若有 { "code": 200, "data": { ... } } 外層封裝,請先讀取 data;直接回傳協定物件的路由則讀取根物件。串流請按事件解析,不能將整段 SSE 當成單一 JSON。使用原生 SDK 前,須確認路由回傳 SDK 所需的直接協定格式。
| 欄位 | 類型 | 說明 |
|---|---|---|
id | string | 聊天回應的唯一識別碼 |
object | string | 物件類型,固定為 chat.completion |
created | integer | 回應建立時的 Unix 時間戳記 |
model | string | 聊天回應使用的模型 |
choices | array | 聊天回應選項清單 |
usage | object | Token 用量統計 |
system_fingerprint | string | 後端設定指紋 |
service_tier | string | 處理請求的服務層級 |
choices[]
| 欄位 | 類型 | 說明 |
|---|---|---|
index | integer | 選項在清單中的索引 |
message | object | 模型產生的訊息 |
finish_reason | string | 模型停止生成 Token 的原因 |
logprobs | object or null | 對數機率資訊 |
message
| 欄位 | 類型 | 說明 |
|---|---|---|
role | string | 固定為 assistant |
content | string or null | 生成的文字內容 |
refusal | string or null | 拒絕訊息(如有) |
finish_reason 的可能值:
| 值 | 說明 |
|---|---|
stop | 自然停止或符合停止序列 |
length | 已達 Token 數量上限 |
content_filter | 因內容篩選而省略內容 |
tool_calls | 模型提出工具呼叫,應用程式須執行並回傳結果 |
usage
| 欄位 | 類型 | 說明 |
|---|---|---|
prompt_tokens | integer | 提示中的 Token 數量 |
completion_tokens | integer | 生成回應中的 Token 數量 |
total_tokens | integer | Token 總用量(提示 + 回應) |
prompt_tokens_details | object | 包含 cached_tokens 的用量明細 |
completion_tokens_details | object | 包含 reasoning_tokens 的用量明細 |
使用範例
基本對話
{
"model": "gpt-5",
"messages": [
{"role": "developer", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"}
]
}系統提示
{
"model": "gpt-5",
"messages": [
{"role": "system", "content": "You are a professional Python programming tutor"},
{"role": "user", "content": "How should I learn programming?"}
]
}多輪對話
{
"model": "gpt-5",
"messages": [
{"role": "user", "content": "What is machine learning?"},
{"role": "assistant", "content": "Machine learning is a branch of artificial intelligence..."},
{"role": "user", "content": "Can you give me a practical example?"}
]
}非串流輸出
{
"model": "gpt-5",
"messages": [
{"role": "user", "content": "Write a poem about spring"}
],
"stream": false
}串流輸出
{
"model": "gpt-5",
"messages": [
{"role": "developer", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"}
],
"stream": true
}請求範例
執行方式見多語言範例說明。先設定 API_KEY,並替換模型、檔案網址及 ID 占位值;四種方式會顯示相同請求的原始回應。
curl --fail-with-body --silent --show-error --max-time 180 \
--request POST \
--url "https://api.tokatlas.ai/v1/chat/completions" \
--header "Authorization: Bearer $API_KEY" \
--header "Content-Type: application/json" \
--data '{
"model": "gpt-5",
"messages": [
{
"role": "developer",
"content": "You are a helpful assistant."
},
{
"role": "user",
"content": "Hello!"
}
],
"stream": false
}'import os
import requests
headers = {
'Authorization': 'Bearer ' + os.environ["API_KEY"],
'Content-Type': 'application/json',
}
payload = {'model': 'gpt-5',
'messages': [{'role': 'developer', 'content': 'You are a helpful assistant.'},
{'role': 'user', 'content': 'Hello!'}],
'stream': False}
response = requests.request(
'POST', 'https://api.tokatlas.ai/v1/chat/completions', headers=headers,
json=payload,
timeout=180,
)
response.raise_for_status()
print(response.text)if (!process.env.API_KEY) throw new Error("Set API_KEY first.");
const response = await fetch("https://api.tokatlas.ai/v1/chat/completions", {
method: "POST",
headers: {
"Authorization": "Bearer " + process.env.API_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({
"model": "gpt-5",
"messages": [
{
"role": "developer",
"content": "You are a helpful assistant."
},
{
"role": "user",
"content": "Hello!"
}
],
"stream": false
}),
signal: AbortSignal.timeout(180_000),
});
if (!response.ok) {
throw new Error(`HTTP ${response.status}: ${await response.text()}`);
}
console.log(await response.text());import java.net.URI;
import java.net.http.*;
import java.time.Duration;
public class Example {
public static void main(String[] args) throws Exception {
String apiKey = System.getenv("API_KEY");
if (apiKey == null || apiKey.isBlank()) {
throw new IllegalArgumentException("Set API_KEY first.");
}
String payload = String.join("\n",
"{",
" \"model\": \"gpt-5\",",
" \"messages\": [",
" {",
" \"role\": \"developer\",",
" \"content\": \"You are a helpful assistant.\"",
" },",
" {",
" \"role\": \"user\",",
" \"content\": \"Hello!\"",
" }",
" ],",
" \"stream\": false",
"}"
);
HttpClient client = HttpClient.newBuilder()
.connectTimeout(Duration.ofSeconds(30)).build();
HttpRequest request = HttpRequest.newBuilder()
.uri(URI.create("https://api.tokatlas.ai/v1/chat/completions"))
.timeout(Duration.ofSeconds(180))
.header("Authorization", "Bearer " + apiKey)
.header("Content-Type", "application/json")
.method("POST", HttpRequest.BodyPublishers.ofString(payload))
.build();
HttpResponse<String> response = client.send(
request, HttpResponse.BodyHandlers.ofString());
if (response.statusCode() < 200 || response.statusCode() >= 300) {
throw new IllegalStateException("HTTP " + response.statusCode() + ": "
+ response.body());
}
System.out.println(response.body());
}
}package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
)
func main() {
url := "https://api.tokatlas.ai/v1/chat/completions"
payload := map[string]interface{}{
"model": "gpt-5",
"messages": []map[string]string{
{"role": "developer", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"},
},
"stream": false,
}
jsonData, _ := json.Marshal(payload)
req, _ := http.NewRequest("POST", url, bytes.NewBuffer(jsonData))
req.Header.Set("Authorization", "Bearer "+os.Getenv("API_KEY"))
req.Header.Set("Content-Type", "application/json")
resp, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer resp.Body.Close()
body, _ := io.ReadAll(resp.Body)
fmt.Println(string(body))
}回應範例
非串流 (stream: false)
{
"code": 200,
"data": {
"id": "chatcmpl-B9MBs8CjcvOU2jLn4n570S5qMJKcT",
"object": "chat.completion",
"created": 1741569952,
"model": "gpt-5",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I assist you today?",
"refusal": null
},
"logprobs": null,
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 19,
"completion_tokens": 10,
"total_tokens": 29,
"prompt_tokens_details": {
"cached_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 0
}
},
"service_tier": "default"
}
}串流 (stream: true)
當 stream 為 true,API 會傳回 Content-Type: text/event-stream 的 SSE 串流。每個事件都是 object 為 chat.completion.chunk 的 JSON 物件。串接各片段的 choices[].delta.content 即可取得完整回應。串流以 data: [DONE] 結束。
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-5","choices":[{"index":0,"delta":{"role":"assistant","content":""},"logprobs":null,"finish_reason":null}]}
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-5","choices":[{"index":0,"delta":{"content":"Hello"},"logprobs":null,"finish_reason":null}]}
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-5","choices":[{"index":0,"delta":{"content":"! How can I assist you today?"},"logprobs":null,"finish_reason":null}]}
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-5","choices":[{"index":0,"delta":{},"logprobs":null,"finish_reason":"stop"}]}
data: [DONE]