Skip to content
繁體中文
Response

Create a Response ​

OpenAI-compatible Responses API for text generation and multi-turn conversations.

POST /v1/responses

Creates a model response from text input. Supports single-turn queries, multi-turn conversations, and streaming or non-streaming output.

  • Uses the OpenAI Responses API request format
  • Plain text and structured message input
  • When the route supports storage, retain responses with store and continue with previous_response_id

For image and file analysis, see File Analysis. For tool use, see Tool Calling.

Endpoint ​

text
https://api.tokatlas.ai/v1/responses

Authentication ​

All endpoints require Bearer Token authentication.

http
Authorization: Bearer YOUR_API_KEY
Content-Type: application/json

Important

Never commit a real API key to a repository or expose it in client-side code.

Request Body ​

ParameterTypeRequiredDefaultDescription
modelstringYes-Model that will generate the response
inputstring or arrayYes-Text input to the model
instructionsstringNo-System (or developer) message inserted into the model's context
max_output_tokensintegerNo-Upper bound for tokens generated, including reasoning tokens
streambooleanNofalseWhether to stream the response via SSE
storebooleanNotrueWhether to store the response for later retrieval
temperaturenumberNo1.0Sampling temperature, range 0–2
top_pnumberNo1.0Nucleus sampling parameter, range 0–1
textobjectNo-Text response configuration (format, verbosity)
reasoningobjectNo-Reasoning configuration (gpt-5 and o-series models)
previous_response_idstringNo-ID of the previous response for multi-turn conversations
metadataobjectNo-Up to 16 key-value pairs attached to the response

model ​

Current official model IDs checked on 2026-10-03 are listed below. First query your available models, then use the exact ID in model. An upstream release does not imply access through your Tokatlas account.

  • gpt-6.1-sol — GPT-6.1 Sol
  • gpt-6-astra — GPT-6 Astra
  • gpt-6-luna — GPT-6 Luna

Official model documentation

Use Responses for tool calls with GPT-6.1 Sol and GPT-6 Astra; do not reuse a Chat tool request unchanged.

input ​

Text input to the model. Can be a plain string (equivalent to a user message) or an array of message items.

Plain text input:

json
"What is the weather like in Boston today?"

Message with text content blocks:

json
[
  {
    "role": "user",
    "content": [
      {"type": "input_text", "text": "Hello, please introduce yourself"}
    ]
  }
]

Multi-turn conversation:

json
[
  {"role": "user", "content": [{"type": "input_text", "text": "Hello"}]},
  {"role": "assistant", "content": [{"type": "output_text", "text": "Hi! How can I help you?"}]},
  {"role": "user", "content": [{"type": "input_text", "text": "Tell me about AI"}]}
]

Each message supports role values: user, assistant, system, or developer. Instructions given with developer or system take precedence over user instructions.

instructions ​

A system (or developer) message inserted into the model's context. When used with previous_response_id, instructions from a previous response are not carried over.

json
{
  "instructions": "You are a professional Python programming tutor"
}

max_output_tokens ​

An upper bound for the number of tokens that can be generated for a response, including visible output tokens and reasoning tokens. Minimum: 16.

stream ​

Whether to incrementally stream the response using Server-Sent Events (SSE).

  • true — Streaming response
  • false — Complete response at once (default)

store ​

Whether to store the generated model response for later retrieval via API. Set to false for stateless requests.

temperature ​

Sampling temperature, range 0–2. Higher values (e.g. 0.8) make output more random; lower values (e.g. 0.2) make it more focused. Recommend using either temperature or top_p, not both.

top_p ​

Nucleus sampling parameter, range 0–1. Controls diversity of generated text.

text ​

Configuration for text responses. Defaults to plain text.

json
{
  "text": {
    "format": {"type": "text"},
    "verbosity": "medium"
  }
}

Optional verbosity: low, medium (default), or high.

reasoning ​

Configuration for reasoning models (gpt-5 and o-series only).

FieldTypeDescription
effortstringReasoning effort: none, minimal, low, medium, high, xhigh, max
summarystringReasoning summary: auto, concise, or detailed
json
{
  "reasoning": {
    "effort": "high"
  }
}

previous_response_id ​

The unique ID of the previous response. Use this to create multi-turn conversations without resending full history.

json
{
  "model": "gpt-5",
  "previous_response_id": "resp_abc123",
  "input": "Can you give me a practical example?"
}

metadata ​

Up to 16 key-value pairs (keys max 64 chars, values max 512 chars) for storing additional information about the response.

Response ​

The table describes the protocol response object. Where the non-streaming examples use a { "code": 200, "data": { ... } } envelope, read the object in data; for routes returning the protocol object directly, read the root object. Parse streaming responses as events, not one JSON document. Native SDKs require a route that returns their expected protocol format directly.

FieldTypeDescription
idstringUnique response identifier
objectstringObject type, always response
created_atintegerCreation timestamp
statusstringResponse status: completed, failed, in_progress, cancelled, queued, incomplete
modelstringModel used for the response
outputarrayArray of output items generated by the model
usageobjectToken usage statistics
errorobject or nullError details if the response failed

output[] ​

For text conversations, the primary output type is message:

json
{
  "type": "message",
  "id": "msg_...",
  "status": "completed",
  "role": "assistant",
  "content": [
    {
      "type": "output_text",
      "text": "Hello! How can I help you?",
      "annotations": []
    }
  ]
}

status ​

ValueDescription
completedResponse finished successfully
failedResponse failed with an error
in_progressResponse is still being generated
cancelledResponse was cancelled
queuedResponse is queued
incompleteResponse ended before completion (e.g. hit max_output_tokens)

usage ​

FieldTypeDescription
input_tokensintegerNumber of input tokens
output_tokensintegerNumber of output tokens
total_tokensintegerTotal tokens used
input_tokens_detailsobjectDetails including cached_tokens
output_tokens_detailsobjectDetails including reasoning_tokens

Usage Examples ​

Basic Text Generation ​

json
{
  "model": "gpt-5",
  "input": "Hello, please introduce yourself"
}

System Instructions ​

json
{
  "model": "gpt-5",
  "instructions": "You are a professional Python programming tutor",
  "input": "How should I learn programming?"
}

Multi-turn Conversation ​

json
{
  "model": "gpt-5",
  "previous_response_id": "resp_abc123",
  "input": "Can you give me a practical example?"
}

Non-streaming Output ​

json
{
  "model": "gpt-5",
  "input": "Write a short essay about artificial intelligence",
  "stream": false
}

Streaming Output ​

json
{
  "model": "gpt-5",
  "input": "Write a short essay about artificial intelligence",
  "stream": true
}

Reasoning Model ​

json
{
  "model": "o3-mini",
  "input": "How much wood would a woodchuck chuck?",
  "reasoning": {
    "effort": "high"
  }
}

Request Examples ​

See language setup. Set API_KEY and replace model, file URL, and ID placeholders first. Each version displays the raw response to the same request.

bash
curl --fail-with-body --silent --show-error --max-time 180 \
  --request POST \
  --url "https://api.tokatlas.ai/v1/responses" \
  --header "Authorization: Bearer $API_KEY" \
  --header "Content-Type: application/json" \
  --data '{
  "model": "gpt-5",
  "input": [
    {
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "Tell me about the history of artificial intelligence."
        }
      ]
    }
  ]
}'
python
import os
import requests

headers = {
    'Authorization': 'Bearer ' + os.environ["API_KEY"],
    'Content-Type': 'application/json',
}
payload = {'model': 'gpt-5',
 'input': [{'role': 'user',
            'content': [{'type': 'input_text',
                         'text': 'Tell me about the history of artificial '
                                 'intelligence.'}]}]}
response = requests.request(
    'POST', 'https://api.tokatlas.ai/v1/responses', headers=headers,
    json=payload,
    timeout=180,
)
response.raise_for_status()
print(response.text)
js
if (!process.env.API_KEY) throw new Error("Set API_KEY first.");
const response = await fetch("https://api.tokatlas.ai/v1/responses", {
  method: "POST",
  headers: {
    "Authorization": "Bearer " + process.env.API_KEY,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
  "model": "gpt-5",
  "input": [
    {
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "Tell me about the history of artificial intelligence."
        }
      ]
    }
  ]
}),
  signal: AbortSignal.timeout(180_000),
});
if (!response.ok) {
  throw new Error(`HTTP ${response.status}: ${await response.text()}`);
}
console.log(await response.text());
java
import java.net.URI;
import java.net.http.*;
import java.time.Duration;

public class Example {
    public static void main(String[] args) throws Exception {
        String apiKey = System.getenv("API_KEY");
        if (apiKey == null || apiKey.isBlank()) {
            throw new IllegalArgumentException("Set API_KEY first.");
        }
        String payload = String.join("\n",
            "{",
            "  \"model\": \"gpt-5\",",
            "  \"input\": [",
            "    {",
            "      \"role\": \"user\",",
            "      \"content\": [",
            "        {",
            "          \"type\": \"input_text\",",
            "          \"text\": \"Tell me about the history of artificial intelligence.\"",
            "        }",
            "      ]",
            "    }",
            "  ]",
            "}"
        );
        HttpClient client = HttpClient.newBuilder()
            .connectTimeout(Duration.ofSeconds(30)).build();
        HttpRequest request = HttpRequest.newBuilder()
            .uri(URI.create("https://api.tokatlas.ai/v1/responses"))
            .timeout(Duration.ofSeconds(180))
            .header("Authorization", "Bearer " + apiKey)
            .header("Content-Type", "application/json")
            .method("POST", HttpRequest.BodyPublishers.ofString(payload))
            .build();
        HttpResponse<String> response = client.send(
            request, HttpResponse.BodyHandlers.ofString());
        if (response.statusCode() < 200 || response.statusCode() >= 300) {
            throw new IllegalStateException("HTTP " + response.statusCode() + ": "
                + response.body());
        }
        System.out.println(response.body());
    }
}
go
package main

import (
    "bytes"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
)

func main() {
    url := "https://api.tokatlas.ai/v1/responses"

    payload := map[string]interface{}{
        "model": "gpt-5",
        "input": []map[string]interface{}{
            {
                "role": "user",
                "content": []map[string]string{
                    {
                        "type": "input_text",
                        "text": "Tell me about the history of artificial intelligence.",
                    },
                },
            },
        },
    }

    jsonData, _ := json.Marshal(payload)

    req, _ := http.NewRequest("POST", url, bytes.NewBuffer(jsonData))
    req.Header.Set("Authorization", "Bearer "+os.Getenv("API_KEY"))
    req.Header.Set("Content-Type", "application/json")

    resp, err := http.DefaultClient.Do(req)
    if err != nil {
        panic(err)
    }
    defer resp.Body.Close()

    body, _ := io.ReadAll(resp.Body)
    fmt.Println(string(body))
}

Response Examples ​

Non-streaming (stream: false) ​

json
{
  "code": 200,
  "data": {
    "id": "resp_686eef60237881a2bd1180bb8b13de430e34c516d176ff86",
    "object": "response",
    "created_at": 1752100704,
    "status": "completed",
    "completed_at": 1752100705,
    "model": "gpt-5",
    "output": [
      {
        "id": "msg_686eef60d3e081a29283bdcbc4322fd90e34c516d176ff86",
        "type": "message",
        "status": "completed",
        "role": "assistant",
        "content": [
          {
            "type": "output_text",
            "text": "The history of artificial intelligence (AI) dates back to the 1950s...",
            "annotations": []
          }
        ]
      }
    ],
    "usage": {
      "input_tokens": 28,
      "output_tokens": 320,
      "total_tokens": 348,
      "input_tokens_details": {
        "cached_tokens": 0
      },
      "output_tokens_details": {
        "reasoning_tokens": 0
      }
    },
    "temperature": 1.0,
    "top_p": 1.0,
    "store": true,
    "metadata": {}
  }
}

Streaming (stream: true) ​

When stream is true, the API returns a Server-Sent Events (SSE) stream with Content-Type: text/event-stream. Each event is a JSON object describing a delta in the response. Concatenate response.output_text.delta events to assemble the full text.

text
event: response.created
data: {"type":"response.created","response":{"id":"resp_...","object":"response","status":"in_progress",...}}

event: response.output_item.added
data: {"type":"response.output_item.added","output_index":0,"item":{"type":"message","id":"msg_...","status":"in_progress","role":"assistant","content":[]}}

event: response.output_text.delta
data: {"type":"response.output_text.delta","item_id":"msg_...","output_index":0,"content_index":0,"delta":"The"}

event: response.output_text.delta
data: {"type":"response.output_text.delta","item_id":"msg_...","output_index":0,"content_index":0,"delta":" history"}

event: response.output_text.done
data: {"type":"response.output_text.done","item_id":"msg_...","output_index":0,"content_index":0,"text":"The history of artificial intelligence (AI) dates back to the 1950s..."}

event: response.completed
data: {"type":"response.completed","response":{"id":"resp_...","status":"completed",...}}