Create a Response
POST /v1/responses
Creates a model response from text input. Supports single-turn queries, multi-turn conversations, and streaming or non-streaming output.
- Uses the OpenAI Responses API request format
- Plain text and structured message input
- When the route supports storage, retain responses with
storeand continue withprevious_response_id
For image and file analysis, see File Analysis. For tool use, see Tool Calling.
Endpoint
https://api.tokatlas.ai/v1/responsesAuthentication
All endpoints require Bearer Token authentication.
Authorization: Bearer YOUR_API_KEY
Content-Type: application/jsonImportant
Never commit a real API key to a repository or expose it in client-side code.
Request Body
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
model | string | Yes | - | Model that will generate the response |
input | string or array | Yes | - | Text input to the model |
instructions | string | No | - | System (or developer) message inserted into the model's context |
max_output_tokens | integer | No | - | Upper bound for tokens generated, including reasoning tokens |
stream | boolean | No | false | Whether to stream the response via SSE |
store | boolean | No | true | Whether to store the response for later retrieval |
temperature | number | No | 1.0 | Sampling temperature, range 0–2 |
top_p | number | No | 1.0 | Nucleus sampling parameter, range 0–1 |
text | object | No | - | Text response configuration (format, verbosity) |
reasoning | object | No | - | Reasoning configuration (gpt-5 and o-series models) |
previous_response_id | string | No | - | ID of the previous response for multi-turn conversations |
metadata | object | No | - | Up to 16 key-value pairs attached to the response |
model
Current official model IDs checked on 2026-10-03 are listed below. First query your available models, then use the exact ID in model. An upstream release does not imply access through your Tokatlas account.
gpt-6.1-sol— GPT-6.1 Solgpt-6-astra— GPT-6 Astragpt-6-luna— GPT-6 Luna
Use Responses for tool calls with GPT-6.1 Sol and GPT-6 Astra; do not reuse a Chat tool request unchanged.
input
Text input to the model. Can be a plain string (equivalent to a user message) or an array of message items.
Plain text input:
"What is the weather like in Boston today?"Message with text content blocks:
[
{
"role": "user",
"content": [
{"type": "input_text", "text": "Hello, please introduce yourself"}
]
}
]Multi-turn conversation:
[
{"role": "user", "content": [{"type": "input_text", "text": "Hello"}]},
{"role": "assistant", "content": [{"type": "output_text", "text": "Hi! How can I help you?"}]},
{"role": "user", "content": [{"type": "input_text", "text": "Tell me about AI"}]}
]Each message supports role values: user, assistant, system, or developer. Instructions given with developer or system take precedence over user instructions.
instructions
A system (or developer) message inserted into the model's context. When used with previous_response_id, instructions from a previous response are not carried over.
{
"instructions": "You are a professional Python programming tutor"
}max_output_tokens
An upper bound for the number of tokens that can be generated for a response, including visible output tokens and reasoning tokens. Minimum: 16.
stream
Whether to incrementally stream the response using Server-Sent Events (SSE).
true— Streaming responsefalse— Complete response at once (default)
store
Whether to store the generated model response for later retrieval via API. Set to false for stateless requests.
temperature
Sampling temperature, range 0–2. Higher values (e.g. 0.8) make output more random; lower values (e.g. 0.2) make it more focused. Recommend using either temperature or top_p, not both.
top_p
Nucleus sampling parameter, range 0–1. Controls diversity of generated text.
text
Configuration for text responses. Defaults to plain text.
{
"text": {
"format": {"type": "text"},
"verbosity": "medium"
}
}Optional verbosity: low, medium (default), or high.
reasoning
Configuration for reasoning models (gpt-5 and o-series only).
| Field | Type | Description |
|---|---|---|
effort | string | Reasoning effort: none, minimal, low, medium, high, xhigh, max |
summary | string | Reasoning summary: auto, concise, or detailed |
{
"reasoning": {
"effort": "high"
}
}previous_response_id
The unique ID of the previous response. Use this to create multi-turn conversations without resending full history.
{
"model": "gpt-5",
"previous_response_id": "resp_abc123",
"input": "Can you give me a practical example?"
}metadata
Up to 16 key-value pairs (keys max 64 chars, values max 512 chars) for storing additional information about the response.
Response
The table describes the protocol response object. Where the non-streaming examples use a { "code": 200, "data": { ... } } envelope, read the object in data; for routes returning the protocol object directly, read the root object. Parse streaming responses as events, not one JSON document. Native SDKs require a route that returns their expected protocol format directly.
| Field | Type | Description |
|---|---|---|
id | string | Unique response identifier |
object | string | Object type, always response |
created_at | integer | Creation timestamp |
status | string | Response status: completed, failed, in_progress, cancelled, queued, incomplete |
model | string | Model used for the response |
output | array | Array of output items generated by the model |
usage | object | Token usage statistics |
error | object or null | Error details if the response failed |
output[]
For text conversations, the primary output type is message:
{
"type": "message",
"id": "msg_...",
"status": "completed",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "Hello! How can I help you?",
"annotations": []
}
]
}status
| Value | Description |
|---|---|
completed | Response finished successfully |
failed | Response failed with an error |
in_progress | Response is still being generated |
cancelled | Response was cancelled |
queued | Response is queued |
incomplete | Response ended before completion (e.g. hit max_output_tokens) |
usage
| Field | Type | Description |
|---|---|---|
input_tokens | integer | Number of input tokens |
output_tokens | integer | Number of output tokens |
total_tokens | integer | Total tokens used |
input_tokens_details | object | Details including cached_tokens |
output_tokens_details | object | Details including reasoning_tokens |
Usage Examples
Basic Text Generation
{
"model": "gpt-5",
"input": "Hello, please introduce yourself"
}System Instructions
{
"model": "gpt-5",
"instructions": "You are a professional Python programming tutor",
"input": "How should I learn programming?"
}Multi-turn Conversation
{
"model": "gpt-5",
"previous_response_id": "resp_abc123",
"input": "Can you give me a practical example?"
}Non-streaming Output
{
"model": "gpt-5",
"input": "Write a short essay about artificial intelligence",
"stream": false
}Streaming Output
{
"model": "gpt-5",
"input": "Write a short essay about artificial intelligence",
"stream": true
}Reasoning Model
{
"model": "o3-mini",
"input": "How much wood would a woodchuck chuck?",
"reasoning": {
"effort": "high"
}
}Request Examples
See language setup. Set API_KEY and replace model, file URL, and ID placeholders first. Each version displays the raw response to the same request.
curl --fail-with-body --silent --show-error --max-time 180 \
--request POST \
--url "https://api.tokatlas.ai/v1/responses" \
--header "Authorization: Bearer $API_KEY" \
--header "Content-Type: application/json" \
--data '{
"model": "gpt-5",
"input": [
{
"role": "user",
"content": [
{
"type": "input_text",
"text": "Tell me about the history of artificial intelligence."
}
]
}
]
}'import os
import requests
headers = {
'Authorization': 'Bearer ' + os.environ["API_KEY"],
'Content-Type': 'application/json',
}
payload = {'model': 'gpt-5',
'input': [{'role': 'user',
'content': [{'type': 'input_text',
'text': 'Tell me about the history of artificial '
'intelligence.'}]}]}
response = requests.request(
'POST', 'https://api.tokatlas.ai/v1/responses', headers=headers,
json=payload,
timeout=180,
)
response.raise_for_status()
print(response.text)if (!process.env.API_KEY) throw new Error("Set API_KEY first.");
const response = await fetch("https://api.tokatlas.ai/v1/responses", {
method: "POST",
headers: {
"Authorization": "Bearer " + process.env.API_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({
"model": "gpt-5",
"input": [
{
"role": "user",
"content": [
{
"type": "input_text",
"text": "Tell me about the history of artificial intelligence."
}
]
}
]
}),
signal: AbortSignal.timeout(180_000),
});
if (!response.ok) {
throw new Error(`HTTP ${response.status}: ${await response.text()}`);
}
console.log(await response.text());import java.net.URI;
import java.net.http.*;
import java.time.Duration;
public class Example {
public static void main(String[] args) throws Exception {
String apiKey = System.getenv("API_KEY");
if (apiKey == null || apiKey.isBlank()) {
throw new IllegalArgumentException("Set API_KEY first.");
}
String payload = String.join("\n",
"{",
" \"model\": \"gpt-5\",",
" \"input\": [",
" {",
" \"role\": \"user\",",
" \"content\": [",
" {",
" \"type\": \"input_text\",",
" \"text\": \"Tell me about the history of artificial intelligence.\"",
" }",
" ]",
" }",
" ]",
"}"
);
HttpClient client = HttpClient.newBuilder()
.connectTimeout(Duration.ofSeconds(30)).build();
HttpRequest request = HttpRequest.newBuilder()
.uri(URI.create("https://api.tokatlas.ai/v1/responses"))
.timeout(Duration.ofSeconds(180))
.header("Authorization", "Bearer " + apiKey)
.header("Content-Type", "application/json")
.method("POST", HttpRequest.BodyPublishers.ofString(payload))
.build();
HttpResponse<String> response = client.send(
request, HttpResponse.BodyHandlers.ofString());
if (response.statusCode() < 200 || response.statusCode() >= 300) {
throw new IllegalStateException("HTTP " + response.statusCode() + ": "
+ response.body());
}
System.out.println(response.body());
}
}package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
)
func main() {
url := "https://api.tokatlas.ai/v1/responses"
payload := map[string]interface{}{
"model": "gpt-5",
"input": []map[string]interface{}{
{
"role": "user",
"content": []map[string]string{
{
"type": "input_text",
"text": "Tell me about the history of artificial intelligence.",
},
},
},
},
}
jsonData, _ := json.Marshal(payload)
req, _ := http.NewRequest("POST", url, bytes.NewBuffer(jsonData))
req.Header.Set("Authorization", "Bearer "+os.Getenv("API_KEY"))
req.Header.Set("Content-Type", "application/json")
resp, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer resp.Body.Close()
body, _ := io.ReadAll(resp.Body)
fmt.Println(string(body))
}Response Examples
Non-streaming (stream: false)
{
"code": 200,
"data": {
"id": "resp_686eef60237881a2bd1180bb8b13de430e34c516d176ff86",
"object": "response",
"created_at": 1752100704,
"status": "completed",
"completed_at": 1752100705,
"model": "gpt-5",
"output": [
{
"id": "msg_686eef60d3e081a29283bdcbc4322fd90e34c516d176ff86",
"type": "message",
"status": "completed",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "The history of artificial intelligence (AI) dates back to the 1950s...",
"annotations": []
}
]
}
],
"usage": {
"input_tokens": 28,
"output_tokens": 320,
"total_tokens": 348,
"input_tokens_details": {
"cached_tokens": 0
},
"output_tokens_details": {
"reasoning_tokens": 0
}
},
"temperature": 1.0,
"top_p": 1.0,
"store": true,
"metadata": {}
}
}Streaming (stream: true)
When stream is true, the API returns a Server-Sent Events (SSE) stream with Content-Type: text/event-stream. Each event is a JSON object describing a delta in the response. Concatenate response.output_text.delta events to assemble the full text.
event: response.created
data: {"type":"response.created","response":{"id":"resp_...","object":"response","status":"in_progress",...}}
event: response.output_item.added
data: {"type":"response.output_item.added","output_index":0,"item":{"type":"message","id":"msg_...","status":"in_progress","role":"assistant","content":[]}}
event: response.output_text.delta
data: {"type":"response.output_text.delta","item_id":"msg_...","output_index":0,"content_index":0,"delta":"The"}
event: response.output_text.delta
data: {"type":"response.output_text.delta","item_id":"msg_...","output_index":0,"content_index":0,"delta":" history"}
event: response.output_text.done
data: {"type":"response.output_text.done","item_id":"msg_...","output_index":0,"content_index":0,"text":"The history of artificial intelligence (AI) dates back to the 1950s..."}
event: response.completed
data: {"type":"response.completed","response":{"id":"resp_...","status":"completed",...}}Related Topics
- File Analysis — Image, file, and document analysis
- Tool Calling — Custom function calling and returning results
