> ## Documentation Index
> Fetch the complete documentation index at: https://speshu.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Responses

> POST /api/v1/responses — Responses API: элементы вывода, инструменты, поток событий

Совместим с OpenAI Responses API: официальный SDK работает без изменений — достаточно поменять `base_url`. В отличие от Chat Completions, ответ собирается в плоский список `output` — сообщения, вызовы функций и вызовы встроенных инструментов лежат в одном массиве и различаются по полю `type`.

<Note>
  Если вам не нужен список элементов вывода, используйте [Запрос чату](/docs/api-reference/chat/completions) — формат проще, а цены и тарификация те же.
</Note>

## Минимальный запрос

<CodeGroup>
  ```bash cURL theme={null} theme={null}
  curl https://speshu.ai/api/v1/responses \
    -H "Authorization: Bearer $SPESHU_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "gpt-4o",
      "input": "Привет!"
    }'
  ```

  ```python Python theme={null} theme={null}
  from openai import OpenAI

  client = OpenAI(base_url="https://speshu.ai/api/v1", api_key="sk-...")

  response = client.responses.create(
      model="gpt-4o",
      input="Привет!",
  )
  print(response.output_text)
  print(response.usage.cost.total_cost, "₽")
  ```

  ```typescript TypeScript theme={null} theme={null}
  import OpenAI from "openai";

  const client = new OpenAI({
    baseURL: "https://speshu.ai/api/v1",
    apiKey: "sk-...",
  });

  const response = await client.responses.create({
    model: "gpt-4o",
    input: "Привет!",
  });
  console.log(response.output_text);
  ```
</CodeGroup>

<Note>
  Идентификатор модели нужно взять из `GET /api/v1/models` и подставить **как есть**. Если модель неизвестна, придёт `404` с кодом `model_not_found`. Идентификатор в примерах ниже конкретный, чтобы их можно было скопировать; в рабочем коде получайте его из каталога — состав моделей меняется.
</Note>

## Поля запроса

### Обязательные

| Поле | Тип | Описание |
| - | - | - |
| `model` | string | Идентификатор из `GET /api/v1/models` |
| `input` | string \| array | Строка с запросом либо массив сообщений |
| `stream` | boolean | Потоковая передача через SSE |

<Warning>
  Если `input` не передан или равен пустой строке, придёт `400 input is required`. `instructions` не заменяет `input` — хотя бы одно из них всегда нужно.
</Warning>

`input` как массив — это обычные сообщения с `role` и `content`. Поддерживаются роли `user`, `assistant`, `system`, `developer`.

### Генерация

| Поле | Тип | Описание |
| - | - | - |
| `instructions` | string | Системная инструкция к текущему ответу |
| `max_output_tokens` | integer | Потолок выходных токенов |
| `temperature` | number | |
| `top_p` | number | |
| `top_logprobs` | integer | Сколько альтернативных токенов возвращать в логвероятностях |
| `truncation` | string | `auto` или `disabled` |
| `service_tier` | string | `auto`, `default`, `flex`, `priority` |
| `metadata` | object | Произвольные данные «ключ → значение» |
| `store` | boolean | |
| `background` | boolean | |
| `conversation` | object | Привязка к диалогу вместо `previous_response_id` |
| `previous_response_id` | string | Продолжение предыдущего ответа |
| `max_tool_calls` | integer | Потолок числа вызовов инструментов |
| `stream_options` | object | |

### Структурированный вывод

Поле `text` задаёт формат ответа и подробность (поле `verbosity`):

```json theme={null} theme={null}
{
  "text": {
    "format": {
      "type": "json_schema",
      "name": "order",
      "description": "Заказ клиента",
      "schema": {
        "type": "object",
        "properties": {
          "sku": { "type": "string" },
          "qty": { "type": "integer" }
        },
        "required": ["sku", "qty"]
      },
      "strict": true
    },
    "verbosity": "low"
  }
}
```

`format.type` — `text`, `json_object` или `json_schema`. Для `json_schema` обязательны `name` и `schema`; `description` и `strict` необязательны.

### Инструменты

```json theme={null} theme={null}
{
  "tools": [
    {
      "type": "function",
      "name": "get_weather",
      "description": "Погода в городе",
      "parameters": {
        "type": "object",
        "properties": { "city": { "type": "string" } },
        "required": ["city"]
      }
    }
  ],
  "tool_choice": "auto"
}
```

| Поле | Тип | Описание |
| - | - | - |
| `tool_choice` | string \| object | `none`, `auto`, `required`, `any` либо объект с конкретным инструментом |
| `parallel_tool_calls` | boolean | Разрешить несколько вызовов за один шаг |

Поддерживаемые `type` в `tools[]`:

| Тип | Что это |
| - | - |
| `function` | Ваша функция |
| `custom` | Ваш инструмент произвольного вида |
| `file_search` | Поиск по загруженным файлам |
| `code_interpreter` | Выполнение кода |
| `web_search`, `web_search_2025_08_26`, `web_search_preview`, `x_search` | Веб-поиск и поиск по X |
| `computer_use_preview` | Управление браузером |
| `mcp` | Внешний MCP-сервер |
| `image_generation` | Генерация изображений |
| `local_shell` | Локальный шелл |
| `tool_search` | Поиск по доступным инструментам |

### Рассуждения

```json theme={null} theme={null}
{
  "reasoning": {
    "effort": "high",
    "summary": "auto",
    "max_tokens": 2048
  }
}
```

`effort` — `none`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max`. Значение, отличное от `none`, включает рассуждение. Текст рассуждений приходит событиями `response.reasoning_summary_text.delta` и попадает в `output_tokens_details.reasoning_tokens`.

### Включить в ответ дополнительные поля

`include` — массив строк. Например:

```json theme={null} theme={null}
{ "include": ["web_search_call.action.sources", "message.output_text.logprobs", "reasoning.encrypted_content"] }
```

## Ответ

```json theme={null} theme={null}
{
  "id": "resp_...",
  "object": "response",
  "created_at": 1756800000,
  "completed_at": 1756800004,
  "status": "completed",
  "error": null,
  "incomplete_details": null,
  "model": "gpt-4o",
  "instructions": null,
  "max_output_tokens": null,
  "output": [
    {
      "id": "msg_...",
      "type": "message",
      "status": "completed",
      "role": "assistant",
      "content": [
        { "type": "output_text", "text": "Привет! Чем помочь?", "annotations": [] }
      ]
    }
  ],
  "usage": {
    "input_tokens": 9,
    "output_tokens": 7,
    "total_tokens": 16,
    "input_tokens_details": { "cached_tokens": 0, "text_tokens": 9, "image_tokens": 0 },
    "output_tokens_details": { "reasoning_tokens": 0, "text_tokens": 7 },
    "cost": { "input_cost": 0.0002, "output_cost": 0.0006, "total_cost": 0.0008 }
  }
}
```

Поле `model` повторяет запрошенный идентификатор. `status` — одно из `completed`, `in_progress`, `incomplete`, `failed`, `cancelled`, `queued`.

<Warning>
  Текст ответа лежит не на верхнем уровне, а внутри `output`. Обычно это `output[0].content[0].text`, но элемент с вызовом функции тоже может стоять первым — ищите текст перебором всех элементов, а не по индексу. В SDK за это отвечает `response.output_text`.
</Warning>

### Стоимость

`usage.cost.total_cost` — сумма, **списанная с вашего кошелька в рублях**, а не себестоимость у провайдера. Именно она совпадает с `cost_context` и `cost_completion` из `GET /api/v1/models`.

## Потоковая передача

При `stream: true` приходит поток SSE. Каждый кадр — безымянное событие `data:`; отдельного поля `event:` сервер не пишет, поэтому клиенты, разбирающие поток по `event:`, работать не будут. Тип события лежит в поле `type` самого объекта.

```
data: {"type":"response.created","sequence_number":0,"response":{"id":"resp_...","object":"response","status":"in_progress","output":[]}}

data: {"type":"response.output_item.added","sequence_number":1,"output_index":0,"item":{"id":"msg_...","type":"message","status":"in_progress","role":"assistant","content":[]}}

data: {"type":"response.output_text.delta","sequence_number":2,"output_index":0,"content_index":0,"item_id":"msg_...","delta":"Привет"}

data: {"type":"response.output_text.delta","sequence_number":3,"output_index":0,"content_index":0,"item_id":"msg_...","delta":"!"}

data: {"type":"response.output_text.done","sequence_number":4,"output_index":0,"content_index":0,"item_id":"msg_...","text":"Привет!"}

data: {"type":"response.output_item.done","sequence_number":5,"output_index":0,"item":{"id":"msg_...","type":"message","status":"completed","role":"assistant","content":[{"type":"output_text","text":"Привет!"}]}}

data: {"type":"response.completed","sequence_number":6,"response":{"id":"resp_...","object":"response","status":"completed","output":[],"usage":{"input_tokens":9,"output_tokens":7,"total_tokens":16,"cost":{"total_cost":0.0008}}}}

data: [DONE]
```

### События

Жизненный цикл ответа:

| Событие | Когда приходит |
| - | - |
| `response.created` | Ответ создан |
| `response.in_progress` | Ответ взят в работу |
| `response.completed` | Ответ успешно завершён |
| `response.failed` | Ответ завершился ошибкой |
| `response.incomplete` | Ответ прерван по лимиту токенов |
| `response.ping` | Проверка живости соединения |
| `error` | Ошибка в середине потока |

Элементы вывода и их части:

| Событие | Когда приходит |
| - | - |
| `response.output_item.added` | Начался новый элемент `output` |
| `response.output_item.done` | Элемент `output` готов целиком |
| `response.content_part.added` | Началась новая часть контента элемента |
| `response.content_part.done` | Часть контента закрыта |

Текст:

| Событие | Когда приходит |
| - | - |
| `response.output_text.delta` | Очередной кусок текста |
| `response.output_text.done` | Текст элемента готов целиком |
| `response.output_text.annotation.added` | Ссылка в тексте (например, источник поиска) |
| `response.refusal.delta` | Модель отказалась отвечать |

Инструменты:

| Событие | Когда приходит |
| - | - |
| `response.function_call_arguments.delta` | Очередной кусок аргументов функции |
| `response.function_call_arguments.done` | Аргументы функции готовы |
| `response.web_search_call.in_progress` | Идёт веб-поиск |
| `response.file_search_call.in_progress` | Идёт поиск по файлам |
| `response.code_interpreter_call.*` | Шаги интерпретатора кода |
| `response.mcp_call.*` | Обращение к MCP-серверу |
| `response.image_generation_call.partial_image` | Промежуточная картинка |

Рассуждения:

| Событие | Когда приходит |
| - | - |
| `response.reasoning_summary_text.delta` | Очередной кусок пересказа рассуждений |
| `response.reasoning_summary_part.added` | Новая часть пересказа рассуждений |

Каждое событие несёт `sequence_number` — номер по порядку. По нему можно отсеять дубли и поймать пропуск кадров. Индексы `output_index` и `content_index` адресуют конкретный элемент и часть контента внутри него.

<Warning>
  Деньги списываются на событии `response.completed`: именно в его поле `response.usage.cost.total_cost` лежит сумма в рублях. Если поток оборвался раньше, списания по этому кадру не будет — сверяйтесь с балансом через `GET /api/v1/balance`.
</Warning>

<Warning>
  После последнего события сервер дописывает кадр `data: [DONE]`. Его **нет** в спецификации OpenAI Responses API: строгий клиент попытается распарсить `[DONE]` как JSON и упадёт. Обработайте эту строку как конец потока и завершайте чтение.
</Warning>

Если провайдер оборвал поток, в середине придёт событие `error` с объектом `error` внутри, а `response.completed` и `[DONE]` не придут вовсе:

```
data: {"type":"error","sequence_number":7,"error":{"message":"...","type":"server_error"}}
```

Обрабатывайте неизвестные типы событий без ошибки: новые события появляются без изменения основного контракта.

## Проверки перед вызовом

Часть ошибок возвращается до обращения к провайдеру — в формате OpenAI:

```json theme={null} theme={null}
{
  "error": {
    "message": "input is required",
    "type": "invalid_request_error",
    "param": "input",
    "code": "invalid_request"
  }
}
```

| Ситуация | Ответ |
| - | - |
| Нет ключа | `401 invalid_api_key` |
| Модель неизвестна | `404 model_not_found` |
| Не хватает денег | `402 insufficient_quota` |
| Нет `input` | `400 input is required` |
| Неподдерживаемый `reasoning.effort` | `400 invalid_request` |
| Входные токены плюс резерв выхода не помещаются в окно | `422 content_too_large` |
| Провайдер недоступен | `503 service_unavailable` |

Полная таблица кодов — на странице [Ошибки](/docs/errors).

## Что дальше

<CardGroup cols={2}>
  <Card title="Запрос чату" icon="message" href="/docs/api-reference/chat/completions">
    Тот же набор моделей в формате Chat Completions
  </Card>

  <Card title="Anthropic Messages" icon="a" href="/docs/api-reference/messages/create">
    Совместимость с Anthropic SDK
  </Card>

  <Card title="Каталог моделей" icon="microchip" href="/docs/models">
    Идентификаторы и цены
  </Card>

  <Card title="Ошибки" icon="triangle-exclamation" href="/docs/errors">
    Полная таблица кодов
  </Card>
</CardGroup>


## OpenAPI

````yaml api-reference/openapi.json POST /api/v1/responses
openapi: 3.1.0
info:
  title: SpeShu.AI API
  version: 1.0.0
  description: >-
    Единый OpenAI-совместимый API для текстовых моделей, генерации изображений,
    видео, аудио и музыки.


    ## Аутентификация


    Все запросы, кроме `GET /api/v1/models` и `GET /api/v1/media/models`,
    требуют API-ключ. Ключ передаётся в заголовке `Authorization: Bearer <ключ>`
    или `X-Api-Key: <ключ>`.


    ## Ошибки


    Текстовые эндпоинты (`/chat/completions`, `/responses`, `/messages`) и
    `/balance` возвращают ошибки в формате OpenAI:


    ```json

    { "error": { "message": "...", "type": "invalid_request_error", "param":
    null, "code": "invalid_request" } }

    ```


    Эндпоинты медиа используют собственный конверт:


    ```json

    { "code": 422, "msg": "...", "data": null }

    ```


    Подробности — на странице [Ошибки](/errors).
servers:
  - url: https://speshu.ai
    description: Production
security:
  - BearerAuth: []
tags:
  - name: Текст
    description: Генерация текста и диалог
  - name: Медиа
    description: Асинхронные задачи генерации изображений, видео, аудио и музыки
  - name: Хранилище
    description: Файлы пользователя
  - name: Справочник
    description: Каталог моделей и баланс
paths:
  /api/v1/responses:
    post:
      tags:
        - Текст
      summary: Responses API
      description: >-
        Совместимость с OpenAI Responses API. Возвращает ответ в виде объекта
        `output` с элементами сообщений и вызовов инструментов.
      operationId: createResponse
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/CreateResponseRequest'
      responses:
        '200':
          description: >-
            Успешный ответ. При `stream: true` — поток SSE с событиями Responses
            API, завершающийся кадром `data: [DONE]`.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ResponseObject'
            text/event-stream:
              schema:
                type: object
                properties:
                  type:
                    type: string
                    description: Тип события, например `response.output_text.delta`.
                  sequence_number:
                    type: integer
                  delta:
                    type: string
                  response:
                    $ref: '#/components/schemas/ResponseObject'
                required:
                  - type
        '400':
          $ref: '#/components/responses/OpenAIError'
        '401':
          $ref: '#/components/responses/OpenAIError'
        '402':
          $ref: '#/components/responses/OpenAIError'
        '422':
          $ref: '#/components/responses/OpenAIError'
        '429':
          $ref: '#/components/responses/OpenAIError'
        '500':
          $ref: '#/components/responses/OpenAIError'
        '503':
          $ref: '#/components/responses/OpenAIError'
components:
  schemas:
    CreateResponseRequest:
      type: object
      properties:
        model:
          type: string
          examples:
            - gpt-4o
        input:
          description: Строка с запросом либо массив сообщений.
          oneOf:
            - type: string
            - type: array
              items:
                type: object
                additionalProperties: true
        stream:
          type: boolean
          default: false
        instructions:
          type: string
        max_output_tokens:
          type: integer
        max_tool_calls:
          type: integer
        temperature:
          type: number
        top_p:
          type: number
        text:
          type: object
          additionalProperties: true
        tools:
          type: array
          items:
            type: object
            additionalProperties: true
        tool_choice:
          oneOf:
            - type: string
              enum:
                - none
                - auto
                - required
                - any
            - type: object
              additionalProperties: true
        parallel_tool_calls:
          type: boolean
        reasoning:
          type: object
          properties:
            effort:
              type: string
              enum:
                - none
                - minimal
                - low
                - medium
                - high
                - xhigh
                - max
            summary:
              type: string
            max_tokens:
              type: integer
        previous_response_id:
          type: string
        conversation:
          type: object
          additionalProperties: true
        truncation:
          type: string
          enum:
            - auto
            - disabled
        include:
          type: array
          items:
            type: string
        store:
          type: boolean
        metadata:
          type: object
          additionalProperties: true
      required:
        - model
        - input
    ResponseObject:
      type: object
      properties:
        id:
          type: string
        object:
          type: string
          examples:
            - response
        created_at:
          type: integer
        completed_at:
          type:
            - integer
            - 'null'
        status:
          type: string
          enum:
            - completed
            - in_progress
            - incomplete
            - failed
            - cancelled
            - queued
        error:
          type:
            - object
            - 'null'
          additionalProperties: true
        model:
          type: string
        output:
          type: array
          items:
            type: object
            properties:
              id:
                type: string
              type:
                type: string
              status:
                type: string
              role:
                type: string
              content:
                type: array
                items:
                  type: object
                  additionalProperties: true
        usage:
          $ref: '#/components/schemas/ResponseUsage'
      required:
        - id
        - object
        - model
        - output
    ResponseUsage:
      type: object
      properties:
        input_tokens:
          type: integer
        output_tokens:
          type: integer
        total_tokens:
          type: integer
        input_tokens_details:
          type: object
          properties:
            cached_tokens:
              type: integer
            text_tokens:
              type: integer
            image_tokens:
              type: integer
        output_tokens_details:
          type: object
          properties:
            reasoning_tokens:
              type: integer
            text_tokens:
              type: integer
        cost:
          type: object
          properties:
            input_cost:
              type: number
            output_cost:
              type: number
            total_cost:
              type: number
              description: Списано с кошелька, в рублях.
      required:
        - total_tokens
    OpenAIError:
      type: object
      properties:
        error:
          type: object
          properties:
            message:
              type: string
            type:
              type: string
              examples:
                - invalid_request_error
              description: >-
                `authentication_error` при 401, `rate_limit_error` при 429,
                `permission_error` при 403, в остальных случаях
                `invalid_request_error`, `api_error` или `internal_error`.
            param:
              type:
                - string
                - 'null'
            code:
              type:
                - string
                - 'null'
              examples:
                - invalid_request
              description: >-
                Машиночитаемый код: `invalid_api_key`, `unauthorized`,
                `insufficient_quota`, `rate_limit_exceeded`,
                `token_limit_exceeded`, `content_too_large`, `limit_exceeded`,
                `model_not_found`, `no_providers`, `service_unavailable`,
                `provider_error`, `invalid_request`, `context_length_exceeded`,
                `internal_error`.
          required:
            - message
            - type
      required:
        - error
  responses:
    OpenAIError:
      description: Ошибка в формате OpenAI.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/OpenAIError'
  securitySchemes:
    BearerAuth:
      type: http
      scheme: bearer
      description: >-
        API-ключ в формате `sk-...`. Альтернативно — заголовок `X-Api-Key:
        sk-...` (совместимо с Anthropic SDK).

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.