1. Resources
  2. API Reference

​
Authentication

All API requests require authentication using your API key in the request headers.

​
Header Format

Authorization
required
string

Bearer token for API authentication. Replace YOUR_API_KEY_HERE with your actual API key.

Content-Type
required
string

Content type of the request. Must be set to application/json.

​
API Reference

​
Base URL

https://api.corethink.ai/v1/code

​
HTTP Method

POST

​
Request Format

​
Parameters

messages
required
array

Array of message objects representing the conversation history. Each message must have a role and content.

model
optional
string

Specifies which model to use. If not specified, defaults to the organization’s configured model.

Available models: gpt-oss-120b deepseek/deepseek-r1 deepseek/deepseek-v3 qwen/qwen3-235b qwen/qwen2.5-max mistral/mistral-small-3-24b

tools
optional
array

Array of tool/function definitions available for the model to use. Each tool follows the OpenAI function calling format.

temperature
optional
default:0.7
number

Number between 0 and 2. Higher values make output more random, lower values more deterministic.

max_tokens
optional
number

Maximum number of tokens to generate. Default varies by model.

stream
optional
default:false
boolean

Enable streaming responses. When true, returns Server-Sent Events (SSE) for real-time token generation.

extra_body
optional
object

Object containing additional provider-specific configuration. This is useful for specifying provider preferences or passing through additional options to downstream services. Common extra_body Options:

Common extra_body Options: provider.order (Array): Preferred provider order for routing rovider.allow_fallbacks (Boolean): Allow fallback to other providers if preferred is unavailable provider.require (Array): Required provider capabilities

​
Token Reduction Parameters

token_limit
optional
number

Token limit for compression (triggers compression when exceeded)

mmr_lambda_nl
optional
default:0.6
number

MMR lambda for natural language compression (0-1)

mmr_lambda_code
optional
default:0.6
number

MMR lambda for code compression (0-1)

embedding_provider
optional
default:gemini
string

Embedding provider for compression

​
Message Object Structure

Each message in the messages array should contain:

role
required
string

The role of the message author. Must be one of: “system”, “user”, “assistant”, “tool”, or “function”

content
optional
string

The content of the message. Can be null for tool/function messages.

Tool Object Structure:

{
  "type": "function",
  "function": {
    "name": "function_name",
    "description": "Description of what the function does",
    "parameters": {
      "type": "object",
      "properties": {
        "param_name": {
          "type": "string",
          "description": "Parameter description"
        }
      },
      "required": ["param_name"]
    }
  }
}

​
Example Request

curl https://api.corethink.ai/v1/code \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "model": "gpt-oss-120b",
    "messages": [
      {
        "role": "user",
        "content": "What is the weather in San Francisco?"
      }
    ],
    "tools": [
      {
        "type": "function",
        "function": {
          "name": "get_weather",
          "description": "Get current weather for a location",
          "parameters": {
            "type": "object",
            "properties": {
              "location": {
                "type": "string",
                "description": "City name"
              }
            },
            "required": ["location"]
          }
        }
      }
    ]
  }'

​
Streaming Example

Enable real-time token generation with streaming:

from openai import OpenAI
import os

client = OpenAI(
    base_url="https://api.corethink.ai/v1/code",
    api_key=os.environ.get("CORETHINK_API_KEY"),
)

# Enable streaming
stream = client.chat.completions.create(
    model="gpt-oss-120b",
    messages=[{"role": "user", "content": "Write a Python function to sort a list"}],
    stream=True
)

# Process streaming response
for chunk in stream:
    if chunk.choices[0].delta.content is not None:
        print(chunk.choices[0].delta.content, end='', flush=True)

print()  # New line after streaming completes

Streaming Response Format:

Each chunk follows this structure:

{
  "id": "chatcmpl-abc123",
  "object": "chat.completion.chunk",
  "created": 1677652288,
  "model": "gpt-oss-120b",
  "choices": [
    {
      "index": 0,
      "delta": {
        "content": "def "
      },
      "finish_reason": null
    }
  ]
}

The final chunk includes "finish_reason": "stop".

​
Additional OpenAI-Compatible Parameters

The API also supports these optional OpenAI-compatible parameters:

tool_choice
optional
string

Controls how the model uses tools. Can be “none”, “auto”, or a specific tool.

top_p
optional
number

Alternative to temperature for nucleus sampling. Considered together with temperature.

n
optional
number

How many chat completion choices to generate for each input message.

stop
optional
string

Up to 4 sequences where the API will stop generating further tokens.

presence_penalty
optional
number

Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far.

frequency_penalty
optional
number

Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far.