Thinkube Models

Call your models from an app or a script

Point the OpenAI or Anthropic client you already use at your gateway, with a Thinkube API token

Level
beginner
Time
30 min
Risk
low
Updated
2026-10-04

LLM GatewayOpenAI APIAnthropic APIAPI tokens

Overview

Basic idea

Every model you load answers at one address, https://llm.<your domain>. The gateway speaks two APIs, so the client libraries you use for cloud models work unchanged:

  • OpenAI API. /v1/chat/completions, /v1/embeddings and /v1/models, for the openai client and everything built on it.

  • Anthropic Messages API. /v1/messages, for the anthropic client.

  • One token. A Thinkube Control API token, which starts with tk_, is accepted as Authorization: Bearer and as x-api-key, so each client sends it the way it always does.

Changing the model is changing one string. The code stays the same.

What you’ll accomplish

You create an API token, call a loaded model with the openai client and with the anthropic client, stream an answer, and unload the model.

What to know before starting

Required

Optional

Supported hardware

  • GPU: not needed where the calls run: a notebook, Thinkube IDE, an app on the cluster, or a computer on your tailnet that reaches llm.<your domain>. The model is served behind the LLM Gateway on a GPU node.

  • Architecture: amd64 or arm64.

Prerequisites

Platform

  • Thinkube running, with Thinkube IDE open. Ask your agent: "what’s running?"

  • Qwen/Qwen3.5-4B mirrored. Ask your agent: "is Qwen/Qwen3.5-4B mirrored?" If it is not, ask "mirror Qwen/Qwen3.5-4B".

  • A notebook server, to run the calls. Ask your agent: "start a notebook server on tkamd1". Replace tkamd1 and tkamd2, here and below, with your own nodes.

Components

  • The optional component vLLM. Ask your agent: "is vLLM installed?" If it is not, ask "install vLLM".

Instructions

Step 1. Load a model

Ask your agent:

› load Qwen/Qwen3.5-4B on tkamd2 with a 16k context

Expected output:

state: loading
message: Loading on tkamd2 with a 16384 token context; this takes some minutes
backend_id: vllm-tkamd2

and then state: available. Load a model on the node and context you choose covers the choices.

Step 2. Get a token

A script or an app outside the platform’s own tools authenticates with an API token.

In Thinkube Control open API Tokens. Under Create New Token type a Token Name, leave Expires in (days) empty for a token with no expiry or give a number of days, and choose Create Token. Copy the token from Token Created Successfully.

Keep the address and the token in two variables:

export LLM_GATEWAY_URL=https://llm.<your domain>
export THINKUBE_API_TOKEN=tk_...

In Thinkube Notebooks both are set for you. Both notebook environments have the openai and anthropic clients installed.

Step 3. Call it with the OpenAI client

The OpenAI client takes the gateway address with /v1.

import os
from openai import OpenAI

client = OpenAI(
    base_url=os.environ["LLM_GATEWAY_URL"] + "/v1",
    api_key=os.environ["THINKUBE_API_TOKEN"],
)

print([m.id for m in client.models.list().data])

response = client.chat.completions.create(
    model="Qwen/Qwen3.5-4B",
    messages=[{"role": "user", "content": "Name the capital of France in one word."}],
    max_tokens=1024,
    extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
print(response.choices[0].message.content)
print(response.usage)

Or ask your agent: "call Qwen/Qwen3.5-4B through the gateway with the OpenAI client from a notebook on tkamd1".

Expected output, from the reference run:

['Qwen/Qwen3-8B', 'Qwen/Qwen3.5-4B', 'unsloth/Qwen3.6-27B-NVFP4', 'Qwen/Qwen3-Embedding-0.6B']
Paris
CompletionUsage(completion_tokens=2, prompt_tokens=21, total_tokens=23, ...)

/v1/models lists loaded and mirrored models; only loaded ones answer. extra_body switches Qwen3.5’s thinking off, so the answer is the whole completion.

Step 4. Stream the answer

An app shows the answer as it is written by asking for a stream.

stream = client.chat.completions.create(
    model="Qwen/Qwen3.5-4B",
    messages=[{"role": "user", "content": "Count from 1 to 5, separated by commas."}],
    max_tokens=1024,
    stream=True,
    extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Expected output, from the reference run, in 13 chunks:

1, 2, 3, 4, 5

Step 5. Call it with the Anthropic client

The Anthropic client takes the gateway address without /v1, and sends the token as x-api-key.

import os
from anthropic import Anthropic

client = Anthropic(
    base_url=os.environ["LLM_GATEWAY_URL"],
    api_key=os.environ["THINKUBE_API_TOKEN"],
)

message = client.messages.create(
    model="Qwen/Qwen3.5-4B",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Name the capital of France in one word."}],
)
for block in message.content:
    print(block.type)
print("".join(b.text for b in message.content if b.type == "text").strip())
print(message.usage.input_tokens, message.usage.output_tokens)

Expected output, from the reference run:

thinking
text
Paris
19 105

The model’s thinking arrives as its own thinking block, before the text block with the answer, as it does from Claude.

Step 6. Unload the model

Ask your agent:

› unload Qwen/Qwen3.5-4B

Or by hand: in LLM Gateway, choose Unload on the model.

Expected output:

state: deployable
message: Model unloaded successfully

Step 7. Next steps

To undo: Unload the model, and delete the token under API Tokens when you no longer use it.

Troubleshooting

Symptom Cause Fix

AuthenticationError: Error code: 401 - {'error': {'message': 'Invalid authentication token', 'type': 'authentication_error', 'code': 'invalid_token'}}

The token is not a valid Thinkube Control API token or Thinkube Identity token

Copy the token again from API Tokens, or create a new one.

NotFoundError: Error code: 404 - {'error': {'message': "Model 'Qwen/Qwen9-1B' not found", …​}}

The name is not a model in the catalogue

Use a name from client.models.list().

InternalServerError: Error code: 502 - {'error': {'message': 'Backend request failed', 'type': 'backend_error', 'code': 'backend_error'}}

The model is mirrored and listed, and not loaded

Load it, and call it once its state is available.