Thinkube Models
Call your models from an app or a script
Point the OpenAI or Anthropic client you already use at your gateway, with a Thinkube API token
- Level
- beginner
- Time
- 30 min
- Risk
- low
- Updated
- 2026-10-04
Overview
Basic idea
Every model you load answers at one address, https://llm.<your domain>. The gateway speaks two APIs, so the client libraries you use for cloud models work unchanged:
-
OpenAI API.
/v1/chat/completions,/v1/embeddingsand/v1/models, for theopenaiclient and everything built on it. -
Anthropic Messages API.
/v1/messages, for theanthropicclient. -
One token. A Thinkube Control API token, which starts with
tk_, is accepted asAuthorization: Bearerand asx-api-key, so each client sends it the way it always does.
Changing the model is changing one string. The code stays the same.
What you’ll accomplish
You create an API token, call a loaded model with the openai client and with the anthropic client, stream an answer, and unload the model.
What to know before starting
Required
-
Serve your first model: load and unload a model.
-
Python, and one of the two client libraries.
Optional
-
LLM serving and models: routes, authentication and routing.
-
How configuration reaches your app: keep the token as a secret in a deployed app.
Supported hardware
-
GPU: not needed where the calls run: a notebook, Thinkube IDE, an app on the cluster, or a computer on your tailnet that reaches
llm.<your domain>. The model is served behind the LLM Gateway on a GPU node. -
Architecture: amd64 or arm64.
Prerequisites
Platform
-
Thinkube running, with Thinkube IDE open. Ask your agent: "what’s running?"
-
Qwen/Qwen3.5-4Bmirrored. Ask your agent: "is Qwen/Qwen3.5-4B mirrored?" If it is not, ask "mirror Qwen/Qwen3.5-4B". -
A notebook server, to run the calls. Ask your agent: "start a notebook server on tkamd1". Replace tkamd1 and tkamd2, here and below, with your own nodes.
Components
-
The optional component vLLM. Ask your agent: "is vLLM installed?" If it is not, ask "install vLLM".
Instructions
Step 1. Load a model
Ask your agent:
› load Qwen/Qwen3.5-4B on tkamd2 with a 16k context
Expected output:
state: loading message: Loading on tkamd2 with a 16384 token context; this takes some minutes backend_id: vllm-tkamd2
and then state: available. Load a model on the node and context you choose covers the choices.
Step 2. Get a token
A script or an app outside the platform’s own tools authenticates with an API token.
In Thinkube Control open API Tokens. Under Create New Token type a Token Name, leave Expires in (days) empty for a token with no expiry or give a number of days, and choose Create Token. Copy the token from Token Created Successfully.
Keep the address and the token in two variables:
export LLM_GATEWAY_URL=https://llm.<your domain>
export THINKUBE_API_TOKEN=tk_...
In Thinkube Notebooks both are set for you. Both notebook environments have the openai and anthropic clients installed.
Step 3. Call it with the OpenAI client
The OpenAI client takes the gateway address with /v1.
import os
from openai import OpenAI
client = OpenAI(
base_url=os.environ["LLM_GATEWAY_URL"] + "/v1",
api_key=os.environ["THINKUBE_API_TOKEN"],
)
print([m.id for m in client.models.list().data])
response = client.chat.completions.create(
model="Qwen/Qwen3.5-4B",
messages=[{"role": "user", "content": "Name the capital of France in one word."}],
max_tokens=1024,
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
print(response.choices[0].message.content)
print(response.usage)
Or ask your agent: "call Qwen/Qwen3.5-4B through the gateway with the OpenAI client from a notebook on tkamd1".
Expected output, from the reference run:
['Qwen/Qwen3-8B', 'Qwen/Qwen3.5-4B', 'unsloth/Qwen3.6-27B-NVFP4', 'Qwen/Qwen3-Embedding-0.6B'] Paris CompletionUsage(completion_tokens=2, prompt_tokens=21, total_tokens=23, ...)
/v1/models lists loaded and mirrored models; only loaded ones answer. extra_body switches Qwen3.5’s thinking off, so the answer is the whole completion.
Step 4. Stream the answer
An app shows the answer as it is written by asking for a stream.
stream = client.chat.completions.create(
model="Qwen/Qwen3.5-4B",
messages=[{"role": "user", "content": "Count from 1 to 5, separated by commas."}],
max_tokens=1024,
stream=True,
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Expected output, from the reference run, in 13 chunks:
1, 2, 3, 4, 5
Step 5. Call it with the Anthropic client
The Anthropic client takes the gateway address without /v1, and sends the token as x-api-key.
import os
from anthropic import Anthropic
client = Anthropic(
base_url=os.environ["LLM_GATEWAY_URL"],
api_key=os.environ["THINKUBE_API_TOKEN"],
)
message = client.messages.create(
model="Qwen/Qwen3.5-4B",
max_tokens=1024,
messages=[{"role": "user", "content": "Name the capital of France in one word."}],
)
for block in message.content:
print(block.type)
print("".join(b.text for b in message.content if b.type == "text").strip())
print(message.usage.input_tokens, message.usage.output_tokens)
Expected output, from the reference run:
thinking text Paris 19 105
The model’s thinking arrives as its own thinking block, before the text block with the answer, as it does from Claude.
Step 6. Unload the model
Ask your agent:
› unload Qwen/Qwen3.5-4B
Or by hand: in LLM Gateway, choose Unload on the model.
Expected output:
state: deployable message: Model unloaded successfully
Step 7. Next steps
To undo: Unload the model, and delete the token under API Tokens when you no longer use it.
-
Load a model on the node and context you choose: serve the model your app needs where it fits.
-
How configuration reaches your app: give a deployed app the token as a secret.
-
Build a research assistant over your papers: a served model answering over your own documents.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
|
The token is not a valid Thinkube Control API token or Thinkube Identity token |
Copy the token again from API Tokens, or create a new one. |
|
The name is not a model in the catalogue |
Use a name from |
|
The model is mirrored and listed, and not loaded |
Load it, and call it once its state is |