Thinkube Models
Serve your first model
Load an open model on your own GPU and call it from a notebook, as you would call a cloud API
- Level
- beginner
- Time
- 30 min
- Risk
- low
- Updated
- 2026-10-04
Overview
Basic idea
A cloud model is an address and a key: you send a request and an answer comes back. Serving your own model gives you the same address and the same API, with the weights on your storage and the work on your GPU. Nothing you send leaves your machines, and you choose the model.
On Thinkube a model goes through three states:
-
Mirrored. The weights are downloaded once from Hugging Face into your own storage. Every later load reads them from there.
-
Loaded. A serving engine, vLLM for this model, holds the weights in a GPU’s memory and answers requests.
-
Called. Every loaded model answers at one address, the LLM Gateway at
llm.<your domain>, through the OpenAI API, so the same code calls any model you load.
What you’ll accomplish
You load Qwen/Qwen3.5-4B on a GPU node, call it from a notebook, and unload it to give the GPU back.
What to know before starting
Required
-
Asking your agent in Thinkube IDE to do things on the platform.
Optional
-
Thinkube Models: parameters, quantization and context, the numbers that decide what fits.
-
LLM serving and models: how the gateway routes, authenticates and sizes a load.
Supported hardware
-
GPU: one NVIDIA GPU with about 14 GB of memory free: what the model needs with its default context.
-
Architecture: amd64 or arm64.
Prerequisites
Platform
-
Thinkube running, with Thinkube IDE open. Ask your agent: "what’s running?"
-
A notebook server. Ask your agent: "start a notebook server on tkamd1". The first start on a node downloads the notebook environments, so it takes longer than later starts. Replace tkamd1 and tkamd2, here and below, with your own nodes.
Components
-
The optional component vLLM. Ask your agent: "is vLLM installed?" If it is not, ask "install vLLM".
Instructions
Step 1. Check the model fits
Before anything is downloaded, the platform knows how much memory a model needs and how much each GPU node has free.
Ask your agent:
› how much GPU memory does Qwen/Qwen3.5-4B need, and which node has room?
Expected output, from the reference run:
Qwen/Qwen3.5-4B: about 13.6 GB with its default context tkamd2: 45.5 GB free across 2 × RTX 3090, 3 of 4 slots free tkspark: 26.1 GB left of its 96 GB budget, 3 of 4 slots free
Step 2. Mirror it
A model already mirrored shows is_downloaded: true in the catalogue; skip to the next step.
The download runs as a job in the background, from Hugging Face into your storage.
Ask your agent:
› mirror Qwen/Qwen3.5-4B
Or by hand: in Thinkube Control open AI Models and choose Mirror on the model.
Expected output, from the reference run:
model_id: Qwen/Qwen3.5-4B status: succeeded created_at: 22:03:22 updated_at: 22:11:02
Step 3. Load it
Loading starts a vLLM engine on the node you name, puts the weights in its GPU memory and registers the model with the gateway.
Ask your agent:
› load Qwen/Qwen3.5-4B on tkamd2
Or by hand: in Thinkube Control open LLM Gateway, choose Load on the model, pick the node in the dialog and choose Load Model.
Expected output:
state: loading message: Loading on tkamd2 with a 32768 token context; this takes some minutes backend_id: vllm-tkamd2
and then, once the engine is ready:
state: available backend_id: vllm-tkamd2
Step 4. Call it from a notebook
tk_llm is installed in both notebook environments and reads the gateway address and your token from the notebook server, so the call needs no configuration.
In Thinkube Notebooks, create a notebook with the fine-tuning kernel and run:
from tk_llm import get_openai_client
client = get_openai_client()
response = client.chat.completions.create(
model="Qwen/Qwen3.5-4B",
messages=[{"role": "user", "content": "Hello"}],
max_tokens=1024,
)
print(response.choices[0].message.content)
print(response.usage)
Or ask your agent: "create first-model.ipynb on tkamd1 with a fine-tuning kernel, call Qwen/Qwen3.5-4B with Hello, and run it".
Expected output, from the reference run:
Hello! 👋 How's your day going? Is there anything I can help you with today? Whether it's solving a problem, answering questions, or just chatting, I'm here to assist! 😊 CompletionUsage(completion_tokens=160, prompt_tokens=11, total_tokens=171, ...)
Qwen3.5 thinks before it answers, and the thinking counts against max_tokens: 160 completion tokens for a two-line answer. Give it room, as here.
Step 5. Unload it
A loaded model holds its GPU memory until you unload it.
Ask your agent:
› unload Qwen/Qwen3.5-4B
Or by hand: in LLM Gateway, choose Unload on the model.
Expected output:
state: deployable message: Model unloaded successfully
The vLLM engine on tkamd2 stopped at once, and the node’s free memory returned to 45.5 GB.
Step 6. Next steps
To undo: Unload the model. The weights stay in your storage for the next load.
-
Load a model on the node and context you choose: choose the node, the backend and the context length for each load.
-
Build a notebook environment: the environments and how to build your own.
-
Build a research assistant over your papers: a served model answering over your own documents.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
The call answers with empty content |
The model spent |
Raise |
The call fails with |
The notebook server started without |
Stop and start the notebook server; it reads the gateway address when it starts. |