Thinkube Models
Load a model on the node and context you choose
Pick the GPU node, the backend and the context length for a load, and what happens when a node is busy
- Level
- intermediate
- Time
- 30 min
- Risk
- low
- Updated
- 2026-10-04
Overview
Basic idea
A load has a few choices. You can set each choice, or let the platform choose. The choices decide where the model runs and how much memory it takes.
-
Node. The GPU node that serves the model. Without it, the platform picks the node with a free slot and the most AI memory, the GPU memory left for models.
-
Backend. A type (
vllm,tensorrt-llm,text-embeddings,ollama) or the id of a running backend, such asvllm-tkspark, which also names its node. -
Tier.
performanceloads through vLLM or TensorRT-LLM.flexibleloads through Ollama, for models the catalogue lists with an Ollama entry. -
Context length. The longest prompt plus answer the model accepts. Memory for the KV cache, the working memory for the conversation, grows with it, so a smaller context lets a larger model fit a GPU. Without it, the load takes the largest of 32k, 16k and 8k tokens that the model allows and the node has memory for.
-
One model per engine. A vLLM, TensorRT-LLM or text-embeddings engine serves one model on its node. Ollama’s engine holds several.
What you’ll accomplish
You load Qwen/Qwen3-8B on tkamd2 with a 16k context, see a second model on the same node refused with the name of the one it serves, and unload the model.
What to know before starting
Required
-
Serve your first model: the mirror, load, call and unload cycle.
Optional
-
LLM serving and models: how a load is sized, and more load options.
-
Thinkube Models: parameters, quantization and context.
Supported hardware
-
GPU: one NVIDIA GPU with about 22 GB of memory free for
Qwen/Qwen3-8Bwith a 32k context: its weights are 15.3 GiB in BF16, and the load options estimate 21.5 GB. A shorter context needs less. -
Architecture: amd64 or arm64.
Prerequisites
Platform
-
Thinkube running, with Thinkube IDE open. Ask your agent: "what’s running?"
-
Qwen/Qwen3-8Bmirrored. Ask your agent: "is Qwen/Qwen3-8B mirrored?" If it is not, ask "mirror Qwen/Qwen3-8B".
Components
-
The optional component vLLM. Ask your agent: "is vLLM installed?" If it is not, ask "install vLLM".
Instructions
Step 1. Read the load options
Before a load, the platform tells you how much memory the model needs and what each GPU node has left.
Ask your agent:
› what are the load options for Qwen/Qwen3-8B?
Expected output, from the reference run:
Qwen/Qwen3-8B: about 21.5 GB with a 32k context; the model allows 131072 tkamd2: 45.5 GB left of 48 GB, 3 of 4 slots free, text-embeddings serving Qwen/Qwen3-Embedding-0.6B tkspark: 26.1 GB left of its 96 GB budget, 3 of 4 slots free, vllm-tkspark serving unsloth/Qwen3.6-27B-NVFP4
tkamd2 has room and no vLLM engine, so the load creates one there. tkspark’s vLLM engine already serves a model.
Step 2. Load it on the node and context you choose
Name the node and the context. vLLM is the only backend for this model.
Ask your agent:
› load Qwen/Qwen3-8B on tkamd2 with vllm and a 16k context
Replace tkamd2 with one of your nodes.
Or by hand: in Thinkube Control open LLM Gateway and choose Load on the model. In the dialog pick tkamd2 under Target GPU Node and 16K under Context Length, then choose Load Model. The dialog shows Backend Type for models the catalogue lists with more than one.
Expected output:
state: loading message: Loading on tkamd2 with a 16384 token context; this takes some minutes backend_id: vllm-tkamd2
and then:
state: available backend_id: vllm-tkamd2
The engine’s log on the reference run:
Maximum concurrency for 16,384 tokens per request: 2.60x
The GPU has room for about 2.6 full-length requests at once.
The platform checks its backends every minute. So the model shows available up to a minute after the engine is ready.
Step 3. Ask for a second model on the same node
The vLLM engine on tkamd2 now serves Qwen/Qwen3-8B. A load of another vLLM model on tkamd2 is refused, and the message names the model the engine serves. The running model keeps running.
Ask your agent:
› load Qwen/Qwen3.5-4B on tkamd2
Expected output, from the reference run:
state: deployable message: vllm-tkamd2 already serves Qwen/Qwen3-8B (available); unload it, or choose another node
The same load with backend: vllm-tkamd2 gives the same answer. While the first model is still loading, the message says (loading).
To serve both, load the second model on another node with room, or unload the first.
Step 4. Unload it
The model holds most of a 24 GB GPU until you unload it.
Ask your agent:
› unload Qwen/Qwen3-8B
Or by hand: in LLM Gateway, choose Unload on the model.
Expected output:
state: deployable message: Model unloaded successfully
On the reference run the vLLM engine on tkamd2 was removed at once, since it served no other model.
Step 5. Next steps
To undo: Unload the model. The weights stay in your storage for the next load.
-
Serve your first model: call a loaded model from a notebook.
-
LLM serving and models: sizing, speculative decoding and routing.
-
Build a research assistant over your papers: a served model answering over your own documents.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
|
The vLLM engine on that node serves one model, and it serves another |
Load on another node with room, or unload the model it names. |
|
A backend id names an engine that runs now; |
Name the type and the node instead: |
|
|
Leave the tier out, or ask for |