Thinkube Models

Serve your first model

Load an open model on your own GPU and call it from a notebook, as you would call a cloud API

Level
beginner
Time
30 min
Risk
low
Updated
2026-10-04

LLM GatewayvLLMNotebooks

Overview

Basic idea

A cloud model is an address and a key: you send a request and an answer comes back. Serving your own model gives you the same address and the same API, with the weights on your storage and the work on your GPU. Nothing you send leaves your machines, and you choose the model.

On Thinkube a model goes through three states:

  • Mirrored. The weights are downloaded once from Hugging Face into your own storage. Every later load reads them from there.

  • Loaded. A serving engine, vLLM for this model, holds the weights in a GPU’s memory and answers requests.

  • Called. Every loaded model answers at one address, the LLM Gateway at llm.<your domain>, through the OpenAI API, so the same code calls any model you load.

What you’ll accomplish

You load Qwen/Qwen3.5-4B on a GPU node, call it from a notebook, and unload it to give the GPU back.

What to know before starting

Required

  • Asking your agent in Thinkube IDE to do things on the platform.

Optional

Supported hardware

  • GPU: one NVIDIA GPU with about 14 GB of memory free: what the model needs with its default context.

  • Architecture: amd64 or arm64.

Prerequisites

Platform

  • Thinkube running, with Thinkube IDE open. Ask your agent: "what’s running?"

  • A notebook server. Ask your agent: "start a notebook server on tkamd1". The first start on a node downloads the notebook environments, so it takes longer than later starts. Replace tkamd1 and tkamd2, here and below, with your own nodes.

Components

  • The optional component vLLM. Ask your agent: "is vLLM installed?" If it is not, ask "install vLLM".

Instructions

Step 1. Check the model fits

Before anything is downloaded, the platform knows how much memory a model needs and how much each GPU node has free.

Ask your agent:

› how much GPU memory does Qwen/Qwen3.5-4B need, and which node has room?

Expected output, from the reference run:

Qwen/Qwen3.5-4B: about 13.6 GB with its default context
tkamd2: 45.5 GB free across 2 × RTX 3090, 3 of 4 slots free
tkspark: 26.1 GB left of its 96 GB budget, 3 of 4 slots free

Step 2. Mirror it

A model already mirrored shows is_downloaded: true in the catalogue; skip to the next step.

The download runs as a job in the background, from Hugging Face into your storage.

Ask your agent:

› mirror Qwen/Qwen3.5-4B

Or by hand: in Thinkube Control open AI Models and choose Mirror on the model.

Expected output, from the reference run:

model_id: Qwen/Qwen3.5-4B
status: succeeded
created_at: 22:03:22   updated_at: 22:11:02

Step 3. Load it

Loading starts a vLLM engine on the node you name, puts the weights in its GPU memory and registers the model with the gateway.

Ask your agent:

› load Qwen/Qwen3.5-4B on tkamd2

Or by hand: in Thinkube Control open LLM Gateway, choose Load on the model, pick the node in the dialog and choose Load Model.

Expected output:

state: loading
message: Loading on tkamd2 with a 32768 token context; this takes some minutes
backend_id: vllm-tkamd2

and then, once the engine is ready:

state: available
backend_id: vllm-tkamd2

Step 4. Call it from a notebook

tk_llm is installed in both notebook environments and reads the gateway address and your token from the notebook server, so the call needs no configuration.

In Thinkube Notebooks, create a notebook with the fine-tuning kernel and run:

from tk_llm import get_openai_client

client = get_openai_client()
response = client.chat.completions.create(
    model="Qwen/Qwen3.5-4B",
    messages=[{"role": "user", "content": "Hello"}],
    max_tokens=1024,
)
print(response.choices[0].message.content)
print(response.usage)

Or ask your agent: "create first-model.ipynb on tkamd1 with a fine-tuning kernel, call Qwen/Qwen3.5-4B with Hello, and run it".

Expected output, from the reference run:

Hello! 👋 How's your day going? Is there anything I can help you with today? Whether it's solving a problem, answering questions, or just chatting, I'm here to assist! 😊
CompletionUsage(completion_tokens=160, prompt_tokens=11, total_tokens=171, ...)

Qwen3.5 thinks before it answers, and the thinking counts against max_tokens: 160 completion tokens for a two-line answer. Give it room, as here.

Step 5. Unload it

A loaded model holds its GPU memory until you unload it.

Ask your agent:

› unload Qwen/Qwen3.5-4B

Or by hand: in LLM Gateway, choose Unload on the model.

Expected output:

state: deployable
message: Model unloaded successfully

The vLLM engine on tkamd2 stopped at once, and the node’s free memory returned to 45.5 GB.

Step 6. Next steps

To undo: Unload the model. The weights stay in your storage for the next load.

Troubleshooting

Symptom Cause Fix

The call answers with empty content

The model spent max_tokens on thinking before it wrote the answer; a 20-token budget left nothing

Raise max_tokens, for example to 1024, or pass extra_body={"chat_template_kwargs": {"enable_thinking": False}}.

The call fails with APIConnectionError: Connection error. and [Errno -2] Name or service not known

The notebook server started without LLM_GATEWAY_URL, so tk_llm could not find the gateway

Stop and start the notebook server; it reads the gateway address when it starts.