Playbooks
Open models on your own GPUs: get one, serve it at one address, work with it in notebooks, fine-tune it, and serve the result the same way.
Ask your agent: "get Qwen/Qwen3-8B and load it on the spark". The model is mirrored once into your storage and served at llm.<your domain> through the OpenAI or Anthropic API. Write your app or your notebook against that address. When the prompt is not enough, fine-tune the model in a notebook, register the result, and load it at the same address.
How it works
-
A model is a file of weights. Its size is counted in parameters. Its context window, counted in tokens, is how much text it can read at one time. Quantization stores the weights in fewer bits: BF16 is the full model, FP8, FP4 and Q4 are smaller copies that fit a smaller GPU. At 16 bits a model needs about two gigabytes of GPU memory per billion parameters, plus room for the context; at 4 bits, a quarter of that. The catalogue carries each model’s size, quantization and context, and the platform says what fits your GPU before you download anything.
-
Mirrored once, served at one address. The weights are downloaded once from Hugging Face into Thinkube Storage. A load starts a serving engine on the node you choose, with the context length you choose, and registers the model with the LLM Gateway. The engines are optional components: vLLM, TensorRT-LLM, Ollama for quantized GGUF models, and text-embeddings for embedding models. A vLLM, TensorRT-LLM or text-embeddings engine serves one model per node; Ollama’s holds several.
-
Notebooks on the same filesystem as the IDE. Thinkube Notebooks is JupyterLab at
notebooks.<your domain>. Two environments come with the platform:fine-tuningwith the training libraries, andagent-devwith LangChain, AG2, CrewAI and LangGraph. Both are built per architecture. Both includetk_llm, a Python helper already set up for your gateway. An environment of your own starts from one of them. The agent opens, runs and reads notebooks through Thinkube Control, in your server or in a server of its own that stops when the run ends. -
The example notebooks. The first server start copies them into
notebooks/examples/:research-assistant/with four notebooks that check the platform, load the models, index papers and run a debate, andzebra-grpo/with the fine-tuning notebook. Each start also puts a fresh copy intemplates/examples/, for you to take when you want it. -
Improving a model. Fine-tuning trains an existing model a little further on your examples. A LoRA adapter (a small add-on file trained beside the model) is a few hundred megabytes and trains on one GPU in hours. When a program can check an answer, the program replaces the examples: the model tries, the checker scores, and the model learns from the score. Distillation lets a large model teach a small one. Thinkube Experiments records every run and the model it produced, and a registered model is served exactly like a mirrored one.
-
Building with models. An embedding model turns text into vectors, and a vector store finds the nearest ones: search by meaning. Retrieval puts the documents found into the prompt, so the model answers from your documents and cites them. Tool calling lets a model ask for a function to be run. An agent repeats that in a loop, and the catalogue marks the models that can do it. Langfuse traces every call, with its prompts, its documents, its tools and its latency, and those traces are the record you fine-tune from next.
Playbooks
Serve your first model
beginnerLoad an open model on your own GPU and call it from a notebook, as you would call a cloud API
2026-10-04 Thinkube Models 30 minLoad a model on the node and context you choose
intermediatePick the GPU node, the backend and the context length for a load, and what happens when a node is busy
2026-10-04 Thinkube Models 30 minCall your models from an app or a script
beginnerPoint the OpenAI or Anthropic client you already use at your gateway, with a Thinkube API token
2026-10-04 Thinkube Models 30 minBuild a notebook environment
beginnerStart from agent-dev or fine-tuning, add the packages you need, and get a kernel built for every architecture in your cluster
2026-10-04 Thinkube Models 30 minRun notebooks from your agent, watched or unattended
beginnerHave your agent run a notebook while you watch it in a Thinkube IDE tab, or on a server of its own that stops when the run ends
2026-10-04 Thinkube Models 30 minBuild a research assistant over your papers
intermediateIndex arXiv papers in your own vector store, ask questions that come back with sources, and let agents argue from the index
2026-10-04 Thinkube Models 9 h 30 minFine-tune a model on rewards a program checks
advancedTrain a 4B model with GRPO against a solver that grades every answer, with no labelled data, and register the result beside its run
2026-10-04 Thinkube Models 30 minServe a model you fine-tuned
intermediateGive a registered fine-tune a catalogue entry, load it on a GPU and call it through the gateway like any other model
2026-10-04Reference
-
LLM serving and models: model states, load options, sizing, routing and authentication at the gateway.
-
Notebook execution over MCP: every notebook operation and what it answers.
-
Components catalog: the inference backends, vector stores and tracing.