Getting started
What Thinkube is
Thinkube is a platform for one developer, on machines you own, that takes an AI project from an idea to a running product: a model served on your GPU, an app that calls it, notebooks to fine-tune it, and a report on every change. You operate it by talking to a coding agent.
Open Thinkube IDE in your browser, start a coding agent, and say what you want: "get Qwen3.5-4B and serve it", "make an app called notes from the web app template", "install Qdrant". Each sentence is a real operation of the platform, and each one is also a button in Thinkube Control. The playbooks are those sentences, one task each, with what you get.
Four areas on one base
| Area | What it gives you |
|---|---|
The base: unmodified upstream Kubernetes, installed with kubeadm on your machines, with identity, storage, registry, database, certificates and GPUs installed and connected, driven from a coding agent. |
|
Push to git, and it runs: applications from templates, built and deployed at an address with a certificate and a login. |
|
Serve, fine-tune and track open models on your own GPUs, from notebooks to one address every app calls. |
|
You and a coding agent working together: the agent does the work, you decide. Thinkube Tandem Chat for a conversation, Thinkube Tandem Workshop for a change built step by step. |
The areas are one platform. A dataset a notebook writes is the file an app reads. The model a notebook fine-tunes is served at the address the app calls. The change the Tandem Workshop delivers is deployed the same way as the app. The base is what they share: one sign-on, one filesystem, one registry, one pipeline, one way to drive it.
How you work with it
-
Thinkube IDE is VS Code in the browser, with two kinds of coding agent installed and connected to the platform. This is where you say things. For professional code, use Claude Code from Anthropic: you sign in with your own Anthropic account, which is a paid service. Thinkube Tandem Chat runs entirely on your hardware, with the models you serve. It keeps everything on your machines and needs no account, but local models are less capable on hard code.
-
The agent reads freely and changes on request. Thinkube Control offers its operations to the agent through MCP (Model Context Protocol, the standard way coding agents call tools). The agent reads the cluster’s state whenever it needs to: what is running, what the GPUs hold, a build’s log. It changes something only when you ask: a deploy, a model load, an install, a restart. The platform draws that line, not the agent.
-
thinkube.yamldescribes an app in a few dozen lines: its containers, routes and the services it uses. The platform generates the Kubernetes objects from it, so you deploy to Kubernetes without writing Kubernetes. Thinkube GitOps shows one. -
Thinkube Control is the control panel: models, the LLM Gateway, templates, components, nodes, secrets, tokens. Everything the agent can do, you can click here, and the health of every service is on its card.
-
Thinkube Notebooks is JupyterLab with two environments built for your hardware,
fine-tuningandagent-dev, on the same filesystem as the IDE. -
The LLM Gateway is one address for every served model,
llm.<your domain>, speaking the OpenAI and Anthropic APIs. -
Optional components arrive with one sentence: vector stores, tracing, search, messaging, caches, inference backends. The components catalog lists them.
What comes out
What you build is a git repository with real code, a Containerfile, tests and a standard image. Your work is portable: the code, the Containerfile and the images are standard, and move with you to any other platform. The platform services it uses (a PostgreSQL database, S3-compatible storage, OpenID Connect sign-in, an OCI image registry) have close, highly compatible equivalents at the main cloud providers.
Why it is built this way
-
Production is a system, so use the one that exists. An address and a certificate, a login, a build that repeats, a rollback, a shared GPU, storage, traces: all of it exists in Kubernetes, maintained by thousands of people. Thinkube installs it unchanged. It chooses the base once, each part under an open licence that cannot be withdrawn. It puts a coding agent on top, so you operate the platform in sentences. Above the application it offers choices: three vector stores, four inference backends.
-
A model you own beats a model you rent. Small open models are good enough for most production tasks, and a fine-tuned small model beats a frontier model on the narrow task it was tuned for. The weights on your disk are the same weights tomorrow, nothing you send them leaves your machines, and once the GPU is bought every experiment is free. One gateway serves any of them behind the same API, so the app stays the same when the model changes.
-
Your data becomes capability. Prompting has limits. Fine-tuning teaches the model your format, vocabulary, tools and tone, so you can drop the long prompt and its cost and latency. Your traces, documents and checkers are yours alone. A program that can verify an answer can train a model with no labelled data, and the traces you record become data you can fine-tune on. The new model is served at the same address. You compare it with the previous one, then keep it or roll it back.
-
One person. One GPU serves one developer without a scheduler or a queue. One login and one set of credentials for one developer. A lab shares it the way lab equipment is shared: by saying who has the GPU this afternoon.
Where it runs
-
Operating system: Ubuntu 24.04, on one to a few machines.
-
Hardware: a DGX Spark (GB10, arm64) and x86-64 workstations with NVIDIA GPUs. The installer accepts any NVIDIA GPU from the Volta generation on, with driver 580 or newer.
-
Architectures: one cluster can mix arm64 and x86-64 machines. vLLM, TensorRT-LLM and Ollama are built for both. The catalogue’s TensorRT-LLM models are NVFP4, which needs a Blackwell GPU such as the DGX Spark’s.
-
Networks: you can move the cluster to a new network. It picks up a new network or a new address on its own (How it works).
How these docs are written
The pages are written for the reference cluster, the test cluster these pages were checked on. It has two amd64 workstations and one DGX Spark. The names of its nodes appear in the prompts and in the output:
| Node | Machine | GPU |
|---|---|---|
tkamd1 |
amd64 workstation |
none |
tkamd2 |
amd64 workstation |
2 × RTX 3090, 24 GB each |
tkspark |
DGX Spark, arm64 |
GB10, 121.7 GB unified |
Your cluster has its own machines and names. Where a page names one of these nodes, use a node of yours that meets the page’s Supported hardware, which states what the task needs on any cluster.
-
A line that starts with
›is what you say to the coding agent in Thinkube IDE; the lines under it are the answer. -
<your domain>is the domain you gave the installer.
Where next
-
Serve your first model: a model answering at your own address.
-
Build a web app from the template: an app at its own address, behind your login.
-
Install Thinkube: what you need, and the order to install in.