Thinkube Models
Fine-tune a model on rewards a program checks
Train a 4B model with GRPO against a solver that grades every answer, with no labelled data, and register the result beside its run
- Level
- advanced
- Time
- 9 h 30 min
- Risk
- medium
- Updated
- 2026-10-04
Overview
Basic idea
Fine-tuning usually needs examples of the right answer. When a program can decide whether an answer is right, the program replaces the examples: the model tries, the program grades, and training pushes the model towards what scored well.
The example is a zebra puzzle: five houses, fifteen clues, fill in the grid.
-
The model formalises. It translates each clue into one formal rule, as JSON.
-
The solver deduces. A constraint solver finds the grid from the rules, and grades it against the puzzle’s answer key.
-
GRPO learns from the grade. GRPO is a training method that learns from scores rather than example answers. For each puzzle the model writes eight answers; the ones that score above their group’s average are made more likely.
-
A LoRA adapter is what trains. Training changes only a small add-on (a LoRA adapter), 0.47% of the model, so it fits on one GPU.
-
Thinkube Experiments (MLflow) keeps the record. The run’s parameters and reward curve, and the registered model version linked to that run.
The same loop fits any task a program can check: text to SQL, documents to records, sentences to a knowledge graph.
What you’ll accomplish
You run zebra_grpo.ipynb on a GPU: it measures the base model, trains it, measures it again on your own puzzles and on the public ZebraLogic benchmark, registers the merged model zebra-rules-qwen35-4b in Thinkube Experiments, and makes it loadable at the LLM Gateway.
What to know before starting
Required
-
Run notebooks from your agent, watched or unattended: this run is long, so it runs unattended.
Optional
-
Thinkube Models: fine-tuning, rewards and distillation, and where the example notebooks are.
Supported hardware
-
GPU: one NVIDIA GPU with about 20 GB of memory free.
-
Architecture: amd64 or arm64.
Prerequisites
Platform
-
Thinkube running, with Thinkube IDE open. Ask your agent: "what’s running?"
-
unsloth/Qwen3.5-4Bmirrored: the notebook reads the base model from your storage. Ask your agent: "is unsloth/Qwen3.5-4B mirrored?" If it is not, ask "mirror unsloth/Qwen3.5-4B". -
A free GPU slot on the node. Ask your agent: "how many GPU slots are free on tkspark?" Replace tkspark with one of your nodes.
-
A private repository named
<your GitHub user>-metadataon GitHub, for the model’s catalogue entry. Ask your agent: "does my GitHub account have a -metadata repository?"
Components
-
The optional component vLLM, for the notebook’s last section. Ask your agent: "is vLLM installed?" If it is not, ask "install vLLM".
Instructions
Step 1. Choose the scale
FULL_RUN in the configuration cell picks the scale:
-
True, the default: 2,000 training puzzles, 200 in each evaluation set, 250 steps. It produces a model worth serving. -
False: 500 training puzzles, 50 in each evaluation set, 60 steps. Every stage runs the same way, on fewer puzzles.
Ask your agent:
› in examples/zebra-grpo/zebra_grpo.ipynb set FULL_RUN to False
The agent changes the line in the configuration cell. Leave it as it is for the full run.
Step 2. Run it unattended on the GPU node
A run of many hours gets a server of its own. The server stops when the run ends and gives the GPU back.
Ask your agent:
› run examples/zebra-grpo/zebra_grpo.ipynb unattended on tkspark with one GPU and the fine-tuning kernel
Or by hand: in Thinkube Notebooks start a server on tkspark with one GPU, open thinkube/notebooks/examples/zebra-grpo/zebra_grpo.ipynb with the fine-tuning kernel, and run all cells.
The reference run is one full run on the test cluster this page was checked on. Expected output from the first cells, from the reference run:
Model: unsloth/Qwen3.5-4B Scale: FULL Train: 2000, Eval: 200, Bench: 200 LoRA rank: 16, LR: 1e-05, Steps: 250 Loading from local MLflow mirror: /home/thinkube/thinkube/mlflow/artifacts/1/.../artifacts/model \\ /| NVIDIA GB10. Num GPUs = 1. Max memory: 121.689 GB. Platform: Linux. Parameters: 4,560,499,200 total, 21,233,664 trainable (0.47%) Train/eval disjoint: confirmed (2000 vs 200 unique) 5-house puzzles for evaluation: 200
Train/eval disjoint: confirmed is the check that the model is measured on puzzles it never trained on.
Step 3. Read the starting point
Before training, the model translates 200 of the generated puzzles and 200 ZebraLogic puzzles. ZebraLogic is written by other people, so it measures whether the skill transfers.
Ask your agent:
› show me the baseline results of the zebra run
Expected output, from the reference run:
Baseline Results: In-dist solved=8% cell_acc=0.313 parse=99% ZebraLogic solved=28% cell_acc=0.527 parse=97%
97 to 99% of the rules are valid. The errors are wrong rules, not broken ones.
Step 4. Follow the training
Each step writes eight answers per puzzle and scores them with the solver: the fraction of the 25 grid cells right, plus 0.05 for one rule per clue and 0.25 for a fully correct grid. The reward shows progress; the loss shows how far the model moved.
Expected output, from the reference run:
Num examples = 2,000 | Num Epochs = 1 | Total steps = 250 Batch size per device = 8 | Gradient accumulation steps = 4 Trainable parameters = 21,233,664 of 4,560,499,200 (0.47% trained) Training completed in 414.2 minutes MLflow run: ec1f9ee6e1fa455b99936a9ef9f32e51 experiment: zebra-grpo (id 5) logged 6256 metric points across 251 steps view: https://experiments.<your domain>/#/experiments/5/runs/ec1f9ee6e1fa455b99936a9ef9f32e51
The run is logged to Thinkube Experiments once training ends, so the whole reward curve is in one run. On the reference run the reward reached its ceiling around step 190.
Step 5. Compare before and after
The same puzzles, the same sampling and the same grader, with the trained adapter.
Expected output, from the reference run:
baseline trained delta ========================================================== in-dist puzzle_acc 8% 99% +92% in-dist cell_acc 0.313 0.995 +0.682 in-dist parse_rate 99% 100% +1% zlogic puzzle_acc 28% 67% +39% zlogic cell_acc 0.527 0.782 +0.255 zlogic parse_rate 97% 100% +3% Results logged to MLflow run ec1f9ee6e1fa455b99936a9ef9f32e51
ZebraLogic is the row that shows real learning: from 28% to 67% of puzzles solved, on clues the generator never wrote. The next section prints one unseen puzzle clue by clue, with the rule the model wrote and the solver’s grid.
Step 6. Register the model and make it loadable
Section 14 of the notebook merges the adapter into the base model. It stores the model, adds it to your catalogue, and waits until the gateway lists it. The model is registered as a version of zebra-rules-qwen35-4b, linked to the training run.
The catalogue entry needs every field, the vLLM backend and the BF16 format. The notebook stops on an entry that asks for anything else.
The cell prints, with Unsloth’s progress bars left out:
Merging the adapter into the base weights: /home/thinkube/thinkube/mlflow/.staging/zebra-rules-qwen35-4b Uploading 9 files to s3://mlflow/artifacts/<experiment>/<run>/artifacts/model model.safetensors-00001-of-00002.safetensors (5.0 GB)... model.safetensors-00002-of-00002.safetensors (3.7 GB)... Upload complete Registered zebra-rules-qwen35-4b version 1, linked to run <run> Catalogue entry written to <your GitHub user>/<your GitHub user>-metadata/models.json Registration recorded for zebra-rules-qwen35-4b zebra-rules-qwen35-4b: ModelState.deployable
In Thinkube Experiments, the version opens onto the run: its parameters and reward curve.
Step 7. Give the GPU back
The notebook’s last section loads the model on vLLM and calls it, as Serve a model you fine-tuned describes. On the reference run it loaded on vllm-tkamd2. The unattended run’s server stopped when the notebook ended; the served model holds its GPU until you unload it.
Ask your agent:
› unload zebra-rules-qwen35-4b
Expected output:
state: deployable message: Model unloaded successfully
Step 8. Next steps
To undo: Delete the registered version in Thinkube Experiments. The base model is untouched.
-
Serve a model you fine-tuned: load the registered model and call it from your code.
-
Load a model on the node and context you choose: the node and the context for the load.
-
Thinkube Models: when rewards, supervised fine-tuning or distillation fit.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
The model cell prints |
|
Stop the run and ask your agent "mirror unsloth/Qwen3.5-4B"; every GPU node then reads the one mirror. |
A compiled kernel fails on a tensor without storage |
An extension was built against a different PyTorch than the image’s |
Run |