Thinkube Models

Fine-tune a model on rewards a program checks

Train a 4B model with GRPO against a solver that grades every answer, with no labelled data, and register the result beside its run

Level
advanced
Time
9 h 30 min
Risk
medium
Updated
2026-10-04

GRPOUnslothThinkube ExperimentsNotebooks

Overview

Basic idea

Fine-tuning usually needs examples of the right answer. When a program can decide whether an answer is right, the program replaces the examples: the model tries, the program grades, and training pushes the model towards what scored well.

The example is a zebra puzzle: five houses, fifteen clues, fill in the grid.

  • The model formalises. It translates each clue into one formal rule, as JSON.

  • The solver deduces. A constraint solver finds the grid from the rules, and grades it against the puzzle’s answer key.

  • GRPO learns from the grade. GRPO is a training method that learns from scores rather than example answers. For each puzzle the model writes eight answers; the ones that score above their group’s average are made more likely.

  • A LoRA adapter is what trains. Training changes only a small add-on (a LoRA adapter), 0.47% of the model, so it fits on one GPU.

  • Thinkube Experiments (MLflow) keeps the record. The run’s parameters and reward curve, and the registered model version linked to that run.

The same loop fits any task a program can check: text to SQL, documents to records, sentences to a knowledge graph.

What you’ll accomplish

You run zebra_grpo.ipynb on a GPU: it measures the base model, trains it, measures it again on your own puzzles and on the public ZebraLogic benchmark, registers the merged model zebra-rules-qwen35-4b in Thinkube Experiments, and makes it loadable at the LLM Gateway.

What to know before starting

Required

Optional

  • Thinkube Models: fine-tuning, rewards and distillation, and where the example notebooks are.

Supported hardware

  • GPU: one NVIDIA GPU with about 20 GB of memory free.

  • Architecture: amd64 or arm64.

Prerequisites

Platform

  • Thinkube running, with Thinkube IDE open. Ask your agent: "what’s running?"

  • unsloth/Qwen3.5-4B mirrored: the notebook reads the base model from your storage. Ask your agent: "is unsloth/Qwen3.5-4B mirrored?" If it is not, ask "mirror unsloth/Qwen3.5-4B".

  • A free GPU slot on the node. Ask your agent: "how many GPU slots are free on tkspark?" Replace tkspark with one of your nodes.

  • A private repository named <your GitHub user>-metadata on GitHub, for the model’s catalogue entry. Ask your agent: "does my GitHub account have a -metadata repository?"

Components

  • The optional component vLLM, for the notebook’s last section. Ask your agent: "is vLLM installed?" If it is not, ask "install vLLM".

Instructions

Step 1. Choose the scale

FULL_RUN in the configuration cell picks the scale:

  • True, the default: 2,000 training puzzles, 200 in each evaluation set, 250 steps. It produces a model worth serving.

  • False: 500 training puzzles, 50 in each evaluation set, 60 steps. Every stage runs the same way, on fewer puzzles.

Ask your agent:

› in examples/zebra-grpo/zebra_grpo.ipynb set FULL_RUN to False

The agent changes the line in the configuration cell. Leave it as it is for the full run.

Step 2. Run it unattended on the GPU node

A run of many hours gets a server of its own. The server stops when the run ends and gives the GPU back.

Ask your agent:

› run examples/zebra-grpo/zebra_grpo.ipynb unattended on tkspark with one GPU and the fine-tuning kernel

Or by hand: in Thinkube Notebooks start a server on tkspark with one GPU, open thinkube/notebooks/examples/zebra-grpo/zebra_grpo.ipynb with the fine-tuning kernel, and run all cells.

The reference run is one full run on the test cluster this page was checked on. Expected output from the first cells, from the reference run:

Model: unsloth/Qwen3.5-4B
Scale: FULL
Train: 2000, Eval: 200, Bench: 200
LoRA rank: 16, LR: 1e-05, Steps: 250
Loading from local MLflow mirror: /home/thinkube/thinkube/mlflow/artifacts/1/.../artifacts/model
   \\   /|    NVIDIA GB10. Num GPUs = 1. Max memory: 121.689 GB. Platform: Linux.
Parameters: 4,560,499,200 total, 21,233,664 trainable (0.47%)
Train/eval disjoint: confirmed (2000 vs 200 unique)
5-house puzzles for evaluation: 200

Train/eval disjoint: confirmed is the check that the model is measured on puzzles it never trained on.

Step 3. Read the starting point

Before training, the model translates 200 of the generated puzzles and 200 ZebraLogic puzzles. ZebraLogic is written by other people, so it measures whether the skill transfers.

Ask your agent:

› show me the baseline results of the zebra run

Expected output, from the reference run:

Baseline Results:
  In-dist    solved=8%  cell_acc=0.313  parse=99%
  ZebraLogic solved=28%  cell_acc=0.527  parse=97%

97 to 99% of the rules are valid. The errors are wrong rules, not broken ones.

Step 4. Follow the training

Each step writes eight answers per puzzle and scores them with the solver: the fraction of the 25 grid cells right, plus 0.05 for one rule per clue and 0.25 for a fully correct grid. The reward shows progress; the loss shows how far the model moved.

Expected output, from the reference run:

Num examples = 2,000 | Num Epochs = 1 | Total steps = 250
Batch size per device = 8 | Gradient accumulation steps = 4
Trainable parameters = 21,233,664 of 4,560,499,200 (0.47% trained)
Training completed in 414.2 minutes
MLflow run: ec1f9ee6e1fa455b99936a9ef9f32e51
  experiment: zebra-grpo (id 5)
  logged 6256 metric points across 251 steps
  view: https://experiments.<your domain>/#/experiments/5/runs/ec1f9ee6e1fa455b99936a9ef9f32e51

The run is logged to Thinkube Experiments once training ends, so the whole reward curve is in one run. On the reference run the reward reached its ceiling around step 190.

Step 5. Compare before and after

The same puzzles, the same sampling and the same grader, with the trained adapter.

Expected output, from the reference run:

                             baseline    trained      delta
==========================================================
in-dist  puzzle_acc                8%        99%       +92%
in-dist  cell_acc               0.313      0.995     +0.682
in-dist  parse_rate               99%       100%        +1%
zlogic  puzzle_acc                28%        67%       +39%
zlogic  cell_acc                0.527      0.782     +0.255
zlogic  parse_rate                97%       100%        +3%
Results logged to MLflow run ec1f9ee6e1fa455b99936a9ef9f32e51

ZebraLogic is the row that shows real learning: from 28% to 67% of puzzles solved, on clues the generator never wrote. The next section prints one unseen puzzle clue by clue, with the rule the model wrote and the solver’s grid.

Step 6. Register the model and make it loadable

Section 14 of the notebook merges the adapter into the base model. It stores the model, adds it to your catalogue, and waits until the gateway lists it. The model is registered as a version of zebra-rules-qwen35-4b, linked to the training run.

The catalogue entry needs every field, the vLLM backend and the BF16 format. The notebook stops on an entry that asks for anything else.

The cell prints, with Unsloth’s progress bars left out:

Merging the adapter into the base weights: /home/thinkube/thinkube/mlflow/.staging/zebra-rules-qwen35-4b
Uploading 9 files to s3://mlflow/artifacts/<experiment>/<run>/artifacts/model
  model.safetensors-00001-of-00002.safetensors (5.0 GB)...
  model.safetensors-00002-of-00002.safetensors (3.7 GB)...
Upload complete
Registered zebra-rules-qwen35-4b version 1, linked to run <run>
Catalogue entry written to <your GitHub user>/<your GitHub user>-metadata/models.json
Registration recorded for zebra-rules-qwen35-4b
zebra-rules-qwen35-4b: ModelState.deployable

In Thinkube Experiments, the version opens onto the run: its parameters and reward curve.

Step 7. Give the GPU back

The notebook’s last section loads the model on vLLM and calls it, as Serve a model you fine-tuned describes. On the reference run it loaded on vllm-tkamd2. The unattended run’s server stopped when the notebook ended; the served model holds its GPU until you unload it.

Ask your agent:

› unload zebra-rules-qwen35-4b

Expected output:

state: deployable
message: Model unloaded successfully

Step 8. Next steps

To undo: Delete the registered version in Thinkube Experiments. The base model is untouched.

Troubleshooting

Symptom Cause Fix

The model cell prints No local mirror; downloading from HuggingFace

unsloth/Qwen3.5-4B is not mirrored, so the notebook would download it into the server

Stop the run and ask your agent "mirror unsloth/Qwen3.5-4B"; every GPU node then reads the one mirror.

A compiled kernel fails on a tensor without storage

An extension was built against a different PyTorch than the image’s

Run import torch, os; print(torch.version, os.path.dirname(torch.file)); the path must be under /usr/local/lib/python3.12/dist-packages, the image’s build.