Thinkube Models
Build a research assistant over your papers
Index arXiv papers in your own vector store, ask questions that come back with sources, and let agents argue from the index
- Level
- intermediate
- Time
- 30 min
- Risk
- low
- Updated
- 2026-10-04
Overview
Basic idea
A research assistant answers from documents you chose and shows which ones it used. It is built from services the platform already runs: a served chat model and embedding model, a vector store, and tracing.
-
Fetch and chunk. Papers come from arXiv and are cut into pieces of about 1,000 characters.
-
Embed and store. Each piece becomes a vector through the LLM Gateway and is stored in Qdrant with what it came from.
-
Retrieve and answer. A question finds the closest pieces, and the chat model answers from them, citing the papers.
-
Trace. Langfuse records every retrieval and every answer, with its latency.
-
Argue. Two AG2 agents take opposite sides of a question and can cite only what a search of your index returns; a judge scores them.
What you’ll accomplish
You run the three research-assistant notebooks in order: 01 checks the models, 02 indexes 146 papers and answers three questions with sources, and 03 runs a debate between two agents and a judge over that index.
What to know before starting
Required
-
Run notebooks from your agent, watched or unattended: how the agent runs a notebook.
-
Load a model on the node and context you choose: serving the two models.
Optional
-
Thinkube Models: embeddings, retrieval, agents and tracing, and where the example notebooks are.
Supported hardware
-
GPU: not needed for the notebooks. The embedding model and the language model are served behind the LLM Gateway on the cluster’s GPU nodes.
-
Architecture: amd64 or arm64.
Prerequisites
Platform
-
Thinkube running, with Thinkube IDE open. Ask your agent: "what’s running?"
-
A notebook server on tkamd1. Ask your agent: "start a notebook server on tkamd1". Replace tkamd1 with one of your nodes.
-
A chat model the catalogue marks for tool use, and an embedding model, both loaded. Ask your agent: "which models are loaded?" The reference run, one run on the test cluster this page was checked on, used
unsloth/Qwen3.6-27B-NVFP4andQwen/Qwen3-Embedding-0.6B. -
Outbound access to arXiv.
Components
-
Qdrant and Langfuse. Ask your agent: "are Qdrant and Langfuse installed?" If not, ask "install Qdrant and Langfuse"; the install brings their requirements, ClickHouse and Valkey.
Instructions
Step 1. Check the models are served
01-load-models.ipynb lists the models and the GPUs, loads a chat model and an embedding model when none is loaded, and tests both. When both are loaded, it says so and moves on.
Ask your agent:
› run examples/research-assistant/01-load-models.ipynb with agent-dev on tkamd1
Expected output, from the reference run:
✅ Chat model already loaded: unsloth/Qwen3.6-27B-NVFP4 ✅ Embedding model already loaded: Qwen/Qwen3-Embedding-0.6B ✅ Embeddings with Qwen/Qwen3-Embedding-0.6B — 1024 dimensions 📐 Semantic Similarity: LoRA ↔ QLoRA: 0.833 (related topics) LoRA ↔ Weather: 0.289 (unrelated) QLoRA ↔ Weather: 0.260 (unrelated)
Texts close in meaning get vectors close to each other. The next notebook uses this to find related text. Leave the models loaded.
Step 2. Index the papers and ask questions
02-langchain-rag.ipynb fetches papers on reasoning post-training from arXiv, chunks them, embeds the chunks, stores them in Qdrant, and answers three questions through a LangChain retrieval chain traced in Langfuse.
Ask your agent:
› run examples/research-assistant/02-langchain-rag.ipynb with agent-dev on tkamd1
Expected output, from the reference run:
✅ Total: 146 papers from ArXiv ℹ️ - Anchor papers: 3 ℹ️ - GRPO/RLVR papers: 50 ℹ️ - Distillation papers: 50 ℹ️ - General reasoning papers: 43 ✅ Created 431 total chunks from 146 papers ✅ All 431 embeddings generated (dimension: 1024) ✅ Created collection 'rl_reasoning_papers' ✅ Uploaded 431 vectors to Qdrant
The index holds abstracts. FETCH_FULL_TEXT = True in the fetch cell also downloads each paper’s LaTeX source and extracts algorithms, results tables and equations.
Each question comes back with an answer and four sources:
| Question | Time | Answer, in short, and sources |
|---|---|---|
What is GRPO and how does it differ from PPO for training language models? |
8.98 s |
GRPO is a widely used algorithm for reinforcement learning with verifiable rewards; the excerpts say nothing about PPO, so the answer says it cannot compare them. Sources: GRPO-SG, CoDistill-GRPO, GRPO-VPS. |
Does reinforcement learning with verifiable rewards create new reasoning ability, or mostly sharpen what the base model already can do? |
11.51 s |
It mostly sharpens what the base model already has. Sources include Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? |
When does distilling reasoning traces from a large teacher beat running RL directly on a small model? |
9.48 s |
Distilling a teacher’s traces generally beats RL on the small model alone. Sources: RLKD, NaturalThoughts and two more. |
Answers are only as good as the retrieved text. The first answer shows the chain saying what its sources do not cover.
Step 3. Read the traces
The last cells of 02 read the traces back from Langfuse and summarise latency per kind of trace.
Expected output, from the reference run:
✅ Fetched 12 traces from Langfuse ℹ️ Excluded 2 trace(s) newer than 30s — their latency is not final yet 🔹 RetrievalQA Traces: 7 Median: 11171 ms (11.17 sec) 💡 Measured: RetrievalQA has the higher median latency at 11.2 s, against 3.2 s for embedding-generation — a factor of 3.5.
Every trace, with its prompts and retrieved pieces, is at https://langfuse.<your domain>.
Step 4. Let two agents argue from the index
03-multi-agent.ipynb gives three AG2 agents a question the literature has not settled: for a small model with a reward checker and a large teacher, post-train with GRPO or distil the teacher’s traces? Each advocate searches the index once, argues, and ends its turn with ARGUMENT_COMPLETE. The judge scores five criteria and declares a winner. The judge writes code for the winning approach.
Ask your agent:
› run examples/research-assistant/03-multi-agent.ipynb with agent-dev on tkamd1
Expected output, from the reference run:
Timing: Round 1 (Openings): 71.9s Round 2 (Rebuttals): 60.0s Judge: 98.6s Total: 230.5s Code Generated: Yes (ok) - Winning approach: Distillation - Algorithm excerpts used: 3 - Paper citations: 2312.10730v2, 2601.18734v3, 2212.00193v2
and the judge’s totals:
GRPO Total: 0.68 Distillation Total: 0.76 The side with the higher total score is Distillation.
Pro-GRPO opened with Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs, DeepSeek-R1 and RLKD; Pro-Distillation with work on chain-of-thought distillation into small models. Tool calling lets the agents search: the model writes a search call, and the server reads it.
Step 5. Close the notebooks
Each open notebook holds a kernel and its memory on the server.
Ask your agent:
› close the research assistant notebooks on tkamd1
Expected output:
saved: true kernel_shut_down: true
The models stay loaded for other work. Unload the ones you loaded for this playbook with "unload <model>".
Step 6. Next steps
To undo: In a notebook, QdrantClient(url=os.environ["QDRANT_URL"]).delete_collection("rl_reasoning_papers").
-
Fine-tune a model on rewards a program checks: train with rewards a program checks, the GRPO side of the debate.
-
Call your models from an app or a script: the same gateway from your own code.
-
Install and remove optional components: another vector store, such as Chroma or Weaviate.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
A chat call answers with no visible text and |
A reasoning model spent |
Raise |