Training LLMs

ai , machine-learning

I don’t train models for a living — I build on top of them. But the models I build on are made in two very different steps, and almost everything I care about as a consumer happens in the second one. Pre-training reads the internet and installs knowledge; post-training turns that into an assistant with a format, a personality, and lately the ability to reason and write working code. These are my working notes on that second step — kept deliberately shallow, enough to know which lever to pull, not to go pull it myself.

AI-augmented, roughly 50% AI (what this means)

I love AI, and I use it on posts like this one partly to get better at using it. The percent is a rough guess: after rounds of prompting, rewriting and editing, it's very hard to say which words came from whom.

How a model gets made

Three stages, and almost everything else is a footnote to them:

  1. Pre-training — show the model a huge pile of text and have it predict the next token, over and over. No labels, just text, which is why it scales: raw text is basically free. This is where the compute and money go — the headline “$X million to train” numbers are almost entirely pre-training — and it’s where facts, skills, and the model’s “world model” come from. If a model doesn’t know something, it usually didn’t see enough of it here. The output is a base model: it has read the internet but isn’t an assistant. Prompt it with a question and it’ll happily autocomplete ten more questions.
  2. Post-training — take that base model and shape its behavior: answer instead of autocomplete, follow instructions, hold a format, refuse the obviously bad stuff, and — the recent addition — think before answering. A rounding error next to pre-training’s cost, and most of what makes a chat model feel like one. It has its own post; the short version is below.
  3. Deployment — shrink and serve the thing so it runs fast and cheap. Briefly below; the full story is /ai-inference.

The one intuition to keep: pre-training installs knowledge; post-training shapes behavior. That single line drives the whole post-training vs RAG vs the harness decision below.

A caveat before the map: a lot of this is still alchemy — recipes kept because they work, not because anyone can say why. Why that is, in the appendix ↓

Post-training

Every post-training method is the same move: pick a behavior you want more of, find a signal that says which outputs have it, and nudge the weights toward it. The methods differ in where the signal comes from, and that’s the lineage:

  1. Demonstrations → SFT. Show it good answers; it imitates them. The first and biggest shift.
  2. Preferences → RLHF, or DPO for the same data with less machinery. Show it two answers and which is better; it learns the taste.
  3. Verifiable rewards → RLVR, with GRPO as the optimizer that made it cheap. Skip the human: let a checker grade the answer. This is what made reasoning and coding models take off.

Each rung stands on the one below — SFT gets the model into the right neighborhood, preference tuning polishes, and verifiable rewards push hard on whatever you can actually grade. Each method — what it optimizes, the data it needs, when to reach for it, how it goes wrong — and the recipes labs actually run (InstructGPT, Constitutional AI, Tülu 3, DeepSeek-R1) have their own post:

Methods at a glance

Method Signal (who or what grades) Separate reward model? Best for Watch-out
SFT Human- or strong-model-written target answers No The first shift: answer instead of autocomplete, hold a format Can’t exceed the demos or learn what not to do
RLHF Humans rank A vs B → reward model → PPO Yes Helpfulness, tone, safety beyond what demos teach Heavy pipeline; reward hacking
DPO The same A-vs-B rankings, fit directly No RLHF’s benefit without the RL loop Bounded by the preference data
RLAIF A model ranks A vs B against written principles Yes RLHF at scale, or when nothing is checkable The judge’s blind spots become the model’s
RLVR A checker: unit tests, math grader, sandbox No — the checker is the reward Reasoning and coding agents, anything checkable Only where checkable; gaming the test
GRPO RLVR’s checker; a group of sampled answers is the baseline No — and no value model either Reasoning RL on a budget; the usual RLVR optimizer Length bias (Dr. GRPO), entropy collapse (DAPO)

Datasets: what you train on vs what you grade on

The data splits into what you train on and what you grade on — and under RLVR the two collapse, because the eval’s checker is the reward.

Training data — the (prompt → good answer) sets behavior gets shaped on:

Evals / benchmarks — held-out tasks you score against, not train on. For coding agents they’re also the RLVR reward target (see Post-training for coding competence):

  • SWE-bench — real GitHub issues paired with the repo they came from; the model must produce a patch that makes the repo’s hidden tests pass. Binary grade, no LLM-judging style. The paper drew from popular Python repos; the subset everyone reports is SWE-bench Verified, 500 tasks each hand-checked by human developers so a correct patch can’t fail on a broken test or ambiguous issue.
  • Terminal-bench — agentic, end-to-end command-line tasks in a real sandbox (build a kernel, stand up a git server, debug a broken system), scored purely by whether the task got done. Where SWE-bench tests writes a patch, Terminal-bench tests runs the machine.

How does post-training differ from RAG and the harness?

There are three places to change how a model behaves, and picking the right one saves enormous effort. Post-training changes the weights; RAG and the harness change things at runtime. Runtime is cheaper, faster to iterate, and updatable — so the order I actually reach for is harness → RAG → post-train, and I rarely get to the third.

Layer What it changes Reach for it when Iterate in
Harness / MCP (runtime) what the model can do and how it acts it needs tools, or a different way of working seconds
RAG (runtime) what the model knows it needs your facts or fresh data seconds
Post-training (weights) the model’s default behavior the change must hold for everyone, every call days
  • New facts — your docs, today’s data, private knowledge → RAG. Retrieve the relevant text at query time and put it in context. You can’t reliably fine-tune facts in; that’s pre-training’s job, and fine-tuning on a handful of examples teaches style, not knowledge — then cheerfully hallucinates the gaps.
  • New actions, or a different way of working — call an API, read a file, search the web, plan before acting, always check its work → the harness: the system prompt, the tools you expose (increasingly through MCP, a standard way to hand a model tools), the agent loop, memory, retries. The weights don’t change; the scaffolding gets bigger. Building this well is most of what /chop and the agent cockpit are about.
  • A new default everywhere — behavior that has to hold for everyone without re-explaining it every call, or that prompting just can’t make reliable → post-training (the methods above).

Rule of thumb: facts → RAG, actions → harness, baked-in defaults → post-train. When in doubt, push the change as far toward runtime as it’ll go.

Post-training for coding competence

The map above is abstract until you watch it chase a target. Coding agents are the cleanest example, because “did it work” is something a computer can check — which is exactly the RLVR setup, pointed at software.

You get what you measure

So first, define the target by the eval. “Coding competence” for an agent isn’t a vibe; it’s two questions a benchmark can answer (both are described above):

  • SWE-bench — can it fix real software? Hand the model a real GitHub issue and its repo; the grade is binary: do the hidden tests pass? The tests decide, not a style judge.
  • Terminal-bench — can it actually operate a computer? Drop the agent in a real terminal sandbox with an end-to-end job, scored by whether it’s done.

Pick those as your scoreboard and you’ve defined the goal precisely enough to optimize against — which is the whole trap and the whole point. You get what you measure, so measure the thing you actually want.

The loop

With the target pinned, post-training is the same lineage, run as a loop:

  1. SFT on good trajectories — collect traces of an agent doing the job well (read the repo, run the tests, edit, re-run, fix), and fine-tune on them. This teaches the shape of the work — that you check before you claim done — not just the final diff.
  2. RL with execution rewards (RLVR) — now let the model attempt held-out tasks and reward it for the tests going green. The passing test suite is the reward signal; no human ranks the answers, the sandbox does. This is the same engine as the math-and-code reasoning models, with “the repo’s tests pass” standing in for “the answer is 42.” The Goblin walkthrough below is a hands-on tour of this exact RL loop.
  3. Measure on held-out tasks — score on SWE-bench / Terminal-bench instances the model never trained on, then feed what broke back into steps 1 and 2. Iterate.

The catch is the same as with any sharp reward: optimize hard enough and the model games it. Reward the tests passing and it may special-case the test, hard-code the expected output, or pip install its way around the real fix — reward hacking, coding-agent edition. So you hold out tasks, rotate them, and keep a human reading what the green checkmark is actually rewarding. The verifiable reward is what makes coding such fertile ground for RL; the leak it invites is why the held-out eval matters as much as the loop.

Deployment: quantization and serving

Not training, but it’s where the model you actually run comes from. Serving is inference, and that story is its own post:

  • Quantization — compress the weights from 16 bits per parameter down to ~4 bits or less (Q4_0, IQ2_XXS, and friends). How the methods work, and a scorecard comparing them. It’s lossy compression, so you have to eval the damage — a common check is comparing the difference in answers via embeddings.
  • GGUF (GPT-Generated Unified Format) — the file format these models are stored in (GGUF is GGML’s successor). It’s what you download when you grab a quantized model to run locally.
  • Introduction to Weight Quantization — the explainer I keep going back to.

See how it works

For the concepts — how a neural network actually learns, then how a transformer turns that into language — nothing beats 3Blue1Brown’s neural networks series. Start with gradient descent and backprop (that is training), then the transformers and GPT chapters.

For the mechanics, Brendan Bycroft’s LLM visualization walks a single token through every layer of a GPT model in 3D — embeddings, attention, the lot.

And for post-training specifically, How to Train Your Goblin (mchen and Will Brown, on Prime Intellect) is a playful scroll-through of RL — it retraces how GPT picked up its accidental “goblin” tic by deliberately RL-training models to overuse the word, hidden trigger reward and all. A concrete look at reward hacking and how RL differs from SFT, with the code and training runs open.

What this post is not about

To keep the map focused, a few neighbors that live elsewhere:

  • Image / diffusion models — a different training story altogether (denoising, not next-token). See /ai-image for generating images of yourself, and /ai-art for the art side.
  • The actual math — I’m staying at the mental-model level on purpose. For the deep version, start with the seminal papers. For the intuition behind language models, see /gpt.
  • Distributed-training infrastructure — the 8,000-GPU engineering problem. Real, hard, and out of scope for me.

For the basics (“what even is an LLM”), /ai-faq is the canonical reference; for putting models to work, see /chop; for checking they actually work, /ai-testing.

Appendix: engineering, science, and alchemy

Training LLMs is three jobs at once. Two have clean definitions: science figures out a true fact about a model that already exists; engineering makes a system hit a target. The same question word tells you which — “why is this true?” is science, “how do I hit this number?” is engineering.

Concrete pairs, same topic on each line:

Topic Science question Engineering question
Scaling Why does loss drop predictably with compute? How do I train this 70B model on 8k GPUs without it crashing?
Inference What is the model actually computing in these layers? How do I serve it at 50ms/token without going broke?
Generalization Why don’t overparameterized nets just memorize? How do I stop this model overfitting on my data?
Alignment Does the model have goals? Can it deceive? How do I make it refuse harmful requests in prod?
Capabilities Why do new abilities appear suddenly at scale? How do I get reliable tool-calling out of what I have?
Data What did the model actually learn from this corpus? How do I dedup, filter, and decontaminate 10T tokens?

Science output is a fact (“loss scales as a power law”). Engineering output is a working thing (“a serving stack that hits the SLA”).

The one insight worth keeping: in AI the usual order is reversed. Normally science comes first — thermodynamics, then engines. In deep learning we built the thing first and are still reverse-engineering why it works, so a lot of AI “science” is closer to biology (dissect an organism you didn’t design) than physics. That’s why the line feels blurry: the systems are running ahead of the explanations.

Which brings in the third job. A lot of training is still alchemy — you curate the data, pick the recipe, run it, and see what comes out, and when it works the explanation usually arrives later, if at all. It can feel like macrodata refinement in Severance: sort the numbers that feel wrong into the bin, without being told why they’re wrong or what the bin is for. The difference is ours ships a product at the end.