Pretraining gives you completion; fine-tuning gives you compliance
Pretraining optimizes exactly one thing: predict the next token, over trillions of tokens of ordinary text. Nothing in that objective knows about "questions" or "answers" or "being helpful." Hand a raw base model the prompt How do I boil an egg? and it is just as likely to continue with How long does it take? What about a soft-boiled one? — more plausible internet text — as it is to answer you. The model is not being difficult. It is doing exactly what it was trained to do: continue the distribution it saw.
Supervised fine-tuning (SFT) — what most people mean when they say "fine-tuning" without qualification — closes that gap with no architecture change at all. Same transformer, same next-token cross-entropy loss, same optimizer. The only thing that changes is the data: instead of "the internet," the model trains on curated (instruction, response) pairs, wrapped in a chat template that marks who is speaking. A few thousand to a few hundred thousand well-written examples are usually enough to shift a model from "completes anything" to "answers you" — SFT is teaching a format and a role, which turns out to need far less data than teaching new knowledge does (a distinction the pitfalls below return to).
One detail that trips people up the first time they write an SFT loop by hand: you do not compute loss on the whole sequence. The prompt tokens are masked out — the model is not being trained to predict your question, only the assistant's response to it. Get this wrong and the model partially learns to imitate users instead of answering them, and burns capacity re-deriving prompts it will never need to generate at inference time.
Base model and instruction-tuned model are usually the exact same architecture at the exact same size; Llama-3-8B and Llama-3-8B-Instruct differ only in what was done to the weights after pretraining finished. Everything else in this module is either a cheaper way to do SFT, or a second stage that runs after it.
<user> How do I boil an egg? <assistant> Bring water to a boil... <eos>
NO NO NO NO NO NO NO NO NO YES YES YES ... YES
\_________ prompt tokens: masked out, zero gradient _________/\___ response: normal loss ___/
REMEMBERA base model completes text; an instruction-tuned model answers you — and the gap between them is a training stage, not an architecture change.