A minimal, heavily-commented GPT-style language model for learning purposes. Every component is implemented from scratch β no HuggingFace, no pre-built transformers.
Most from-scratch LLM repos optimize for one thing: getting the training loss down. Karpathy's nanoGPT is the gold standard for that β but reproducing a GPT-2-quality run out of it means renting cloud A100s, which realistically costs on the order of $100+ and hours of babysitting a training job you can't iterate on casually.
TinyLLM optimizes for the other thing: understanding. Every stage a real LLM pipeline goes through β tokenizer training, pretraining on raw text, instruction fine-tuning, reasoning fine-tuning with tool-use, and a chat interface to actually talk to the result β is here, small enough to run end-to-end on a single consumer GPU you already own, in minutes to hours instead of days. You're not spectating a loss curve on someone else's infrastructure; you're stepping through the same pipeline GPT/LLaMA use, just at a scale where you can read every line, change it, and watch what breaks.
| Resource | Minimum | Notes |
|---|---|---|
| GPU | None required | Falls back to CPU automatically (train.py/pretrain.py) β slower, but every script runs |
| GPU (recommended) | RTX 3060 (12 GB) or similar Ampere+ card | Ampere+ is needed to actually benefit from bfloat16 mixed precision (see Mixed Precision); the default "Tiny" config trains comfortably within a few GB of VRAM |
| RAM | 16 GB | Covers dataset downloads (Alpaca/Dolly/GSM8K/WikiText-2 are all small, low hundreds of MB combined) and tokenizer training |
| Disk | ~2 GB free | Raw datasets + checkpoints + optional W&B logs |
| Python | 3.10+ | With PyTorch (CUDA build optional, see Quick Start) |
No multi-GPU setup is required β DDP support exists for when you have more than one GPU, not as a prerequisite for anything in this repo to work.
- Chat UI
- What You'll Learn
- Datasets & Data Pipeline
- Quick Start
- Limitations & Caveats
- Experiment Ideas
- Key Papers
This is a real, unedited session in webchat.py against a 50M-param checkpoint
(train_reasoning.py --dataset_format chatml) trained jointly on multi-turn small talk and
1-/2-/3-hop synthetic word problems, pooled as equals rather than one being "replay" for the
other β see data_utils.prepare_multitask_data. Turns are rendered ChatML-style
(<|im_start|>{role}...<|im_end|>, same scheme Qwen/GPT use), and the loss is masked to
assistant spans only.
"Good morning" and "Can you tell me a joke?" are answered exactly as trained β this is the
point of the demo, not a limitation: the goal of this stage isn't teaching the model to
generalize, it's taking it from outputting nonsense to reliably answering things close to
what it was trained on. The word problem is different: it's pulled from
data/reasoning_heldout.json, held out and never trained on. The reasoning trace is real β
each <CALC> call is the model deciding when and what to compute, with the actual
arithmetic result injected rather than guessed (model/calculator.py, model/generate.py)
β and renders as a collapsible "Thoughts" section, Claude/ChatGPT-style, instead of inline
with the answer.
Once you have a checkpoint, there are two ways to talk to it:
chat_demo.pyβ CLI demo that feeds a fixed sequence of questions to prove the pipeline works end-to-end.webchat.pyβ a real chat UI in your browser, for typing arbitrary follow-ups yourself and watching the answer stream in token by token, the same way a real LLM actually generates:
python -m demo.webchat \
--checkpoint checkpoints/multitask_chatml/final.pt \
--tokenizer_path checkpoints/tokenizer.json \
--format chatml
# opens http://127.0.0.1:8765It's a tiny stdlib-only HTTP server (no extra pip install beyond training) serving a
static chat page. /chat streams newline-delimited JSON as generation proceeds β
model/generate.py's on_token hook fires once per token actually appended to the
sequence, so the browser renders each token as it's produced instead of blocking until
the whole reply is ready. It's stateless by design: the browser resends the full
conversation history with every request and the server rebuilds the prompt from scratch
each time β the same mechanic real chat APIs use, and it means what you see is exactly
what the model does with the conversation so far, not some server-side memory trick.
This model is small enough to generate a full reply in well under a second on either
CPU or GPU β much faster than real LLM APIs, whose pace is normally set by network +
far larger model compute. The GIF above was recorded with --stream_delay_ms 70, an
artificial per-token pause meant for demos/recordings; it defaults to 0 (as fast as
the model actually generates) for real use.
Byte Pair Encoding is how GPT-2/GPT-4 turn raw text into numbers.
The algorithm:
- Start with characters as vocabulary
- Count all adjacent token pairs in the corpus
- Merge the most frequent pair into a new token
- Repeat until vocab size is reached
π‘ Key insight:
"lower"β["βlow", "er"],"lowest"β["βlow", "est"]. Common subwords get merged; rare words stay as characters. Handles OOV words gracefully.
Special tokens:
| Token | ID | Purpose |
|---|---|---|
<PAD> |
0 | Padding |
<UNK> |
1 | Unknown token |
<BOS> |
2 | Beginning of sequence |
<EOS> |
3 | End of sequence β model stops here |
Attention(Q, K, V) = softmax(QK^T / sqrt(d_k)) * V
| Symbol | Role |
|---|---|
| Q (Query) | "What am I looking for?" |
| K (Key) | "What information do I have?" |
| V (Value) | "What do I actually return?" |
- Division by
sqrt(d_k)prevents softmax from saturating in high dimensions - Causal mask: future tokens get
-infβ 0 probability (autoregressive) - Multiple heads: each head learns different relationship types (syntax, semantics, coreferenceβ¦)
FFN(x) = GELU(xWβ + bβ)Wβ + bβ
Applied position-wise. Acts like "memory" β stores factual associations.
| Component | What it does |
|---|---|
Pre-norm (x + sublayer(LayerNorm(x))) |
More stable gradients than post-norm (GPT-2 style) |
Residual connections (x = x + sublayer(x)) |
Prevent vanishing gradients in deep networks |
| Weight tying | Embedding matrix and LM head share weights β fewer params |
Objective: Causal Language Modeling β given [t1, t2, t3, t4], predict [t2, t3, t4, t5].
Warmup β lr = max_lr Γ (iter / warmup_iters)
Cosine β lr = min_lr + (max_lr - min_lr) Γ 0.5 Γ (1 + cos(Ο Γ progress))
torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)Prevents sudden loss spikes by scaling down large gradients.
Simulates larger batches without extra memory:
for micro_step in range(grad_accumulation_steps):
loss = model(batch) / grad_accumulation_steps
loss.backward()
optimizer.step() # Only step once per "logical" batch| Format | Bits | Range | Precision | Benefit |
|---|---|---|---|---|
| float32 | 32 | Β±3.4Γ10Β³βΈ | ~7 digits | Training stable |
| bfloat16 | 16 | Β±3.4Γ10Β³βΈ | ~3 digits | 2-4Γ faster, ~50% less VRAM |
bfloat16 > float16 for training β same dynamic range, no loss scaling needed.
GPU 0: model copy β forward(batch_shard_0) β backward β gradients ββ
GPU 1: model copy β forward(batch_shard_1) β backward β gradients ββ€
β
All-Reduce (NCCL): avg gradients
β
Both GPUs update weights identically
torchrun sets these environment variables automatically:
| Variable | Meaning |
|---|---|
RANK |
Global process index (0 = master) |
LOCAL_RANK |
GPU index on this node |
WORLD_SIZE |
Total number of processes |
Autoregressive loop: feed tokens β sample next token β append β repeat β stop at <EOS>
| Strategy | How | Effect |
|---|---|---|
| Greedy (temp=0) | Always pick argmax | Deterministic, can be repetitive |
Temperature T < 1 |
Sharpen distribution | More confident, less creative |
Temperature T > 1 |
Flatten distribution | More random, more creative |
| Top-k | Sample from top-k tokens only | Blocks very unlikely tokens |
| Top-p (nucleus) | Sample from smallest set with cumulative prob β₯ p | Adaptive to model certainty |
At this scale, throwing generic text at the model isn't enough β both what the shared tokenizer sees and what skill each stage is asked to learn have to be deliberately scoped down to something learnable at 10Mβ65M parameters. TinyLLM mixes real corpora with purpose-built synthetic data for exactly that reason.
python -m data_pipeline.prepare_pipeline_data --force_retrain_tokenizer # first run, or to fold in a new corpusDownloads real datasets, builds one shared BPE tokenizer used by every later stage, and writes the per-stage files everything else reads from:
| Source | Used for | Output |
|---|---|---|
| WikiText-2 (real Wikipedia prose) | Stage 1 pretraining β Alpaca/GSM8K text alone has essentially no world-knowledge exposure | data/raw_text/corpus.txt |
| Alpaca + Dolly-15k | Stage 2 instruction SFT | data/sft_dataset.json |
| GSM8K (train/test) | Stage 3 reasoning SFT (real chain-of-thought math) | data/reasoning_dataset.json |
Re-running is safe and cheap β downloads are skipped if already present, and the tokenizer is reused (not retrained) by default so changing the SFT mix later doesn't invalidate an existing pretrained checkpoint's vocabulary.
Real datasets like GSM8K are open-domain and hard for a 10M-param model to make progress on quickly. TinyLLM also ships generators for narrower, controlled tasks it can actually learn:
generate_synthetic_reasoning.pyβdata/synthetic_reasoning_all_hops.jsonβ templated 1-, 2-, and 3-hop arithmetic word problems. The<CALC>tool already guarantees correct arithmetic (see Limitations), so this isolates the actual hard part: correctly reading which numbers and operation a problem calls for.generate_smalltalk_multiturn.py+build_smalltalk_demo_dataset.pyβdata/smalltalk_multiturn.json/data/smalltalk_demo.jsonβ multi-turn small-talk conversations built by cross-combining independent "opener" and "follow-up" pools, so every opener is followed by many different follow-ups. This forces the model to actually read turn 2 instead of memorizing one fixed script per opener.
Both feed train_reasoning.py --dataset_format chatml, which pools chat and reasoning
conversations as equals through data_utils.prepare_multitask_data β see the
Chat UI section above for what that checkpoint can do.
You can also train TinyLLM on your own question-answer data using a simple JSON format.
[
{
"id": 1,
"category": "Identity",
"question": "Who are you?",
"answer": "I am TinyLLM, a small but capable language model here to help you!"
},
{
"id": 2,
"category": "Identity",
"question": "What is your name?",
"answer": "My name is TinyLLM."
}
]| Field | Required | Description |
|---|---|---|
id |
No | Unique identifier (ignored during training) |
category |
No | Grouping label (ignored during training) |
question |
β Yes | The input question text |
answer |
β Yes | The expected answer text |
Each Q&A pair is automatically wrapped with boundary tokens before training:
<BOS> Question: Who are you?
Answer: I am TinyLLM, a small but capable language model here to help you! <EOS>
This teaches the model where answers end β without <EOS> boundaries, the model would answer a question and then immediately ask itself another one and keep going indefinitely.
from data_utils import prepare_custom_data, create_dataloader
train_ds, val_ds, tokenizer = prepare_custom_data(
json_path="data/my_dataset.json",
vocab_size=2000,
context_length=128,
force_retrain_tokenizer=True, # retrain so BOS/EOS appear in the corpus
)
loader = create_dataloader(train_ds, batch_size=8)Your generation loop must honour the <EOS> token:
for _ in range(max_new_tokens):
logits = model(input_ids)
next_token = logits[:, -1, :].argmax(dim=-1)
if next_token.item() == tokenizer.eos_id:
break # β stop here!
input_ids = torch.cat([input_ids, next_token.unsqueeze(0)], dim=1)pip install torch --index-url https://download.pytorch.org/whl/cu121python -m training.traintorchrun --nproc_per_node=2 -m training.traintorchrun --nproc_per_node=2 -m training.train \
--batch_size 32 \
--max_iters 5000 \
--d_model 512 \
--n_heads 8 \
--n_layers 8d_model must be divisible by n_heads (model/config.py validates this) β the
default n_heads is 6, which doesn't divide 512, so raising d_model means picking
a compatible n_heads too.
Both train.py (SFT) and pretrain.py support optional W&B logging of train/val
loss, perplexity, learning rate, tokens/sec, and gradients β off by default, opt in with --use_wandb.
pip install wandb
wandb login
python -m training.train --use_wandb --wandb_project tinyllm-sft --wandb_run_name my-run
python -m training.pretrain --use_wandb --wandb_project tinyllm-pretrainpython -m demo.inference \
--checkpoint checkpoints/latest.pt \
--prompt "Who are you?" \
--temperature 0.8 \
--top_k 50 \
--max_tokens 300TinyLLM is a learning tool, not a production model β knowing where it falls short is part of understanding how it works:
- No world knowledge. The default config is ~10M parameters (bump
d_model/n_layersvia CLI flags and it scales to tens of millions β seemodel/config.py), trained on a small corpus. That's nowhere near enough capacity or data to have memorized facts the way GPT-3/4 does. It will confidently make things up (hallucinate) outside what it was trained on. - Short context window.
context_lengthis 256 tokens by default β a few paragraphs. Long documents or long conversations will simply fall off the front of the window. Compare to production LLMs, which use 32Kβ1M+ token windows. - Small vocabulary.
vocab_size=10000(vs. ~100K+ for GPT-4-class tokenizers) means more tokens per word on average, especially for rare words or non-English text. - The
<CALC>tool only does arithmetic.model/calculator.pysupports+ - * /and parentheses β no algebra, no unit conversion, no comparisons. It's real (not guessed) arithmetic, but the model still has to correctly decide which numbers and operation to plug in, which is the actual hard part (seegenerate_synthetic_reasoning.py). - Reasoning is on synthetic, templated word problems, not open-domain math (like GSM8K) or general reasoning. This narrows the skill on purpose so it's learnable at this scale β it does not mean the model can reason broadly.
- Demo examples are held-out but few. The
<CALC>reasoning examples shown in the demo are unseen during training, but the held-out set is small β treat the demo as a qualitative illustration, not a statistically rigorous benchmark result. - "Multi-turn" here means multiple exchanges, not real conversational memory.
data/smalltalk_multiturn.jsonis built by cross-combining independent, self-contained turns (seegenerate_smalltalk_multiturn.py's docstring) specifically so each turn is answerable on its own β none of them require the model to actually recall or reuse something from an earlier turn. Teaching a model to genuinely track state across a conversation (remember a name, resolve "it"/"that" back to something said earlier, follow a multi-step task) takes large volumes of real or carefully human-curated dialogue data β the kind of thing OpenAI/Anthropic/Google collect from actual product usage or pay annotators to write at scale. That's a resource gap a small/independent project can't synthesize its way around with templates alone, so it's out of scope here rather than something this repo currently attempts.
Once the base model is training, try these:
| # | Experiment | Where |
|---|---|---|
| 1 | Sinusoidal vs learned positional embeddings | --use_learned_pos_emb False |
| 2 | Swap GELU for ReLU or SiLU | model.py FFN block |
| 3 | RoPE positional encoding (used in LLaMA) | Add to model.py |
| 4 | Flash Attention (drop-in, better memory) | Replace MultiHeadAttention |
| 5 | Different datasets β Bible, Project Gutenberg | data_utils.py |
| 6 | Scaling laws β train on 10Γ more data | Watch val loss curve |
| 7 | SwiGLU activation (SwiGLU(x) = (xW+b) Γ Ο(xV+c)) |
model.py FFN block |
| Paper | Authors | Why read it |
|---|---|---|
| Attention Is All You Need | Vaswani et al., 2017 | Original Transformer |
| Language Models are Unsupervised Multitask Learners | Radford et al., 2019 | GPT-2 |
| An Image is Worth 16Γ16 Words | Dosovitskiy et al., 2020 | ViT β shows transformers work everywhere |
| Training Compute-Optimal LLMs | Hoffmann et al., 2022 | Chinchilla scaling laws |
