Deep learning
Tiny Transformer
A small language-model training setup for inspecting attention, token predictions, and the learning process.
The link opens my GitHub profile until a project repository is added.
The problem
High-level training libraries can hide the operations that make a language model work. This example project reduces the scope to a small decoder so tensor shapes, masking, and the loss are easy to inspect.
How it would work
Begin with a bigram baseline on a small, openly licensed text corpus. Define train and validation splits at document boundaries before making token windows.
Implement embeddings, causal self-attention, a position-wise feed-forward network, and residual connections. Log tensor shapes and check that a token cannot attend to future positions.
Save the tokenizer, configuration, random seed, and checkpoint together. A short generation script would compare temperatures using the same prompt, while an attention view would help inspect one layer at a time.
What to evaluate
A proposed evaluation plan for this example:
- Overfit one tiny batch to check the training loop before a larger run.
- Compare validation cross-entropy against the bigram baseline using the same tokenizer and split.
- Track training and validation loss together; inspect samples without treating them as a benchmark.