← All projects

Deep learning

Tiny Transformer

A small language-model training setup for inspecting attention, token predictions, and the learning process.

GitHub (placeholder) (opens in a new tab)Example project

The link opens my GitHub profile until a project repository is added.

Tiny Transformer concept: token sequences alongside a triangular causal attention grid.
Concept illustration · example project

The problem

High-level training libraries can hide the operations that make a language model work. This example project reduces the scope to a small decoder so tensor shapes, masking, and the loss are easy to inspect.

How it would work

Begin with a bigram baseline on a small, openly licensed text corpus. Define train and validation splits at document boundaries before making token windows.

Implement embeddings, causal self-attention, a position-wise feed-forward network, and residual connections. Log tensor shapes and check that a token cannot attend to future positions.

Save the tokenizer, configuration, random seed, and checkpoint together. A short generation script would compare temperatures using the same prompt, while an attention view would help inspect one layer at a time.

What to evaluate

A proposed evaluation plan for this example:

  • Overfit one tiny batch to check the training loop before a larger run.
  • Compare validation cross-entropy against the bigram baseline using the same tokenizer and split.
  • Track training and validation loss together; inspect samples without treating them as a benchmark.
← Back to all projects