Build lab

Experiment 01 · Models from scratch

Did my tiny AI learn—or just memorize?

Build a small text model, watch its output change, and design a test that training examples cannot answer for it.

Free project brief · Video planned · Python foundations

The experiment

Build a deliberately small text generator and compare it with a simple baseline. The goal is to make learning visible, then investigate what the result actually proves. This is an experiment brief to build along with; a recorded walkthrough and measured results have not been published yet.

Keep the first version small

Use a public or synthetic text collection you have permission to use. Begin with a character-frequency or character-pair baseline. Save a fixed training split and a separate test split before changing the model. A baseline is useful even if you later replace it with a small neural network.

Build a simple interface with three panels: the input data, generated samples, and evaluation results. Keep the seed, training settings and dataset version next to each run.

Run three comparisons

  1. Before learning: save output from the initial system with a fixed seed.
  2. After learning: generate samples using the same settings and record what changed.
  3. On unseen examples: evaluate held-out text and check how much generated output repeats training passages.

Include a deliberately tiny dataset to make memorization easier to observe. Compare it with a larger, more varied dataset. Keep the evaluation prompts separate from the examples used to choose settings.

Make the story visible

Open with the initial output. Show one working change at a time. When output looks impressive, pause and test it. End with the evidence: what improved, what repeated, and what still fails. A good-looking sentence is a sample, not a complete evaluation.

What to publish

  • A README with the setup, data source and exact run command.
  • A baseline and a held-out evaluation split.
  • Example outputs with seeds and settings.
  • A short results table and at least one failure case.
  • A clear distinction between a frequency baseline and any neural model you build.

Get the foundations first

Start with the engineering foundations, then build a measurable AI feature. Use the evaluation plan to decide what counts as progress before tuning anything.