The experiment
Take a small app with an existing test suite. Compare how a coding assistant performs with no project memory and with a concise memory written after earlier tasks. This is a build brief; there are no published results or finished episode yet.
Define useful memory
Create a project note with five sections: how to run the app, important design decisions, known failures, commands that verify the work, and the next open task. Include file references and dates. Keep credentials, private customer data and speculative claims out of the note.
Every entry should answer a future question. Prefer “the API expects UTC timestamps; see the date conversion test” to a long transcript of the previous conversation.
Make the comparison fair
- Choose three small tasks with acceptance criteria and hidden checks.
- Start each trial from the same repository commit.
- Use the same model, tool access and budget in both conditions.
- Give one condition the project note and the other the normal setup documentation.
- Repeat enough trials to notice variation instead of declaring a winner after one result.
Record task completion, regressions, repeated mistakes, time and tool usage. Review the patches yourself. If the model changes during the experiment, record that and keep the comparisons separate.
Test when memory becomes wrong
Change one project assumption in a controlled branch. Does the assistant verify the note against the code, or follow a stale instruction? Update the note after discovering the mismatch. A memory that cannot be corrected becomes another source of errors.
What to publish
Show the task, the two patches, their test results, and the exact memory used. Explain the limits of the comparison. Do not present one successful trial as proof that memory always helps.
Continue with bounded tools and agent workflows, or use the skills library to explore reusable project instructions.