Build the core skills · 06 / 15
Build one LLM feature you can measure
Add a suggested category and short summary to your ticket intake service. Treat the model response as untrusted data that must satisfy your application’s schema. Keep the original ticket so the suggestion can be corrected.
Understand these six ideas
- Token: a piece of text used by the model; token usage affects context and often billing.
- Context: the input available during a request, with a finite limit.
- Inference: using a trained model to generate a result.
- Embedding: a numeric representation useful for comparing meaning; similarity does not prove truth.
- Structured output: a response format your application validates, such as a fixed JSON schema.
- Evaluation: a repeatable check of whether the system handles a defined task well.
Build in small steps
- Start with 10 representative, non-sensitive tickets you wrote yourself. Assign expected categories before calling the model.
- Make a keyword baseline. It gives you something concrete to improve on.
- Add one model call behind a function. Keep the key on the server in an environment variable, outside the repository.
- Request a category from a fixed allowed set and a short summary. Validate the response; handle refusal, invalid output and timeout explicitly.
- Log a request ID, timing, model/prompt version and usage counters. Avoid retaining unnecessary user text.
- Run the same examples after changing the prompt. Record both successes and failures.
Watch, then build
Full Stack LLM Bootcamp: LLM Foundations ↗ gives the conceptual background. Its 2023 examples are historical; use the current provider documentation for API details. The optional DeepLearning.AI: Building Systems with the ChatGPT API ↗ lab can provide more structure, but you can complete this exercise with the free lecture and your provider’s documentation.
My suggested practice budget is a small, explicit ceiling you choose before making live calls. Mock responses during development, limit retries and output length, and inspect the provider’s current billing controls. No paid course or GPU purchase is necessary to start.
Your tests exercise a correct result, invalid output, refusal, timeout and unavailable provider. You can explain the baseline comparison, cost estimate and one failure that remains.
Save the first 10 examples. They become the seed of your evaluation set.
Your lesson resources
Download these files to follow along and put the lesson into practice.