Resume Project Steps · 05 / 7
4. AI System Architecture
You are going to design an actual AI system with real engineering decisions. This is what separates a top 1% candidate from the other 99%.
Phase 4: AI System Architecture (quick-start)
This is the fast, do-it-now version of the architecture step. This is where you design the whole system on paper and make the engineering decisions that separate a top one percent candidate from everyone else. You are not building yet. You are building it mentally first and locking in the model. You build it with your team in the next section. Use the attached doc for the full process.
The core idea before you start: the difference between a good project and a top one percent project is not the code, it is the decisions. This phase is where you make them and learn to defend them.
Step 1: Frame the problem as a system
Before any AI decisions, define the boundaries. Answer: who is the user and what job are they doing, what are the inputs and outputs, what does success concretely look like, what are the constraints (latency, cost, data, privacy, scale), and what is explicitly out of scope. Write a one-paragraph problem statement and a short list of functional and non-functional requirements.
Step 2: Decide whether you even need AI, and how much
Knowing when not to reach for an LLM is a senior signal. Use AI only where it adds real value (language, generation, reasoning over messy data), and keep the rest of the system simple. Reaching for AI everywhere is a junior tell.
Step 3: Choose your core AI pattern
Pick the pattern that fits your problem, and be ready to explain why it beats the alternatives:
Direct prompting: task is self-contained and the model already knows enough.
RAG (retrieval-augmented generation): the answer depends on specific, private, or current data. Retrieve first, then generate grounded answers.
Agents and tool use: the task needs the model to take actions or reason across steps. Powerful but harder to make reliable.
Fine-tuning: you need a specific behavior prompting cannot give and you have data. Usually not the first move.
Combinations: real systems mix these. Route, chain, and ground as needed.
Step 4: Make the model and infrastructure decisions
Model: hosted API (OpenAI, Anthropic, Google) vs open-weight self-hosted (Llama, Mistral). Weigh capability, latency, cost, privacy, and control. Bonus maturity: use a cheap model for simple steps and a strong one for hard steps.
Data and retrieval (if RAG): what data, ingestion, chunking, embeddings, vector store, hybrid search, reranking, context assembly. Chunking and retrieval quality make or break it.
Surrounding stack: vector store, framework or none, backend, frontend, database, deployment. Every tool earns its place. Do not cargo-cult a framework you do not need.
Step 5: Design for reliability, cost, and evaluation (this is the top 1% part)
Almost no portfolio project thinks about this, which is exactly why it makes you stand out.
Evaluation, from the start: build a small eval set, define metrics (retrieval quality, answer correctness and faithfulness, task success), and decide how you measure them (human review, model-as-judge, automated checks). Having an eval plan at all beats most candidates.
Failure modes: hallucination, wrong tool calls, malformed output, latency, rate limits, downtime. Ground with retrieval, validate and structure outputs, add retries and fallbacks, and keep a human in the loop where stakes are high.
Cost and latency: cache, right-size models, limit context, stream responses. Know what a request costs and where it is slow.
Security and privacy: data handling, PII, prompt injection, no leaking sensitive data into prompts or logs. This matters more in regulated industries.
Step 6: Draw the architecture diagram
One clear diagram showing the user, the app and backend, the AI components and pattern (retrieval, model calls, tools), the data stores, external services, and the path a request takes end to end. Legible beats fancy. Use Excalidraw, draw.io, or Mermaid.
Step 7: Write the design doc
Write it up like a real engineering design doc: problem and goals, requirements, architecture overview, key decisions and tradeoffs, evaluation plan, risks and failure modes, cost and scalability, and out of scope. For every major decision, document it as: Decision, Options considered, Choice, Why, Tradeoffs. That decision log is your interview script, written down in advance.
Step 8: Validate before you build
Trace a real user scenario through the whole system. Confirm it is buildable in one to two months. Test your single riskiest assumption with a tiny throwaway experiment (do not start the real build). Make sure you can explain every decision out loud, and get a second set of eyes if you can.
What you should have when you are done
A crisp problem statement and requirements.
A chosen AI pattern and model, with tradeoffs written down.
A data/retrieval or agent design, an evaluation plan, and a reliability/cost/failure-mode plan.
A clear architecture diagram and a written design doc with a decision log.
The ability to walk someone through the whole system and defend every decision.
Do not move on until you can do that last one without notes.
Avoid these
Reaching for AI where something simpler is better.
Jumping to a framework or tool before understanding the problem.
Skipping evaluation. This is the biggest miss.
Ignoring cost, latency, and failure modes.
Designing something you cannot build in one to two months.
Copying a tutorial architecture you cannot justify.
Not writing it down, or starting to build before the design is clear.
Go deeper: the full doc
This is the quick-start. The attached document breaks down each pattern and when to use it, the model and retrieval decisions in detail, the full evaluation and reliability thinking, and the design doc and decision-log templates. Use it whenever you want the depth.
Your lesson resources
Download these files to follow along and put the lesson into practice.