amitshekhariitbhu/ai-system-design
AI System Design - Learn how to design AI systems built on LLMs, RAG, and AI Agents step by step.
About amitshekhariitbhu/ai-system-design
amitshekhariitbhu/ai-system-design is an open-source project on GitHub, mainly written in Markdown. AI System Design - Learn how to design AI systems built on LLMs, RAG, and AI Agents step by step. It currently holds 514 stars and 64 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).
Project Overview
AI Homed tracks it on the Today's Trending board.
GitHub Repository Details
README
AI System Design
AI System Design - A complete guide to learn AI System Design step by step - from LLM inference, GPUs, KV Cache, and caching to RAG, Vector Databases, AI Agents, MCP, Multi-Agent Systems, Voice AI, Guardrails, Evaluation, Observability, Cost Optimization, and a step-by-step framework to crack any AI System Design interview. Everything in one place, explained in simple words, with detailed blogs for every deep dive.
This AI System Design guide is helpful for anyone who wants to become:
> - AI Engineer
- Gen AI Engineer
- LLM Engineer
- Agentic AI Engineer
- AI Agent Engineer
- Forward Deployed Engineer
- AI Solutions Architect
- AI Platform Engineer
- Applied AI Engineer
- Machine Learning Engineer
- MLOps Engineer
- LLMOps Engineer
- Backend Engineer building AI products
---
Prepared and maintained by the Founder of Outcome School: Amit Shekhar
Follow Amit Shekhar
Follow Outcome School
I teach at Outcome School
---Note: AI System Design is moving very fast, so this guide will continue to grow as I write more blogs on new topics. Bookmark it and come back whenever you want a refresher. Keep learning.
---
Table of Contents
- About This AI System Design Guide
- What is AI System Design?
- Who is This AI System Design Guide For?
- What Will We Learn in This AI System Design Guide?
- How to Use This AI System Design Guide
- AI System Design Learning Path
- Why study AI System Design?
- How is AI System Design different from regular System Design?
- LLM Recap
- The Big Picture
- Inference Server
- Choosing an Inference Engine
- AI Hardware: GPU, TPU, and LPU
- Tokens, Latency, and Throughput
- Time to First Token (TTFT)
- Tokens per Second (TPS)
- Throughput
- Cost per Token
- Prefill and Decode: The Two Phases of LLM Inference
- Prefill
- Decode
- Chunked Prefill
- Prefill-Decode Disaggregation
- Scaling AI Systems
- Vertical Scaling
- Horizontal Scaling
- Tensor Parallelism
- Pipeline Parallelism
- How to choose
- Auto Scaling for AI
- Back-of-the-envelope Estimation for AI
- Token Estimation
- GPU Estimation
- Cost Estimation
- Load Balancing for LLM Servers
- Least Outstanding Tokens
- Prefix-Aware Routing
- Sticky Sessions
- Caching in AI
- KV Cache
- KV Cache Compression
- Prompt Cache
- Semantic Cache
- Embedding Cache
- LLM Routing (Model Routing)
- Vector Database
- Vector Index Types
- RAG (Retrieval Augmented Generation)
- Document Parsing and Ingestion
- Chunking Strategies
- Hybrid Search
- Query Transformation with HyDE
- Reranking
- ColBERT and Late Interaction
- Agentic RAG
- GraphRAG
- Vectorless RAG
- Context Window Management
- Context Rot, Lost in the Middle, and RoPE Decay
- Context Engineering
- Truncation
- Summarization
- Sliding Window
- Compaction
- Hierarchical Memory
- Streaming Responses
- Server-Sent Events (SSE)
- WebSockets
- Async Processing for Long AI Tasks
- Message Queues in AI Systems
- Rate Limiting in AI
- Tokens Per Minute (TPM)
- Requests Per Minute (RPM)
- Concurrent Requests
- Cost-Based Rate Limiting
- AI Gateway
- Embeddings Pipeline
- Choosing an Embedding Model
- Matryoshka Embeddings
- AI Agents and Agentic Systems
- The Five Core Parts
- How an AI Agent Works End to End
- Types of AI Agents
- Computer Use and Browser Agents
- Common Failure Modes
- AI Orchestration vs AI Agents
- Loop Engineering
- Graph Engineering
- Tool Calling
- Model Context Protocol (MCP)
- How MCP works
- Why MCP matters for AI System Design
- When to use MCP
- When MCP is overkill
- Agent Skills
- Structured Output
- Memory for AI Agents
- The Memory Stack
- The Four Core Operations
- How Memory Flows at Runtime
- What to Store and What Not to Store
- Multi-Agent Systems
- The Three Pillars
- Common Agent Roles
- A Concrete Example: Customer Support
- Coordination Patterns
- AI SubAgents
- Trade-offs
- A2A (Agent2Agent Protocol)
- How A2A works
- How MCP and A2A fit together
- Multimodal Systems
- Storage
- Pre-processing Pipeline
- Token Cost
- Output Modalities
- Latency
- Voice and Realtime APIs
- Edge AI and On-Device Inference
- Guardrails and Safety
- Input Guardrails
- Output Guardrails
- Prompt Injection
- AI Red Teaming
- LLM Watermarking
- Data Privacy and Compliance
- PII Redaction
- Data Residency
- No-Train and BAA Clauses
- Voice and Multimodal Privacy
- Audit Logs
- Observability in AI Systems
- Evaluation Pipeline
- LLM as a Judge
- Evaluating AI Agents
- Prompt Management
- Programmatic Prompting with DSPy
- Cost Optimization
- 1. Use a cheaper model when possible
- 2. Use prompt caching
- 3. Use semantic cache
- 4. Shorter outputs
- 5. Self-host smaller models
- 6. Batch inference
- 7. Better retrieval (for RAG)
- Multi-Tenancy
- Fine-Tuning Infrastructure
- Training Cluster
- Training Data Pipeline
- Experiment Tracking
- Model Registry
- Evaluation
- Inference Optimization
- Quantization
- Continuous Batching
- Speculative Decoding
- Test-time Compute (Inference-time Scaling)
- Flash Attention
- Mixture of Experts (MoE)
- Model Distillation
- Grouped Query Attention (GQA)
- Fault Tolerance
- Timeouts
- Retries with Backoff
- Fallback Models
- Graceful Degradation
- Output Validation
- How to Solve Any AI System Design Problem
- Step 1: Requirements
- Step 2: AI Objective
- Step 3: Data Preparation
- Step 4: Architecture Design
- Step 5: Model Selection and Prompting
- Step 6: Evaluation
- Step 7: Deployment and Serving
- Step 8: Monitoring
- Real-World AI System Case Studies
- Case Study 1: How Claude Code Works
- Case Study 2: How Cursor Works
- Case Study 3: Design a Real-Time Voice AI Agent
- AI System Design Interview Questions
- Quick Summary
- AI System Design Key Concepts Glossary
- AI System Design FAQs
- License
About This AI System Design Guide
In this guide, we will learn about AI System Design, the discipline of putting GPUs, inference servers, caches, vector databases, AI agents, gateways, guardrails, and evals together into one system that is fast, cheap, reliable, and safe. We will also see how an LLM actually runs on a GPU, how prefill and decode shape latency, how caching, routing, and batching cut the cost, how RAG and AI Agents are built for production, how we keep the system safe and measurable, and a step-by-step framework to solve any AI System Design problem in an interview or in real production.
When we use a product like ChatGPT, Cursor, Perplexity, or Claude Code, we see a simple chat box and a streamed response. Behind that simple interface, there is a lot more happening - GPUs, inference servers, vector databases, agent loops, caches, gateways, guardrails, and a long list of design decisions that all have to work together.
This guide is everything we need in one place. We start with the basics like tokens and the inference server, build up through hardware, prefill and decode, scaling, caching, RAG, agentic systems, multi-agent systems, multimodal and voice systems, safety, observability, evaluation, and inference optimization, and finish with a step-by-step framework to solve any AI System Design problem.
Every section is:
- Written for beginners. No jargon. No assumptions. Every term is explained before it is used.
- Practical. Real numbers, real trade-offs, real tools, and real architecture diagrams.
- Connected to a deep dive. Wherever a topic deserves more depth, we link to a detailed blog that explains it from the ground up.
What is AI System Design?
AI System Design is the discipline of designing the complete system around an AI model, especially a Large Language Model (LLM), so that it can serve real users in a way that is fast, cheap, reliable, safe, and measurable.
In simple words:
AI System Design = System Design + The new constraints of AI models.
The new constraints are GPUs, tokens, long and streamed responses, non-deterministic output, and a real cost on every single request. AI System Design is how we design caches, queues, databases, gateways, retrieval, agents, guardrails, and evals around these constraints.
Let's say we want to build a customer support chatbot. Calling an LLM API in a script takes ten lines of code. But serving 100,000 users with an answer that starts in under a second, stays grounded in our own documents, never leaks private data, and fits within a monthly budget - that is AI System Design.
Who is This AI System Design Guide For?
This AI System Design guide is for:
- Software Engineers who want to move into AI Engineering.
- Backend, Mobile, and Frontend Developers who want to build AI-powered products.
- Machine Learning Engineers and Data Scientists who want to take models to production.
- Engineering Managers, Tech Leads, and Architects who want to understand how modern AI systems are built.
- Students and freshers who want to start a career in AI.
- Anyone preparing for AI System Design interviews, AI Engineer interviews, and GenAI Engineer interviews.
What Will We Learn in This AI System Design Guide?
In this AI System Design guide, we will learn:
- The foundations: how AI System Design differs from regular System Design, tokens, the inference server, choosing an inference engine, and the hardware (GPU, TPU, LPU).
- LLM inference: prefill vs decode, TTFT, TPOT, throughput, chunked prefill, and prefill-decode disaggregation.
- Scaling: vertical and horizontal scaling, tensor parallelism, pipeline parallelism, auto scaling with warm pools, back-of-the-envelope estimation, and load balancing for LLM servers.
- Caching: KV Cache, Paged Attention, KV Cache Compression, Prompt Cache, Semantic Cache, and Embedding Cache.
- LLM Routing: rule-based, classifier-based, embedding-based, LLM-as-router, and cascade routing.
- Retrieval: embeddings, vector databases, vector indexes, RAG, document parsing, chunking, hybrid search, HyDE, reranking, ColBERT, Agentic RAG, GraphRAG, and Vectorless RAG.
- Context: context window management, context rot, lost in the middle, context engineering, and context compaction.
- Serving patterns: token streaming, async processing, message queues, rate limiting, and the AI Gateway.
- AI Agents: the five core parts, the agent loop, the harness, AI Orchestration vs AI Agents, Loop Engineering, Graph Engineering, ReAct, Plan-and-Execute, Reflection, and computer-use agents.
- Tools and knowledge: tool calling, MCP, Agent Skills, structured output, and agent memory.
- Multi-Agent Systems: the three pillars, agent roles, coordination patterns, SubAgents, and A2A.
- Multimodal and Voice AI: voice AI agents, the latency budget, barge-in, cloud vs on-device deployment, and edge AI.
- Safety: guardrails, prompt injection, AI red teaming, LLM watermarking, and data privacy and compliance.
- Quality: observability with traces and spans, LLM evaluation, LLM as a Judge, and AI Agent evaluation.
- Operations: prompt management, DSPy, cost optimization, multi-tenancy, fine-tuning infrastructure, and fault tolerance.
- Inference optimization: quantization, continuous batching, speculative decoding, test-time compute, Flash Attention, Mixture of Experts, distillation, and Grouped Query Attention.
- Interviews: a step-by-step framework to solve any AI System Design problem, real-world case studies, and common AI System Design interview questions.
How to Use This AI System Design Guide
- If we are new to AI System Design, we read it from top to bottom. Each section builds on top of the previous one.
- If we are preparing for an interview, we read How to Solve Any AI System Design Problem first, and then come back to the building blocks.
- If we want to go deep into any topic, we open the linked blog. Every linked blog explains one concept from the ground up.
- After every section, we try to explain it to a friend in our own words. If we can explain it, we have learned it.
AI System Design Learning Path
flowchart TD
A[Foundations: Tokens, Inference Server, Hardware] --> B[LLM Inference: Prefill and Decode]
B --> C[Scaling, Estimation, and Load Balancing]
C --> D[Caching and LLM Routing]
D --> E[Embeddings, Vector Databases, and RAG]
E --> F[Context Window Management]
F --> G[Streaming, Queues, Rate Limiting, AI Gateway]
G --> H[AI Agents, Tools, MCP, and Memory]
H --> I[Multi-Agent Systems]
I --> J[Multimodal and Voice AI]
J --> K[Guardrails, Safety, and Privacy]
K --> L[Observability and Evaluation]
L --> M[Cost, Fine-Tuning, and Inference Optimization]
M --> N[How to Solve Any AI System Design Problem]
I am Amit Shekhar, Founder @ Outcome School, I have taught and mentored many developers, and their efforts landed them high-paying tech jobs, helped many tech companies in solving their unique problems, and created many open-source libraries being used by top companies. I am passionate about sharing knowledge through open-source, blogs, and videos.
I teach AI and Machine Learning at Outcome School.
Let's get started.
Why study AI System Design?
When most of us start with AI, we just pick an OpenAI or Anthropic API key, write a small script that calls the API, and get a response back. We feel that our AI app is built.
For a small project or a personal demo, this is good enough.
But the real world is very different.
In the real world, an AI product serves millions of users. Each user sends long prompts. Each response is streamed token by token. Models are slow. GPUs are expensive. Costs add up fast. Hallucinations creep in. Latency matters.
A single API call cannot handle all of this.
To make our AI product reliable, fast, cheap, and safe, we have to think about many things together. This is where AI System Design comes into the picture.
How is AI System Design different from regular System Design?
In regular system design, we deal with CPU, RAM, disk, databases, and network.
In AI System Design, we deal with all of these, and on top of that, we deal with:
- GPUs: LLMs run on GPUs, not CPUs. GPUs are expensive and limited.
- Tokens: LLMs do not work with characters. They work with tokens. Both input and output are billed per token.
- Long requests: A single LLM call can take 30 seconds or even minutes. Regular APIs return in milliseconds.
- Streaming: LLM responses are streamed token by token. Not a single big response.
- Non-deterministic output: The same input can give different outputs. We cannot just unit-test like regular code.
- Cost per request: Every API call costs real money. A bug in a loop can burn thousands of dollars overnight.
LLM Recap
Before we go into AI System Design, let's quickly recap what an LLM is.
LLM stands for Large Language Model. It is a model that takes some text as input and predicts the next token. It does this again and again until the full response is generated.
The text is not directly given to the model. It is first broken into smaller pieces called tokens. Think of it like a chocolate bar. The full bar is the sentence. Each small part we break off is a token. The model processes multiple tokens at a time. Most modern LLMs use BPE (Byte Pair Encoding) to do this tokenization. One token is roughly 4 characters in English. So short common words like "AI" or "Hello" are one token. Longer words like "Bangalore" are 2 tokens. Compound names like "ChatGPT" are 2 to 3 tokens depending on the tokenizer.
LLMs generate text one token at a time. This is called autoregressive generation, which means the model keeps feeding its own output back as the new input. Let's say we give the model this input:
"I love"
The model looks at "I" and "love", and predicts the next token: "teaching". The full sequence becomes "I love teaching". The model now looks at "I", "love", and "teaching" and predicts the next token: "AI". The full sequence becomes "I love teaching AI". This process continues, one token at a time, until the model decides to stop.
At every step, the model does not know the next token for sure. It gives a probability to every possible token, and then one token is picked. Settings like Temperature and Top-k and Top-p Sampling control how this pick happens. This is also why the same prompt can give different outputs, which is one of the biggest reasons AI System Design is different from regular System Design.
Internally, the LLM is a Transformer - a stack of attention and feed-forward layers. The single most important idea inside it is the attention mechanism, where each token converts itself into three vectors - Query (Q), Key (K), and Value (V) - and uses them to figure out which previous tokens matter most for predicting the next one. We have a detailed blog on the math behind Attention - Q, K, and V that goes into the math step by step.
When we use an API like OpenAI, Anthropic, or Google, we send a prompt, and we get a streamed response back, one token at a time.
Examples of LLMs: GPT-5.5, Claude Opus 4.8, Gemini 3.5, Llama 4, Mistral.
If we wan