harshuljain13/llm-inference-at-scale
A Practitioner handbook for production llm serving.
About harshuljain13/llm-inference-at-scale
harshuljain13/llm-inference-at-scale is an open-source project on GitHub, mainly written in Jupyter Notebook. A Practitioner handbook for production llm serving. It currently holds 238 stars and 43 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).
Project Overview
AI Homed tracks it on the Today's Trending board, currently at rank #40 with 19 new stars today.
GitHub Repository Details
README
The definitive guide to serving large language models in production.
Quick Start • Contents • Equations • Contributing
---
Why This Exists
Serving an LLM is not like serving a traditional ML model. With traditional ML, you send a request, the model runs one forward pass, and you get a result. Fixed time. Fixed memory. Simple.
With LLMs, every request is different. A short reply takes 100ms. A long one takes 10 seconds. The model generates one word at a time, and each word requires reading the entire model from memory again. The longer the conversation, the more GPU memory it consumes. There is no fixed cost per request.
This makes LLM inference expensive, unpredictable, and hard to scale.
We wrote this handbook because the knowledge to solve these problems exists, but it is scattered across research papers, blog posts, source code comments, and tribal knowledge. No single resource connected the full picture: from how GPU memory works, to why decode is slow, to how production systems like vLLM actually solve it.
This is that resource.
---
🚀 Quick Start
git clone https://github.com/harshuljain13/llm-inference-at-scale.git
cd llm-inference-at-scale
pip install -e .
Start reading with Chapter 00: The Transformer at Inference Time.
---
📚 Table of Contents
12 chapters, ~55 modules. Each module is a focused 8-10 minute read with a companion lab.
Part I: Foundations
Chapter 00: The Transformer at Inference Time (4 modules)
| # | Module | Path | |---|--------|------| | 0.1 | Transformer Architecture | transformer_architecture.md | | 0.2 | What Happens During Inference | what_happens_during_inference.md | | 0.3 | Attention and KV Cache | attention_and_kv_cache.md | | 0.4 | Why LLM Inference is Different | why_llm_inference_is_different.md |
Chapter 01: GPU Hardware for Inference (2 modules)
| # | Module | Path | |---|--------|------| | 1.1 | GPU Memory Hierarchy | gpu_memory.md | | 1.2 | The Roofline Model | roofline_fundamentals.md |
Chapter 02: Sizing and Serving (1 module)
| # | Module | Path | |---|--------|------| | 2.1 | Capacity Planning | capacity_planning.md |
Part II: Optimizations
Chapter 03: Attention Variants (6 modules)
| # | Module | Path | |---|--------|------| | 3.1 | Multi-Head Attention (MHA) | mha.md | | 3.2 | MQA and GQA | mqa_gqa.md | | 3.3 | GQA Deep Dive | gqa_deep_dive.md | | 3.4 | Multi-Latent Attention (MLA) | multi_latent_attention.md | | 3.5 | FlashAttention | flash_attention.md |
Chapter 04: KV Cache Engineering (5 modules)
| # | Module | Path | |---|--------|------| | 4.1 | PagedAttention | paged_attention.md | | 4.2 | KV Cache Compression | kv_cache_compression.md | | 4.3 | Smart KV Caching | smart_kv_caching.md | | 4.4 | LMCache | lmcache.md | | 4.5 | Prefix Caching | prefix_caching.md |
Chapter 05: Optimization Techniques (6 modules)
| # | Module | Path | |---|--------|------| | 5.1 | Quantization | quantization.md | | 5.2 | TurboQuant | turboquant.md | | 5.3 | Continuous Batching | continuous_batching.md | | 5.4 | Speculative Decoding | speculative_decoding.md | | 5.5 | Chunked Prefill | chunked_prefill.md | | 5.6 | Inference-Time Compute | inference_time_compute.md |
Part III: Engines and Scaling
Chapter 06: Inference Engines (4 modules)
| # | Module | Path | |---|--------|------| | 6.1 | vLLM | vllm.md | | 6.2 | SGLang | sglang.md | | 6.3 | TensorRT-LLM | tensorrt_llm.md | | 6.4 | NVIDIA Dynamo | nvidia_dynamo.md |
Chapter 07: Scaling (3 modules)
| # | Module | Path | |---|--------|------| | 7.1 | Tensor Parallelism | tensor_parallelism.md | | 7.2 | Mixture-of-Experts Inference | moe_inference.md | | 7.3 | Distillation for Serving | distillation.md |
Part IV: Production
Chapter 08: Serving Infrastructure (7 modules)
| # | Module | Path | |---|--------|------| | 8.1 | Ray Serve | ray_serve.md | | 8.2 | EKS and KServe | eks_kserve.md | | 8.3 | SageMaker | sagemaker.md | | 8.4 | Disaggregated Serving | disaggregated_serving.md | | 8.5 | Cold Start Optimization | cold_start.md | | 8.6 | Cache-Aware Routing | cache_aware_routing.md | | 8.7 | Kubernetes Inference Infrastructure | kubernetes_inference_infrastructure.md |
Chapter 09: Operations (6 modules)
| # | Module | Path | |---|--------|------| | 9.1 | Benchmarking and Metrics | benchmarking.md | | 9.2 | Structured Output and Guided Decoding | structured_output.md | | 9.3 | Edge Deployment | edge_deployment.md | | 9.4 | Inference Metrics and Monitoring | inference_metrics.md | | 9.5 | Multi-Region KV Locality | multi_region_kv_locality.md | | 9.6 | Custom Silicon | custom_silicon.md |
Chapter 10: Production Stories (3 modules)
| # | Module | Path | |---|--------|------| | 10.1 | Meta's Inference Platform | meta_inference_platform.md | | 10.2 | Databricks Multi-Tenant Serving | databricks_multi_tenant.md | | 10.3 | Mixed Workload Management | mixed_workload_management.md |
Chapter 11: System Designs (5 modules)
| # | Module | Path | |---|--------|------| | 11.1 | ChatGPT-Scale Chatbot | chatgpt_scale_chatbot.md | | 11.2 | Code Copilot | code_copilot.md | | 11.3 | Enterprise RAG | enterprise_rag.md | | 11.4 | Multi-Model Gateway | multi_model_gateway.md | | 11.5 | Agentic Workload | agentic_workload.md |
Total: 12 chapters, ~55 modules.
---
🤝 Contributing
Contributions are welcome. This is a living document.
- Fix errors — typos, outdated information, incorrect formulas
- Improve clarity — better explanations, additional examples
- Add content — new modules, labs, or reference materials
---
👤 About the Author
Harshul Jain is a Senior ML Infrastructure Engineer specializing in real-time ML systems, feature stores, and LLM serving infrastructure. He builds and operates ML platforms serving millions of users, mentors 300+ engineers through an eMentoring program, and is a recurring speaker at ML infrastructure conferences.
- GitHub: @harshuljain13
- Newsletter: The Engineer's Digest
⚠️ Disclaimer
The views, techniques, and opinions expressed in this handbook are solely those of the author and do not represent the views of any employer or affiliated organization. No proprietary or confidential information has been included. All content is based on publicly available research, open-source tooling, and independent analysis.
---
📄 License
© 2026 Harshul Jain. All rights reserved.
No part of this work may be reproduced, distributed, modified, or used without prior written permission from the author. To request permission, open a GitHub Issue.
---
🙏 Acknowledgments
This handbook builds on the work of many researchers and engineers:
- The vLLM team for PagedAttention and continuous batching
- The SGLang team for RadixAttention
- Tri Dao for FlashAttention
- The authors of foundational papers: Attention Is All You Need, GQA, Medusa, EAGLE, and many others
Built with care for the ML infrastructure community