harshuljain13/llm-inference-at-scale

▲ 19 stars today★ 238⑂ 43

A Practitioner handbook for production llm serving.

About harshuljain13/llm-inference-at-scale

harshuljain13/llm-inference-at-scale is an open-source project on GitHub, mainly written in Jupyter Notebook. A Practitioner handbook for production llm serving. It currently holds 238 stars and 43 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the Today's Trending board, currently at rank #40 with 19 new stars today.

GitHub Repository Details

Repository harshuljain13/llm-inference-at-scale · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

https://github.com/harshuljain13/llm-inference-at-scale/blob/HEAD/LLM Inference at Scale

The definitive guide to serving large language models in production.

Quick Start • Contents • Equations • Contributing

https://github.com/harshuljain13/llm-inference-at-scale/blob/HEAD/Python 3.10+ https://github.com/harshuljain13/llm-inference-at-scale/blob/HEAD/vLLM https://github.com/harshuljain13/llm-inference-at-scale/blob/HEAD/SGLang https://github.com/harshuljain13/llm-inference-at-scale/blob/HEAD/TensorRT-LLM https://github.com/harshuljain13/llm-inference-at-scale/blob/HEAD/PyTorch https://github.com/harshuljain13/llm-inference-at-scale/blob/HEAD/CUDA https://github.com/harshuljain13/llm-inference-at-scale/blob/HEAD/All Rights Reserved

---

Why This Exists

Serving an LLM is not like serving a traditional ML model. With traditional ML, you send a request, the model runs one forward pass, and you get a result. Fixed time. Fixed memory. Simple.

With LLMs, every request is different. A short reply takes 100ms. A long one takes 10 seconds. The model generates one word at a time, and each word requires reading the entire model from memory again. The longer the conversation, the more GPU memory it consumes. There is no fixed cost per request.

This makes LLM inference expensive, unpredictable, and hard to scale.

We wrote this handbook because the knowledge to solve these problems exists, but it is scattered across research papers, blog posts, source code comments, and tribal knowledge. No single resource connected the full picture: from how GPU memory works, to why decode is slow, to how production systems like vLLM actually solve it.

This is that resource.

---

🚀 Quick Start

git clone https://github.com/harshuljain13/llm-inference-at-scale.git
cd llm-inference-at-scale
pip install -e .

Start reading with Chapter 00: The Transformer at Inference Time.

---

📚 Table of Contents

12 chapters, ~55 modules. Each module is a focused 8-10 minute read with a companion lab.

Part I: Foundations

Chapter 00: The Transformer at Inference Time (4 modules)

| # | Module | Path | |---|--------|------| | 0.1 | Transformer Architecture | transformer_architecture.md | | 0.2 | What Happens During Inference | what_happens_during_inference.md | | 0.3 | Attention and KV Cache | attention_and_kv_cache.md | | 0.4 | Why LLM Inference is Different | why_llm_inference_is_different.md |

Chapter 01: GPU Hardware for Inference (2 modules)

| # | Module | Path | |---|--------|------| | 1.1 | GPU Memory Hierarchy | gpu_memory.md | | 1.2 | The Roofline Model | roofline_fundamentals.md |

Chapter 02: Sizing and Serving (1 module)

| # | Module | Path | |---|--------|------| | 2.1 | Capacity Planning | capacity_planning.md |

Part II: Optimizations

Chapter 03: Attention Variants (6 modules)

| # | Module | Path | |---|--------|------| | 3.1 | Multi-Head Attention (MHA) | mha.md | | 3.2 | MQA and GQA | mqa_gqa.md | | 3.3 | GQA Deep Dive | gqa_deep_dive.md | | 3.4 | Multi-Latent Attention (MLA) | multi_latent_attention.md | | 3.5 | FlashAttention | flash_attention.md |

Chapter 04: KV Cache Engineering (5 modules)

| # | Module | Path | |---|--------|------| | 4.1 | PagedAttention | paged_attention.md | | 4.2 | KV Cache Compression | kv_cache_compression.md | | 4.3 | Smart KV Caching | smart_kv_caching.md | | 4.4 | LMCache | lmcache.md | | 4.5 | Prefix Caching | prefix_caching.md |

Chapter 05: Optimization Techniques (6 modules)

| # | Module | Path | |---|--------|------| | 5.1 | Quantization | quantization.md | | 5.2 | TurboQuant | turboquant.md | | 5.3 | Continuous Batching | continuous_batching.md | | 5.4 | Speculative Decoding | speculative_decoding.md | | 5.5 | Chunked Prefill | chunked_prefill.md | | 5.6 | Inference-Time Compute | inference_time_compute.md |

Part III: Engines and Scaling

Chapter 06: Inference Engines (4 modules)

| # | Module | Path | |---|--------|------| | 6.1 | vLLM | vllm.md | | 6.2 | SGLang | sglang.md | | 6.3 | TensorRT-LLM | tensorrt_llm.md | | 6.4 | NVIDIA Dynamo | nvidia_dynamo.md |

Chapter 07: Scaling (3 modules)

| # | Module | Path | |---|--------|------| | 7.1 | Tensor Parallelism | tensor_parallelism.md | | 7.2 | Mixture-of-Experts Inference | moe_inference.md | | 7.3 | Distillation for Serving | distillation.md |

Part IV: Production

Chapter 08: Serving Infrastructure (7 modules)

| # | Module | Path | |---|--------|------| | 8.1 | Ray Serve | ray_serve.md | | 8.2 | EKS and KServe | eks_kserve.md | | 8.3 | SageMaker | sagemaker.md | | 8.4 | Disaggregated Serving | disaggregated_serving.md | | 8.5 | Cold Start Optimization | cold_start.md | | 8.6 | Cache-Aware Routing | cache_aware_routing.md | | 8.7 | Kubernetes Inference Infrastructure | kubernetes_inference_infrastructure.md |

Chapter 09: Operations (6 modules)

| # | Module | Path | |---|--------|------| | 9.1 | Benchmarking and Metrics | benchmarking.md | | 9.2 | Structured Output and Guided Decoding | structured_output.md | | 9.3 | Edge Deployment | edge_deployment.md | | 9.4 | Inference Metrics and Monitoring | inference_metrics.md | | 9.5 | Multi-Region KV Locality | multi_region_kv_locality.md | | 9.6 | Custom Silicon | custom_silicon.md |

Chapter 10: Production Stories (3 modules)

| # | Module | Path | |---|--------|------| | 10.1 | Meta's Inference Platform | meta_inference_platform.md | | 10.2 | Databricks Multi-Tenant Serving | databricks_multi_tenant.md | | 10.3 | Mixed Workload Management | mixed_workload_management.md |

Chapter 11: System Designs (5 modules)

| # | Module | Path | |---|--------|------| | 11.1 | ChatGPT-Scale Chatbot | chatgpt_scale_chatbot.md | | 11.2 | Code Copilot | code_copilot.md | | 11.3 | Enterprise RAG | enterprise_rag.md | | 11.4 | Multi-Model Gateway | multi_model_gateway.md | | 11.5 | Agentic Workload | agentic_workload.md |

Total: 12 chapters, ~55 modules.

---

🤝 Contributing

Contributions are welcome. This is a living document.

Fork the repo, create a branch, submit a PR. To report errors, open a GitHub Issue.

---

👤 About the Author

Harshul Jain is a Senior ML Infrastructure Engineer specializing in real-time ML systems, feature stores, and LLM serving infrastructure. He builds and operates ML platforms serving millions of users, mentors 300+ engineers through an eMentoring program, and is a recurring speaker at ML infrastructure conferences.

---

⚠️ Disclaimer

The views, techniques, and opinions expressed in this handbook are solely those of the author and do not represent the views of any employer or affiliated organization. No proprietary or confidential information has been included. All content is based on publicly available research, open-source tooling, and independent analysis.

---

📄 License

© 2026 Harshul Jain. All rights reserved.

No part of this work may be reproduced, distributed, modified, or used without prior written permission from the author. To request permission, open a GitHub Issue.

---

🙏 Acknowledgments

This handbook builds on the work of many researchers and engineers:

---

Built with care for the ML infrastructure community

GitHub Stars & Activity

238Stars
43Forks
0Open issues
Jupyter NotebookLanguage

GitHub Popularity

GitHub stars238
Forks43
Open issues0
Primary languageJupyter Notebook
License-
Stars gained today19
Created-
Last pushed-

Trending History

Daily boardrank #40 · ▲ 19 stars

Related AI Projects

1

microsoft / generative-ai-for-beginners

Jupyter Notebook★ 121,033⑂ 63,699▲ 45 stars
→
2

rasbt / LLMs-from-scratch

Jupyter Notebook★ 106,072⑂ 0
→
3

microsoft / ai-agents-for-beginners

Jupyter Notebook★ 76,494⑂ 25,144▲ 79 stars
→
4

Lordog / dive-into-llms

Jupyter Notebook★ 55,697⑂ 6,643▲ 51 stars
→
5

karpathy / nn-zero-to-hero

Jupyter Notebook★ 24,654⑂ 3,594▲ 9 stars
→
6

AI4Finance-Foundation / FinGPT

Jupyter Notebook★ 21,351⑂ 3,028▲ 17 stars
→
7

DataTalksClub / machine-learning-zoomcamp

Jupyter Notebook★ 14,689⑂ 3,267▲ 11 stars
→
8

ageron / handson-ml3

Jupyter Notebook★ 14,272⑂ 5,343▲ 9 stars
→

More AI Rankings