Open reference library / 523 lessons
Study the mechanics. Build the system.
Explore the full English lesson sequence from AI Engineering from Scratch, alongside our governed agent field course. The original curriculum spans mathematics, machine learning, language models, tools, agents, infrastructure, safety, and capstones.
Reference material by Rohit Ghumare and contributors, MIT licensed. Imported at revision bf7791e14076. Runnable code and outputs remain linked to the source repository.
Connect your agent to this foundation ↗ — public search and lesson retrieval, with a reusable onboarding guide.
28 lessons
17 / infrastructure and production Managed LLM Platforms — Bedrock, Vertex AI, Azure OpenAI ~60 minutes ↗17 / infrastructure and production Inference Platform Economics — Fireworks, Together, Baseten, Modal, Replicate, Anyscale ~60 minutes ↗17 / infrastructure and production GPU Autoscaling on Kubernetes — Karpenter, KAI Scheduler, Gang Scheduling ~75 minutes ↗17 / infrastructure and production Serving Engine Internals — PagedAttention, Continuous Batching, Chunked Prefill ~75 minutes ↗17 / infrastructure and production EAGLE-3 Speculative Decoding in Production ~60 minutes ↗17 / infrastructure and production Prefix-Cache Serving — RadixAttention and KV Reuse ~75 minutes ↗17 / infrastructure and production Hardware-Specialized Inference Compilation — FP8 and NVFP4 on Blackwell ~75 minutes ↗17 / infrastructure and production Inference Metrics — TTFT, TPOT, ITL, Goodput, P99 ~60 minutes ↗17 / infrastructure and production Production Quantization — AWQ, GPTQ, GGUF K-quants, FP8, MXFP4/NVFP4 ~75 minutes ↗17 / infrastructure and production Cold Start Mitigation for Serverless LLMs ~60 minutes ↗17 / infrastructure and production Multi-Region LLM Serving and KV Cache Locality ~60 minutes ↗17 / infrastructure and production Edge Inference — Apple Neural Engine, Qualcomm Hexagon, WebGPU/WebLLM, Jetson ~60 minutes ↗17 / infrastructure and production LLM Observability Stack Selection ~60 minutes ↗17 / infrastructure and production Prompt Caching and Semantic Caching Economics ~60 minutes ↗17 / infrastructure and production Batch APIs — the 50% Discount as Industry Standard ~45 minutes ↗17 / infrastructure and production Model Routing as a Cost-Reduction Primitive ~60 minutes ↗17 / infrastructure and production Disaggregated Prefill/Decode — NVIDIA Dynamo and llm-d ~75 minutes ↗17 / infrastructure and production Production Serving Stack — KV Offloading and Cache-Aware Routing ~60 minutes ↗17 / infrastructure and production AI Gateways — LiteLLM, Portkey, Kong AI Gateway, Bifrost ~60 minutes ↗17 / infrastructure and production Shadow Traffic, Canary Rollout, and Progressive Deployment for LLMs ~60 minutes ↗17 / infrastructure and production A/B Testing LLM Features — GrowthBook, Statsig, and the Vibes Problem ~60 minutes ↗17 / infrastructure and production Load Testing LLM APIs — Why k6 and Locust Lie ~75 minutes ↗17 / infrastructure and production SRE for AI — Multi-Agent Incident Response, Runbooks, Predictive Detection ~60 minutes ↗17 / infrastructure and production Chaos Engineering for LLM Production ~60 minutes ↗17 / infrastructure and production Security — Secrets, API Key Rotation, Audit Logs, Guardrails ~60 minutes ↗17 / infrastructure and production Compliance — SOC 2, HIPAA, GDPR, PCI-DSS, EU AI Act, ISO 42001 ~60 minutes ↗17 / infrastructure and production FinOps for LLMs — Unit Economics and Multi-Tenant Attribution ~60 minutes ↗17 / infrastructure and production Self-Hosted Serving Selection — Matching Engine to Hardware and Scale ~45 minutes ↗