Reading List
Being less concrete further out, the reading list is being incrementally updated to include more papers as we go.
Scaling laws
Training Compute-Optimal Large Language Models (☑️ already covered in lecture)
Scaling Laws for Neural Language Models (☑️ already covered in lecture)
Attention and LLM kernel optimization
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness [NIPS 2022]
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving [MLSys 2025]
Mirage: A Multi-Level Superoptimizer for Tensor Programs [OSDI 2025]
LLM serving systems and memory management
Efficient Memory Management for Large Language Model Serving with PagedAttention [ACM SOSP 2023]
H2O: heavy-hitter oracle for efficient generative inference of large language models [NIPS 2023]
FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU [ICML 2023]
InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management [USENIX OSDI 2024]
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving [USENIX OSDI 2024]
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve [USENIX OSDI 2024]
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU [ACM SOSP 2024]
KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models [ACM SOSP 2025]
LLM in a Flash: Efficient Large Language Model Inference with Limited Memory [ACL 2024]
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving [SIGCOMM 2024]
FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution
A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing
Speculative and accelerated decoding
Accelerating Large Language Model Decoding with Speculative Sampling (☑️ will be covered in lecture)
DFlash: Block Diffusion for Flash Speculative Decoding
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
Large-scale training, post-training, and fault tolerance
MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs [USENIX OSDI 2024]
Safeguarding LLM Training at Scale: Online SDC Detection and Insights from 35 Million GPU Hours [USENIX OSDI 2026]
Understanding Stragglers in Large Model Training Using What-if Analysis [USENIX OSDI 2025]
RobustRL: Role-Based Fault Tolerance System for RL Post-Training [USENIX OSDI 2026]
Agentic systems
Murakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud Platforms [USENIX OSDI 2026]
DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale
LLM-as-a-Verifier: A General-Purpose Verification Framework
AI for GPU kernel generation and optimization
Hawkeye: Hardware-Aware GPU Kernel Optimization with Minimal Supervision
KernelBench: Can LLMs Write Efficient GPU Kernels? [ICML 2025]
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation