AI Infra

The premise

Most inference isn't compute-bound. It's waiting on memory.

A GPU only reaches its headline TFLOP/s when a kernel does enough arithmetic per byte it moves. Single-sequence decode does about one FLOP per byte — under 1 % of what an A100 needs to climb off the memory roof. Nearly every serving win starts there.

Roofline model for an NVIDIA A100-80GBAttainable throughput rises with arithmetic intensity along a 1.935 terabyte-per-second memory bandwidth slope until it flattens at the 312 TFLOP/s compute ceiling. The ridge point sits at 161 FLOP per byte. Single-sequence decode sits far to the left at roughly 1 FLOP per byte, deep in the memory-bound region. Batching moves it up the slope: batch 8, batch 64, and at batch 256 it reaches the compute roof.14166425618641.935 TB/s HBMmemory-bound312 TFLOP/s · compute-boundridge point · 161 FLOP/bytedecode, batch = 1 · AI ≈ 1 · 0.6 % of peakbatch = 8 · 15 TFLOP/s · still freebatch = 64 · 124 TFLOP/sbatch = 256 · on the roofarithmetic intensity — FLOP per byte of HBM trafficTFLOP/s attained

365-day study log

365 天 AI Infra 自学课程

从 transformer 零基础开始,每天一篇。已写到 Day 30:M1 总复习:一页笔记、20 道全月自测、错题本汇总与 M2 预告

目录 →