CUDA Graph:捕获 / 回放 / 失败降级
谁该读这一篇? 想搞清"CUDA Graph 究竟省了什么 / 什么时候不应该开"的工程师。 前置阅读:
03-code-walkthrough/04-model-runner.md§9。 耗时: 30 分钟 学完能: 1. 解释 CUDA Graph 的核心收益(消除 kernel launch 开销); 2. 说出 SGLang 在哪一步捕获 / 回放; 3. 知道什么变量会导致 graph fallback; 4. 调整 decode/prefill 分阶段的 graph backend、batch size 与上限; 5. 排查 "Capture cuda graph 卡 60s+" 类启动问题。
1. 一个 kernel launch 也是不便宜的
PyTorch 跑一次 forward 要发起几百到几千次 CUDA kernel launch,每次大约 5-10 μs。 一次 decode 的"算"时间可能只有 1.5 ms,但这些 launch 累计起来就有 0.5-1 ms,不算小数。
CUDA Graph:把"这一串 launch 的拓扑"录下来,下次直接提交整张图,省掉 host ↔ device 的多次同步。 省下来的就是上面那 0.5-1 ms。
2. SGLang 怎么用 CUDA Graph
源码:runner/decode_cuda_graph_runner.py、runner/prefill_cuda_graph_runner.py(统一接口在 runner/base_cuda_graph_runner.py)。
2.1 启动时捕获
ModelRunner.init_cuda_graphs → capture_cuda_graphs(model_runner_components/cuda_graph_setup.py:95)→ decode/prefill runner 的 capture。
def capture(self):
bs_to_capture = get_batch_sizes_to_capture(self.model_runner)
# 通常: [1, 2, 4, 8, 16, 32, 64, 128, 256] 或自定义
# runner 内部按 batch size 建立静态输入和 graph
...
runner 的 capture 实现(decode :1027,prefill :1342)会为每个 batch size
建立静态输入、attention metadata 和 output buffer。下面是概念化流程:
def capture_for_batch_size(self, bs): # 概念化伪代码
dummy_batch = self.make_dummy_decode_batch(bs)
graph = torch.cuda.CUDAGraph()
with torch.cuda.graph(graph, pool=self.global_pool):
out = self.model_runner.model.forward(dummy_batch.input_ids, ..., dummy_batch)
self.graph_pool[bs] = (graph, out)
pool=self.global_pool 让多个 graph 共享一块 memory pool,避免显存放大。
启动日志里 "Capture cuda graph begin/end" 就是这一步,耗时 5-30s。
2.2 运行时回放
ModelRunner.forward → decode runner replay(model_runner.py:1563)。
def replay(self, forward_batch):
bs = forward_batch.batch_size
graph, out_buf = self.graph_pool[bs]
self.replay_prepare(forward_batch) # 把输入塞进 input buffer
graph.replay()
return out_buf # output 在固定 buffer 里
replay_prepare 做两件事:
- copy
forward_batch.input_ids等到捕获时分配的 input buffer。 - copy 当前 batch 的 attention metadata(page table 等)到捕获时分配的元数据 buffer。
关键不变量:所有"形状可变"的张量都必须用捕获时分配的 buffer,运行时只 copy 内容、不重新分配。 这是 CUDA Graph 的硬约束(捕获时记录的是地址)。
3. 什么 batch size 被捕获
get_batch_sizes_to_capture(runner/base_cuda_graph_runner.py:64):
def get_batch_sizes_to_capture(model_runner, num_tokens_per_bs=1):
# 默认指数序列 [1, 2, 4, 8, 16, ..., max_running_requests]
# 可以用 --cuda-graph-bs-decode 强制
...
可调参数:
| 参数 | 含义 |
|---|---|
--cuda-graph-bs-decode |
显式列出 decode capture batch size,如 1 4 16 64 128 |
--cuda-graph-max-bs-decode |
decode 最大 capture bs |
--cuda-graph-backend-prefill |
breakable / tc_piecewise / disabled 等 prefill backend |
--cuda-graph-backend-decode |
full / breakable / disabled 等 decode backend |
--disable-cuda-graph-padding |
关掉 batch size padding |
Padding 机制
如果 batch size 是 5,没有 5 的 graph,会 padding 到下一个有的(8):
real batch_size = 5
padded to 8
graph_replay(bs=8)
output[:5] 取前 5 个有效
代价:算了 3 个 dummy 请求。但比 fallback 到 eager 还划算。 关 padding 后小 batch 走 eager 路径。
4. Fallback 触发条件
CUDA Graph 不能用时回退到 eager forward。触发条件由 ModelRunner 和具体 runner 的 batch-shape 检查共同决定:
| 情况 | 原因 |
|---|---|
| batch 含 prefill 请求 | Graph 不支持变长输入 |
| batch_size > max_bs | 没捕获过 |
| 使用了 LoRA 但 adapter 没 captured | LoRA 路径未支持 graph |
| 受约束解码切换 grammar | grammar 元数据不在 graph 里 |
| Speculative draft tree 形状改变 | tree attention 不固定 |
可观察 metrics:
sglang:cuda_graph_passes_total{mode="decode_cuda_graph"}
sglang:cuda_graph_passes_total{mode="decode_none"}
用 mode 标签比较 graph 与 eager 次数;持续大量 decode_none 或
prefill_none 时,再结合 batch shape 和 runner 日志排查。
5. Memory Pool 与显存
捕获每个 graph 都需要中间张量和 allocator scratch。
SGLang 让 decode / prefill graph 共享进程级 memory pool,并让串行的 CUDA capture pass
复用同一条 capture stream(runner_utils/pool.py:87):
pool = get_or_create_global_graph_memory_pool(torch.cuda)
stream = get_or_create_global_graph_capture_stream()
with graph_capture(stream=stream):
capture_all_shapes(pool=pool)
decode 与 prefill 不会并发 replay,因此 pool 只需保留两阶段较大的 footprint; 共享 capture stream 又避免 allocator scratch 为每条 capture stream 各留一份。 效果是 N 个 graph 不会让显存翻 N 倍,只是共享池增大一些。 但仍有开销,启动日志里 "Memory pool end" 之后会看到显存被多吃几 GB。
6. CUDA Graph + Overlap 调度
Overlap 模式下 graph 在 forward_stream 上 replay;
准备工作(replay_prepare)在 schedule_stream 上做。
两个 stream 同时跑,graph replay 直接拿"上一轮 prepare 好的 buffer"。
实现复杂,但收益是 TPOT 再降 10-15%。详见 decode runner 的执行入口(decode_cuda_graph_runner.py:1449)。
7. 与 torch.compile 的关系
# 启动时
def set_torch_compile_config():
torch._dynamo.config.cache_size_limit = ...
torch._inductor.config.coordinate_descent_tuning = True
SGLang 支持 torch.compile 编译模型(compilation/)。
两者关系:
torch.compile在 op 层做 fusion,产出更少的 kernel。- CUDA Graph 在更上层把"一串 kernel launch" 录下来。
- 配合用:先 compile(少 kernel)再 capture(去 launch 开销)。
启动:
--enable-torch-compile # 开 torch.compile
--torch-compile-max-bs 128 # compile 的 batch size 上限
8. 启动慢排查
"Capture cuda graph"卡 60s+:
- 减
--cuda-graph-bs-decode列表。 - 关
--enable-torch-compile(compile 也会慢)。 - 模型超大(70B+)单 bs capture 就要 10s,多 bs 累计很慢。
- 调试期间把 decode/prefill backend 都设为
disabled跳过。
9. 生产部署建议
| 场景 | 配置 |
|---|---|
| 主流 7B-70B 模型 | 默认即可 |
| 模型超大 + 启动时间敏感 | 自定义 --cuda-graph-bs-decode 1 32 128(少 bs) |
| LoRA 多 adapter | 用 cuda_graph_passes_total 验证 eager 比例后再决定是否分流/关 graph |
| Disaggregated decoder | 默认开 |
| Debug 模式 | --cuda-graph-backend-decode disabled --cuda-graph-backend-prefill disabled |
10. 小结
- CUDA Graph 录一串 kernel launch,下次 replay 省 host launch 开销;TPOT 收益依赖模型、batch 和硬件。
- SGLang 启动时给一组常见 batch size 各捕获一份;decode/prefill 共用 memory pool 与 capture stream,运行时 batch_size + 形状匹配才能 replay。
- Padding 机制:batch_size 不匹配时填到下一档。
- Prefill / LoRA / spec draft 形状变化触发 fallback。
- 关掉 graph 启动快、运行慢;生产建议保留默认开。
11. 自检
- CUDA Graph 省的是什么开销?
答案
**Kernel launch 开销**——每次 `kernel<<- 为什么普通 prefill 通常不用 graph?
答案
CUDA Graph 录制时把张量地址和形状都"烙进"图里,replay 时形状必须完全一致。 Prefill 每请求的 prompt 长度不同,input shape 很难复用同一张图,因此普通 extend 通常走 eager。 不是绝对不能用:固定 chunk、满足 piecewise runner 条件的 extend 或 target verify 可由 `prefill_cuda_graph_runner.py` capture/replay。- Padding 和 fallback 哪个代价大?
答案
**通常 fallback 代价大**。 - Padding 到下一档(如 real bs=5 → graph bs=8):算了 3 个 dummy 请求的算力(GPU 浪费 ~37%),但仍享 graph 的 launch 省开销。整体时延仅比 graph bs=5 高一点点,比 eager 快得多。 - Fallback 到 eager:失去 graph 的 host launch 优势,但避免 dummy 计算。 哪个更快取决于 padding 比例、模型层数和 GPU 利用率,应按实际 batch 分布 benchmark。 关闭 padding(`--disable-cuda-graph-padding`)时小 batch 走 eager 路径,建议保留 padding 默认开。- Graph 和 torch.compile 重复吗?
答案
**不重复,互补**。 torch.compile 在算子层做 fusion + 代码生成:把 elementwise + matmul + layernorm 一系列小 op 融合成少数几个大 kernel(Inductor 产出)。**产出更少 kernel**。 CUDA Graph 在更上层把"一串 kernel launch" 录下来。**让剩下的 launch 不花 CPU 时间**。 两者顺序:先 compile(少 kernel)再 capture(去 launch 开销),可以叠加但收益需实测。 SGLang 启动 `--enable-torch-compile --cuda-graph-max-bs-decode 128` 同时开。 代价:torch.compile 本身首次 capture 慢(30-60s),总启动时间长。- 启动 capture 卡 90s,怎么诊断 + 缓解?
答案
诊断(按概率): (a) **graph batch size 列表太长**:默认 9 个(1,2,4,...,256),70B 模型每个 5-10s,累计 1.5 分钟。 (b) **同时开 torch.compile**:每个 bs 多花 5-20s。 (c) **TP 大** + NCCL 慢:每个 graph 录制都要走 AllReduce,节点间 NCCL 慢拖累。 (d) **模型超大**:DeepSeek-V3 671B 每 bs capture 几十秒正常。 缓解: - `--cuda-graph-bs-decode 1 32 128` 减档(牺牲部分 batch size 的 padding 表现)。 - `--cuda-graph-max-bs-decode 128` 上限切低。 - 关 torch.compile(debug 阶段)。 - 把模型 ramdisk + NCCL 调好。 - 接受 90s:用 K8s `startupProbe initialDelaySeconds=120s`,扩容预热池避免突发流量等待。12. 下一步
03-code-walkthrough/04-model-runner.md— Graph 的上层用户。01-flashinfer.md— Attention backend 的 graph 兼容性。- 源码:
runner/decode_cuda_graph_runner.py、runner/prefill_cuda_graph_runner.py。
上游源码:
sglang/python/sglang/srt/model_executor/runner/、model_runner_components/cuda_graph_setup.py。