预计阅读 8 分钟

CUDA Graph:捕获 / 回放 / 失败降级

谁该读这一篇? 想搞清"CUDA Graph 究竟省了什么 / 什么时候不应该开"的工程师。 前置阅读: 03-code-walkthrough/04-model-runner.md §9。 耗时: 30 分钟 学完能: 1. 解释 CUDA Graph 的核心收益(消除 kernel launch 开销); 2. 说出 SGLang 在哪一步捕获 / 回放; 3. 知道什么变量会导致 graph fallback; 4. 调整 decode/prefill 分阶段的 graph backend、batch size 与上限; 5. 排查 "Capture cuda graph 卡 60s+" 类启动问题。


1. 一个 kernel launch 也是不便宜的

PyTorch 跑一次 forward 要发起几百到几千次 CUDA kernel launch,每次大约 5-10 μs。 一次 decode 的"算"时间可能只有 1.5 ms,但这些 launch 累计起来就有 0.5-1 ms,不算小数。

CUDA Graph:把"这一串 launch 的拓扑"录下来,下次直接提交整张图,省掉 host ↔ device 的多次同步。 省下来的就是上面那 0.5-1 ms。


2. SGLang 怎么用 CUDA Graph

源码:runner/decode_cuda_graph_runner.py、runner/prefill_cuda_graph_runner.py(统一接口在 runner/base_cuda_graph_runner.py)。

2.1 启动时捕获

ModelRunner.init_cuda_graphs → capture_cuda_graphs(model_runner_components/cuda_graph_setup.py:95)→ decode/prefill runner 的 capture。

def capture(self):
    bs_to_capture = get_batch_sizes_to_capture(self.model_runner)
    # 通常: [1, 2, 4, 8, 16, 32, 64, 128, 256] 或自定义
    # runner 内部按 batch size 建立静态输入和 graph
    ...

runner 的 capture 实现(decode :1027,prefill :1342)会为每个 batch size 建立静态输入、attention metadata 和 output buffer。下面是概念化流程:

def capture_for_batch_size(self, bs):  # 概念化伪代码
    dummy_batch = self.make_dummy_decode_batch(bs)
    graph = torch.cuda.CUDAGraph()
    with torch.cuda.graph(graph, pool=self.global_pool):
        out = self.model_runner.model.forward(dummy_batch.input_ids, ..., dummy_batch)
    self.graph_pool[bs] = (graph, out)

pool=self.global_pool 让多个 graph 共享一块 memory pool,避免显存放大。

启动日志里 "Capture cuda graph begin/end" 就是这一步,耗时 5-30s。

2.2 运行时回放

ModelRunner.forward → decode runner replay(model_runner.py:1563)。

def replay(self, forward_batch):
    bs = forward_batch.batch_size
    graph, out_buf = self.graph_pool[bs]
    self.replay_prepare(forward_batch)         # 把输入塞进 input buffer
    graph.replay()
    return out_buf                              # output 在固定 buffer 里

replay_prepare 做两件事:

  • copy forward_batch.input_ids 等到捕获时分配的 input buffer。
  • copy 当前 batch 的 attention metadata(page table 等)到捕获时分配的元数据 buffer。

关键不变量:所有"形状可变"的张量都必须用捕获时分配的 buffer,运行时只 copy 内容、不重新分配。 这是 CUDA Graph 的硬约束(捕获时记录的是地址)。


3. 什么 batch size 被捕获

get_batch_sizes_to_capture(runner/base_cuda_graph_runner.py:64):

def get_batch_sizes_to_capture(model_runner, num_tokens_per_bs=1):
    # 默认指数序列 [1, 2, 4, 8, 16, ..., max_running_requests]
    # 可以用 --cuda-graph-bs-decode 强制
    ...

可调参数:

参数 含义
--cuda-graph-bs-decode 显式列出 decode capture batch size,如 1 4 16 64 128
--cuda-graph-max-bs-decode decode 最大 capture bs
--cuda-graph-backend-prefill breakable / tc_piecewise / disabled 等 prefill backend
--cuda-graph-backend-decode full / breakable / disabled 等 decode backend
--disable-cuda-graph-padding 关掉 batch size padding

Padding 机制

如果 batch size 是 5,没有 5 的 graph,会 padding 到下一个有的(8):

real batch_size = 5
padded to 8
graph_replay(bs=8)
output[:5] 取前 5 个有效

代价:算了 3 个 dummy 请求。但比 fallback 到 eager 还划算。 关 padding 后小 batch 走 eager 路径。


4. Fallback 触发条件

CUDA Graph 不能用时回退到 eager forward。触发条件由 ModelRunner 和具体 runner 的 batch-shape 检查共同决定:

情况 原因
batch 含 prefill 请求 Graph 不支持变长输入
batch_size > max_bs 没捕获过
使用了 LoRA 但 adapter 没 captured LoRA 路径未支持 graph
受约束解码切换 grammar grammar 元数据不在 graph 里
Speculative draft tree 形状改变 tree attention 不固定

可观察 metrics:

sglang:cuda_graph_passes_total{mode="decode_cuda_graph"}
sglang:cuda_graph_passes_total{mode="decode_none"}

用 mode 标签比较 graph 与 eager 次数;持续大量 decode_none 或 prefill_none 时,再结合 batch shape 和 runner 日志排查。


5. Memory Pool 与显存

捕获每个 graph 都需要中间张量和 allocator scratch。 SGLang 让 decode / prefill graph 共享进程级 memory pool,并让串行的 CUDA capture pass 复用同一条 capture stream(runner_utils/pool.py:87):

pool = get_or_create_global_graph_memory_pool(torch.cuda)
stream = get_or_create_global_graph_capture_stream()
with graph_capture(stream=stream):
    capture_all_shapes(pool=pool)

decode 与 prefill 不会并发 replay,因此 pool 只需保留两阶段较大的 footprint; 共享 capture stream 又避免 allocator scratch 为每条 capture stream 各留一份。 效果是 N 个 graph 不会让显存翻 N 倍,只是共享池增大一些。 但仍有开销,启动日志里 "Memory pool end" 之后会看到显存被多吃几 GB。


6. CUDA Graph + Overlap 调度

Overlap 模式下 graph 在 forward_stream 上 replay; 准备工作(replay_prepare)在 schedule_stream 上做。 两个 stream 同时跑,graph replay 直接拿"上一轮 prepare 好的 buffer"。

实现复杂,但收益是 TPOT 再降 10-15%。详见 decode runner 的执行入口(decode_cuda_graph_runner.py:1449)。


7. 与 torch.compile 的关系

# 启动时
def set_torch_compile_config():
    torch._dynamo.config.cache_size_limit = ...
    torch._inductor.config.coordinate_descent_tuning = True

SGLang 支持 torch.compile 编译模型(compilation/)。 两者关系:

  • torch.compile 在 op 层做 fusion,产出更少的 kernel。
  • CUDA Graph 在更上层把"一串 kernel launch" 录下来。
  • 配合用:先 compile(少 kernel)再 capture(去 launch 开销)。

启动:

--enable-torch-compile           # 开 torch.compile
--torch-compile-max-bs 128       # compile 的 batch size 上限

8. 启动慢排查

"Capture cuda graph"卡 60s+:

  • 减 --cuda-graph-bs-decode 列表。
  • 关 --enable-torch-compile(compile 也会慢)。
  • 模型超大(70B+)单 bs capture 就要 10s,多 bs 累计很慢。
  • 调试期间把 decode/prefill backend 都设为 disabled 跳过。

9. 生产部署建议

场景 配置
主流 7B-70B 模型 默认即可
模型超大 + 启动时间敏感 自定义 --cuda-graph-bs-decode 1 32 128(少 bs)
LoRA 多 adapter 用 cuda_graph_passes_total 验证 eager 比例后再决定是否分流/关 graph
Disaggregated decoder 默认开
Debug 模式 --cuda-graph-backend-decode disabled --cuda-graph-backend-prefill disabled

10. 小结

  • CUDA Graph 录一串 kernel launch,下次 replay 省 host launch 开销;TPOT 收益依赖模型、batch 和硬件。
  • SGLang 启动时给一组常见 batch size 各捕获一份;decode/prefill 共用 memory pool 与 capture stream,运行时 batch_size + 形状匹配才能 replay。
  • Padding 机制:batch_size 不匹配时填到下一档。
  • Prefill / LoRA / spec draft 形状变化触发 fallback。
  • 关掉 graph 启动快、运行慢;生产建议保留默认开。

11. 自检

  1. CUDA Graph 省的是什么开销?
答案 **Kernel launch 开销**——每次 `kernel<<>>(...)` 调用 host 都要做:(a) 参数打包;(b) 写入 GPU command queue;(c) 设置 stream 同步。每次 launch 平均 5-10 μs。 一次 LLM forward 有几百到几千个 kernel(每层 attention + MLP + layernorm + ...),累计 launch 开销 0.5-1 ms。 Graph 录制把"这串 launch 的拓扑"录下来,replay 时一次 `graph.replay()` 把整个序列直接提交,省掉 host ↔ device 的多次同步。 省的是 host-side overhead,不是 GPU 算力;所以单 token decode(GPU 算力本就少)收益最大。
  1. 为什么普通 prefill 通常不用 graph?
答案 CUDA Graph 录制时把张量地址和形状都"烙进"图里,replay 时形状必须完全一致。 Prefill 每请求的 prompt 长度不同,input shape 很难复用同一张图,因此普通 extend 通常走 eager。 不是绝对不能用:固定 chunk、满足 piecewise runner 条件的 extend 或 target verify 可由 `prefill_cuda_graph_runner.py` capture/replay。
  1. Padding 和 fallback 哪个代价大?
答案 **通常 fallback 代价大**。 - Padding 到下一档(如 real bs=5 → graph bs=8):算了 3 个 dummy 请求的算力(GPU 浪费 ~37%),但仍享 graph 的 launch 省开销。整体时延仅比 graph bs=5 高一点点,比 eager 快得多。 - Fallback 到 eager:失去 graph 的 host launch 优势,但避免 dummy 计算。 哪个更快取决于 padding 比例、模型层数和 GPU 利用率,应按实际 batch 分布 benchmark。 关闭 padding(`--disable-cuda-graph-padding`)时小 batch 走 eager 路径,建议保留 padding 默认开。
  1. Graph 和 torch.compile 重复吗?
答案 **不重复,互补**。 torch.compile 在算子层做 fusion + 代码生成:把 elementwise + matmul + layernorm 一系列小 op 融合成少数几个大 kernel(Inductor 产出)。**产出更少 kernel**。 CUDA Graph 在更上层把"一串 kernel launch" 录下来。**让剩下的 launch 不花 CPU 时间**。 两者顺序:先 compile(少 kernel)再 capture(去 launch 开销),可以叠加但收益需实测。 SGLang 启动 `--enable-torch-compile --cuda-graph-max-bs-decode 128` 同时开。 代价:torch.compile 本身首次 capture 慢(30-60s),总启动时间长。
  1. 启动 capture 卡 90s,怎么诊断 + 缓解?
答案 诊断(按概率): (a) **graph batch size 列表太长**:默认 9 个(1,2,4,...,256),70B 模型每个 5-10s,累计 1.5 分钟。 (b) **同时开 torch.compile**:每个 bs 多花 5-20s。 (c) **TP 大** + NCCL 慢:每个 graph 录制都要走 AllReduce,节点间 NCCL 慢拖累。 (d) **模型超大**:DeepSeek-V3 671B 每 bs capture 几十秒正常。 缓解: - `--cuda-graph-bs-decode 1 32 128` 减档(牺牲部分 batch size 的 padding 表现)。 - `--cuda-graph-max-bs-decode 128` 上限切低。 - 关 torch.compile(debug 阶段)。 - 把模型 ramdisk + NCCL 调好。 - 接受 90s:用 K8s `startupProbe initialDelaySeconds=120s`,扩容预热池避免突发流量等待。

12. 下一步

上游源码:sglang/python/sglang/srt/model_executor/runner/、model_runner_components/cuda_graph_setup.py。