Token的一生:调度器详解

admin 2026-08-25 05:02:01 网络安全文章 来源:ZONE.CI 全球网 0 阅读模式

文章总结: 本文解析SGLang调度器核心决策逻辑,包括Token预算的三个维度(remtotaltokens、remchunktokens、reminputtokens)、PrefillAdder的请求准入控制与AddReqResult枚举、ChunkedPrefill将长prompt切块以降低短请求TTFT的机制、extend_range精确控制计算范围、Decode阶段批量打包、请求结束与KVCache释放,以及fcfs/lpm/dfs-weight等调度策略的适用场景。 综合评分: 88 文章分类: 其他


Token 的一生:调度器详解

原创

zouyee zouyee

DCOS

2026年8月19日 08:00 上海

在小说阅读器读本章

去阅读

系列定位:中级。理解 SGLang 调度器如何分配 Token 预算、控制并发,以及 Chunked Prefill 机制。

Summary

本篇解析调度器的核心决策逻辑。内容涵盖:Token 预算的三个维度(rem_total_tokens、rem_chunk_tokens、rem_input_tokens);PrefillAdder.add_one_req 的准入控制流程与 AddReqResult 枚举;Chunked Prefill 如何把长 prompt 切块以降低短请求的 TTFT;extend_range 精确控制每轮计算范围;Decode 阶段的批量打包;请求结束与 KV Cache 释放;以及 fcfs/lpm/dfs-weight 等调度策略的适用场景。附录预览了 DFlash 投机解码、DSpark 半自回归草稿、DSA 稀疏注意力三大进阶机制。


1. 调度器面临的核心问题

GPU 的显存是有限的,KV Cache 占用大量显存。调度器需要在每个推理步(forward pass)前回答:

这一步能跑哪些请求?各跑多少 token?

如果贪心地把所有等待请求都塞进去,显存会爆(OOM);如果过于保守,GPU 利用率低,吞吐浪费。

SGLang 的调度策略核心是 Token 预算(Token Budget)。


2. Token 预算的三个维度

schedule_policy.py 里,PrefillAdder(负责本轮 prefill 准入的核心类)维护三个预算计数器:

# python/sglang/srt/managers/schedule_policy.py(精简,真实类名为 PrefillAdder)
class PrefillAdder:
    def __init__(self, ..., rem_input_tokens: int, rem_chunk_tokens: Optional[int], ...):
        self.rem_input_tokens   # 本轮 prefill 阶段输入 token 上限
        self.rem_chunk_tokens   # 本轮 chunked prefill 可用的 token 数(Optional[int])

    @property
    def rem_total_tokens(self):  # 注意:这是一个 @property,动态计算
        # 返回 KV Pool 可用+可驱逐容量 - 已预留偏移量
        ...

| 预算 | 含义 | 典型来源 | | — | — | — | | rem_total_tokens | KV Cache 中还剩多少槽位(@property,动态查询) | 显存大小 / KV block 大小 | | rem_chunk_tokens | 本轮 chunked prefill 的 token 配额 | --chunked-prefill-size | | rem_input_tokens | 本轮最多处理多少输入 token | --max-prefill-tokens |


3. 请求准入:add_one_req

每次调度循环,SchedulePolicy.add_one_req() 决定一个等待中的请求能不能进入本轮:

# 简化示意(非完整代码,省略了边界检查和 LoRA/EOS 等细节)
def add_one_req(self, req, has_chunked_req, truncation_align_size):
    max_new = min(req.sampling_params.max_new_tokens, CLIP_MAX_NEW_TOKENS)
    cand_extend_input_len = len(req.full_untruncated_fill_ids) - len(req.prefix_indices)
    total_tokens = cand_extend_input_len + max_new + self.page_size

    if total_tokens >= self.rem_total_tokens:
        return AddReqResult.NO_TOKEN

    input_tokens = self.ceil_paged_tokens(cand_extend_input_len)

    # 无 chunked prefill(rem_chunk_tokens is None)时:
    # 若已有其他请求在本轮,则检查 prefill token 上限
    if self.rem_chunk_tokens is None:
        if len(self.can_run_list) != 0 and input_tokens >= self.rem_input_tokens:
            return AddReqResult.OTHER
        # 整个请求一次处理完
        req.set_extend_range(len(req.prefix_indices), len(req.full_untruncated_fill_ids))
        self.can_run_list.append(req)
&nbsp; &nbsp;&nbsp;elif&nbsp;input_tokens <=&nbsp;self.rem_chunk_tokens:
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;# 有 chunked prefill 且本轮配额够:一次处理完
&nbsp; &nbsp; &nbsp; &nbsp; req.set_extend_range(len(req.prefix_indices),&nbsp;len(req.full_untruncated_fill_ids))
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;self.can_run_list.append(req)
&nbsp; &nbsp;&nbsp;else:
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;# 有 chunked prefill 且配额不够:只处理一块
&nbsp; &nbsp; &nbsp; &nbsp; trunc_len =&nbsp;self.rem_chunk_tokens //&nbsp;self.page_size *&nbsp;self.page_size
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;if&nbsp;trunc_len <=&nbsp;0:
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;return&nbsp;AddReqResult.OTHER
&nbsp; &nbsp; &nbsp; &nbsp; req.set_extend_range(len(req.prefix_indices),&nbsp;len(req.prefix_indices) + trunc_len)
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;self.can_run_list.append(req)
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;self.new_chunked_req = req

&nbsp; &nbsp;&nbsp;return&nbsp;self.budget_state() &nbsp;# 返回 CONTINUE / NO_TOKEN / OTHER

AddReqResult 有三种结果:

| 结果 | 含义 | | — | — | | NO_TOKEN | KV Cache 满了,请求必须等待 | | OTHER | 本轮 prefill 配额用完,请求推迟 | | CONTINUE | 进入本轮,继续追加更多请求 |

枚举定义见 schedule_policy.py 第 422 行:class AddReqResult(Enum): CONTINUE / NO_TOKEN / OTHER


4. Chunked Prefill(分块预填充)

问题背景

假设有两个请求同时到达:

  • 请求 A:prompt 长度 = 10,000 tokens(长文档摘要)
  • 请求 B:prompt 长度 = 10 tokens(简单问答)

如果按序处理,请求 A 的 prefill(处理全部 10,000 tokens)会独占 GPU 很长时间,导致请求 B 的**首 token 延迟(TTFT)**极高。

Chunked Prefill 的解法

把长请求的 prefill 切成小块,和 decode 请求混合执行:

不使用 Chunked Prefill:
&nbsp; Round 1: [A prefill: 10000 tokens] &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ← B 等待
&nbsp; Round 2: [A decode: 1 token] [B prefill: 10 tokens] [B decode: 1 token]

使用 Chunked Prefill(chunk_size=512):
&nbsp; Round 1: [A prefill chunk 1: 512] [B prefill: 10] [running decodes]
&nbsp; Round 2: [A prefill chunk 2: 512] [running decodes]
&nbsp; ...
&nbsp; Round 20: [A prefill chunk 20: 512] [A decode: 1] [running decodes]

请求 B 的 TTFT 从”等 A 跑完 prefill”变成”等一个 chunk”,大幅降低。

代码位置

# python/sglang/srt/managers/scheduler.py
def&nbsp;init_chunked_prefill(self):
&nbsp; &nbsp;&nbsp;self.chunked_prefill_size =&nbsp;self.server_args.chunked_prefill_size
&nbsp; &nbsp;&nbsp;# chunked_prefill_size=None 表示禁用
&nbsp; &nbsp;&nbsp;# chunked_prefill_size=512 &nbsp;表示每轮最多处理 512 个 prefill tokens
&nbsp; &nbsp;&nbsp;self.chunked_req =&nbsp;None&nbsp;&nbsp;# 当前正在分块处理的请求

# 在每轮调度中:
if&nbsp;self.chunked_req&nbsp;is&nbsp;not&nbsp;None:
&nbsp; &nbsp;&nbsp;# 优先继续上一轮未完成的分块请求
&nbsp; &nbsp; schedule_policy.rem_chunk_tokens =&nbsp;self.chunked_prefill_size
&nbsp; &nbsp;&nbsp;# 把 chunked_req 排在队列最前面

extend_range:精确控制本轮计算范围

Req.extend_range 是 Range(start, end),告诉 TP Worker 本轮只需要计算哪段 tokens 的注意力:

origin_input_ids: [t0, t1, t2, t3, t4, t5, t6, t7, t8, t9]
prefix_indices (cached): [t0, t1, t2] &nbsp; ← 已在 KV Cache 里,不用算

第一轮 extend_range = Range(3, 7) &nbsp; → 计算 t3, t4, t5, t6
第二轮 extend_range = Range(7, 10) &nbsp;→ 计算 t7, t8, t9(完成 prefill)
第三轮 extend_range = ... &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; → decode 阶段,每轮只算 1 个新 token

5. Decode 阶段的调度

Decode 阶段每轮只生成 1 个新 token(或在投机解码下生成多个),相比 prefill 计算量小得多。调度器把所有正在 decode 的请求打包成一个批次(batching),一次 forward pass 同时处理多个请求,充分利用 GPU 并行性。

Decode 批次(假设 4 个请求同时在 decode 阶段):
&nbsp; req_001: origin=[...] + output=[t_out_0, t_out_1, t_out_2] &nbsp;→ 预测 t_out_3
&nbsp; req_002: origin=[...] + output=[t_out_0] &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; → 预测 t_out_1
&nbsp; req_003: origin=[...] + output=[t_out_0, ..., t_out_99] &nbsp; &nbsp; → 预测 t_out_100
&nbsp; req_004: origin=[...] + output=[] &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;→ 预测第一个输出 token

一次 forward pass,同时为 4 个请求各生成 1 个新 token

6. 请求结束与 KV Cache 释放

当请求满足停止条件(达到 max_new_tokens、生成了 EOS、命中 stop_str),finished_reason 被设置:

class&nbsp;Req:
&nbsp; &nbsp; finished_reason:&nbsp;Optional[BaseFinishReason]
&nbsp; &nbsp;&nbsp;# 子类(定义在 schedule_batch.py):
&nbsp; &nbsp;&nbsp;# &nbsp; FINISH_LENGTH &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; — 达到 max_new_tokens
&nbsp; &nbsp;&nbsp;# &nbsp; FINISH_MATCHED_TOKEN &nbsp; &nbsp;— 生成了 stop_token_id
&nbsp; &nbsp;&nbsp;# &nbsp; FINISH_MATCHED_STR &nbsp; &nbsp; &nbsp;— 文本命中 stop 字符串
&nbsp; &nbsp;&nbsp;# &nbsp; FINISHED_MATCHED_REGEX &nbsp;— 文本命中 stop_regex(注意:源码里是 FINISHED_ 而非 FINISH_)
&nbsp; &nbsp;&nbsp;# &nbsp; FINISH_ABORT &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;— 客户端断开或内部错误

请求结束后,调度器释放它在 KV Cache 里独占的槽位,空出来给下一个请求。已计算过的共享前缀(例如 system prompt)则留在 Radix Cache 里,供后续请求复用。


7. 调度策略

SGLang 支持多种调度策略,通过 --schedule-policy 指定:

| 策略 | 类别 | 选请求的方式 | 适用场景 | | — | — | — | — | | fcfs | 缓存无关 | 先来先服务(默认) | 通用 | | lof | 缓存无关 | 最长输出优先(longest output first) | 长输出请求批次 | | random | 缓存无关 | 随机 | 测试/基准 | | routing-key | 缓存无关 | 按 routing_key 路由到同一 worker | 多租户隔离 | | lpm | 缓存感知 | 最长前缀匹配优先 | 重复前缀多(RAG、Agent) | | dfs-weight | 缓存感知 | 深度优先 + 权重 | 树状请求(beam search、tree of thought) |

优先级调度(按 request.priority 字段排序)不是独立的 --schedule-policy 值,而是通过 --enable-priority-scheduling 标志叠加在 fcfs 或 lof 之上启用:--schedule-policy fcfs --enable-priority-scheduling。

LPM 策略和 Radix Cache 配合效果最好:同一个 system prompt 的请求被优先调度在一起,KV Cache 命中率最高。


小结

每轮调度循环:
&nbsp; ① 检查 Radix Cache,更新各请求的 prefix_indices
&nbsp; ② 遍历等待队列,用 add_one_req 做准入控制:
&nbsp; &nbsp; &nbsp;- 检查 rem_total_tokens(显存)
&nbsp; &nbsp; &nbsp;- 检查 rem_chunk_tokens(分块配额)
&nbsp; ③ 构建 ScheduleBatch:
&nbsp; &nbsp; &nbsp;- 正在 decode 的请求(全部)
&nbsp; &nbsp; &nbsp;- 可以 prefill 的请求(受预算限制)
&nbsp; ④ 发给 TP Worker 执行 forward pass
&nbsp; ⑤ 收结果,更新 output_ids,检查终止条件,释放显存

下一篇:Radix Cache 是 SGLang 高效复用 Token 计算的核心,我们来看它是怎么用前缀树管理 KV Cache 的。


附录:Token 生成加速的三大进阶机制

投机解码(Speculative Decoding)背景

标准 decode 每轮只生成 1 个 token,GPU 是 memory-bound 而非 compute-bound——大量时间花在搬运 KV Cache 权重而非真正计算上。投机解码的核心思路:

传统 decode:
&nbsp; Round 1 → 1 token(KV 读 1 次)
&nbsp; Round 2 → 1 token(KV 读 1 次) &nbsp;← 大量时间在等内存

投机解码:
&nbsp; 草稿阶段:小模型快速生成 N 个候选 token(廉价)
&nbsp; 验证阶段:大模型一次 forward 验证所有候选(贵但值)
&nbsp; → 每次大模型 forward 实际产出 ~k 个 token(k ≤ N)
&nbsp; → 吞吐量提升正比于平均接受长度 k

投机解码不改变输出分布(lossless),生成质量与标准 decode 完全等价。


DFlash:混合注意力草稿模型

DFlash 是 SGLang 已实现的投机解码方案(sglang/srt/speculative/dflash_worker_v2.py),草稿模型架构设计为全注意力层与滑窗注意力层交替排列:

DFlash 草稿模型层结构:
&nbsp; Layer 0: sliding_window_attention(只看最近 W 个 token,KV 开销小)
&nbsp; Layer 1: full_attention(看全部上下文,精度高)
&nbsp; Layer 2: sliding_window_attention
&nbsp; Layer 3: full_attention
&nbsp; ...

滑窗层降低草稿阶段的 KV 读写量,全注意力层保证草稿质量。调度器对 DFlash 请求的处理分两步:

# 草稿阶段:
# &nbsp; draft_model.forward(input_ids) → draft_tokens [batch, N]
# &nbsp; 构建 DFlashVerifyInput(draft_token=..., draft_token_num=N)

# 验证阶段(ForwardMode.TARGET_VERIFY):
# &nbsp; target_model.forward(draft_tokens) → logits [batch, N, vocab_size]
# &nbsp; tree_speculative_sampling_target_only() → 接受/拒绝每个 draft token
# &nbsp; accepted_tokens → 追加到 req.output_ids

相关源码:dflash_worker_v2.py / dflash_info.py / models/dflash.py


DSpark:置信度调度的半自回归投机解码

DSpark(arxiv:2606.19348,DeepSeek + 北京大学,2026-06-27)是比 DFlash 更新的投机解码方案,在 acceptance length 上比 DFlash 高出 16.3%–30.9%。

核心创新 1:半自回归草稿生成

传统投机解码的草稿生成要么完全并行(精度差,后续 token 与前驱脱耦),要么完全自回归(速度慢)。DSpark 引入混合架构:

并行骨干网络(Parallel Backbone)
&nbsp; ↓ 一次 forward 同时生成 block 内所有 draft tokens(快)
&nbsp; ↓
轻量顺序模块(Sequential Module / Markov Head)
&nbsp; ↓ 在 block 内引入相邻 token 的依赖关系(防止"后缀衰减")
&nbsp; ↓
Block 内 draft tokens 质量显著高于纯并行方案

“后缀衰减(suffix decay)”指的是:纯并行草稿生成时,block 尾部的 token 因为看不到前驱 token 的预测结果,精度会随位置递减。轻量顺序模块以极小代价修复这个问题。

核心创新 2:置信度调度验证(Confidence-Scheduled Verification)

不同请求的草稿接受率不同。DSpark 根据”前缀存活概率”估算每个请求在当前轮的预期接受长度,动态决定草稿块大小:

高置信度请求(模型对上下文很"笃定")→ 发较长草稿,高接受率
低置信度请求(分叉较多,如创意写作) → 发较短草稿,减少验证浪费

这避免了固定草稿长度带来的两种浪费:对难请求生成太长草稿、对易请求生成太短草稿。

关键指标(DeepSeek 官方测试)

| 指标 | V4-Flash | V4-Pro | | — | — | — | | 单用户生成速度提升 | 60–85% | 57–78% | | 系统总吞吐提升 | 51–400% | — | | vs DFlash acceptance length | +16.3–30.9% | — |

当前 SGLang 支持状态

DSpark 论文于 2026-06-27 发布,SGLang 主干尚未合并 DSpark 支持(截至本文写作时)。DFlash 是当前可用的 SGLang 原生投机解码方案。如需跟踪 DSpark 集成进展,可关注 SGLang 仓库的 speculative/ 目录和相关 Issue。


DSA(Dynamic Sparse Attention):稀疏 KV 访问

上面两种方案解决的是”如何减少大模型 forward 次数”,DSA 解决的是”每次 forward 里如何减少 KV 读取量”。

对于超长上下文(1M tokens),每个 decode 步都读全部 KV Cache 带宽压力极大。DSA 在注意力计算时只访问 top-k 个重要 KV 块:

标准 decode attention:读取 seqlen × kv_heads × head_dim 的全部 KV
DSA decode attention:index → top-k 重要 KV 块(k << seqlen)→ 只读这些块

核心组件在 sglang/srt/layers/attention/dsa/:

  • dsa_indexer.py:维护每个请求的 KV 重要性索引,动态更新
  • dsa_topk_backend.py:基于 top-k 索引的稀疏注意力 kernel

三者关系:Radix Cache 在请求间共享前缀 KV(减少重复计算);DFlash/DSpark 在时间维度批量验证多个 token(减少 forward 次数);DSA 在空间维度稀疏访问 KV(减少单次 forward 带宽)。三者正交,可以叠加启用。


小结

每轮调度循环(完整版):
&nbsp; ① Radix Cache 前缀匹配,更新 prefix_indices
&nbsp; ② add_one_req 做 Token 预算准入控制
&nbsp; ③ 构建 ScheduleBatch:
&nbsp; &nbsp; &nbsp;- prefill 请求(受 chunked prefill 分块)
&nbsp; &nbsp; &nbsp;- decode 请求(全部打包,含投机解码)
&nbsp; ④ TP Worker forward pass:
&nbsp; &nbsp; &nbsp;- DFlash 请求:draft forward → verify forward(TARGET_VERIFY 模式)
&nbsp; &nbsp; &nbsp;- DSA 请求:decode forward 只访问 top-k KV 块
&nbsp; ⑤ 收结果,更新 output_ids,检查终止条件,释放显存

下一篇:Radix Cache 是 SGLang 高效复用 Token 计算的核心,我们来看它是怎么用前缀树管理 KV Cache 的。


免责声明:

本文所载程序、技术方法仅面向合法合规的安全研究与教学场景,旨在提升网络安全防护能力,具有明确的技术研究属性。

任何单位或个人未经授权,将本文内容用于攻击、破坏等非法用途的,由此引发的全部法律责任、民事赔偿及连带责任,均由行为人独立承担,本站不承担任何连带责任。

本站内容均为技术交流与知识分享目的发布,若存在版权侵权或其他异议,请通过邮件联系处理,具体联系方式可点击页面上方的联系我。

本文转载自:DCOS zouyee zouyee《Token 的一生:调度器详解》

[EDU]从0到1的的接管 网络安全文章

[EDU]从0到1的的接管

文章总结: 本文分享了一次渗透测试实战经验,从发现公开的accesstoken入手,通过翻找JS文件找到接口并构造请求包获取登录凭证,进而获取应用配置和会话列表
评论:0   参与:  0