返回专栏
Agent SDK/05 · DSH/5.3

DSH 核心机制与插件开发

轮次与步骤:一个步骤(step)是一次模型请求,加上它触发的工具调用;

预计阅读
20分钟
全文字数
3,655字
资料截至
2026-10-09
Agent SDK · DSH5.3
版本与时效声明
  • 机制部分(会话日志、工具流水线、沙箱)依据的是仓库 commit 5badb15 中的中文官方文档和源码。插件示例依据的是 npm 上的 rc 包,笔者已在本地实际运行过。
  • DSH 是开发者预览版,会话格式已经演进到 v4,接口还在快速变化中1。笔者发现 npm 上的 rc 包落后于仓库源码,具体差异见第 8 节的踩坑记录。
  • 资料截至 2026-10-09,请以 仓库文档 为准。
本文要点
  1. 轮次与步骤:一个**步骤(step)是一次模型请求,加上它触发的工具调用;一个轮次(turn)**包含零个或多个步骤,从领取第一条输入时开始,到不再有待完成的工作时结束2。
  2. 会话日志是唯一的事实来源:模型看到的上下文由 deriveMessages() 从日志中投影出来。规则是“模型可见即已记录”:每一次模型请求都必须能从日志中重建2。
  3. 工具执行流水线:先记录 tool/call;再依次经过 tools/pre-execute waterfall(hooks、权限、沙箱)、单调守卫、审批、tools/execute waterfall(超时、重试、指标)、工具本身、tools/post-execute;最后写入 tool/result3。
  4. fail-closed 沙箱:有 read-only、workspace-write、danger-full-access 三档模式。如果后端无法强制执行所请求的模式,调用会以 SANDBOX_UNAVAILABLE 失败,而绝不会在不受限的状态下运行4。

1. 一个轮次是怎样运行的#

下图根据官方的 Agent 轮次与步骤生命周期 时序图做了简化5:

Session 日志ctx.toolsctx.llmHook 监听器Driver(agent-loop 插件)Agent(ctx.agents)用户 / SDKSession 日志ctx.toolsctx.llmHook 监听器Driver(agent-loop 插件)Agent(ctx.agents)用户 / SDKloop[每个工具调用]工具结果还需要再请求一次模型?那就开始下一个 stepfollowup(content)排队的工作唤醒 driverturn/startagent/pre-step(waterfall:拒绝,或接纳这批消息)step/startagent/request(waterfall:可以替换调用配置)system/message、user/message、request/header 等从日志派生并冻结本次请求llm/stream(waterfall)流式 chunkassistant/message(嵌入完整的 stream)tool/callpre → execute → posttool/resultstep/endagent/turn-stopping(serial:最后一次检查)turn/end
Session 日志ctx.toolsctx.llmHook 监听器Driver(agent-loop 插件)Agent(ctx.agents)用户 / SDKSession 日志ctx.toolsctx.llmHook 监听器Driver(agent-loop 插件)Agent(ctx.agents)用户 / SDKloop[每个工具调用]工具结果还需要再请求一次模型?那就开始下一个 stepfollowup(content)排队的工作唤醒 driverturn/startagent/pre-step(waterfall:拒绝,或接纳这批消息)step/startagent/request(waterfall:可以替换调用配置)system/message、user/message、request/header 等从日志派生并冻结本次请求llm/stream(waterfall)流式 chunkassistant/message(嵌入完整的 stream)tool/callpre → execute → posttool/resultstep/endagent/turn-stopping(serial:最后一次检查)turn/end
  • turn/*、step/*、system/message、user/message、assistant/message、assistant/attempt、tool/* 都是持久化的会话事件,其余是实时的扩展点2。
  • agent/pre-step、agent/request、llm/stream,以及三个 tools/* 事件都是 waterfall,监听器必须调用 next() 才会把控制权交给下游。agent/turn-stopping 是 serial 事件2。
术语提醒

DSH 的“轮次”和“步骤”与其他 SDK 的用法不同。Claude Agent SDK 中的一个 turn 相当于 DSH 的一个 step。详见 1.1 从 LLM API 到 Agent Harness 的分层 › 5. 关键术语速查。


2. 会话日志:唯一的事实来源#

2.1 核心规则#

模型可见即已记录

每一次模型请求都必须能从日志中重建。新增任何模型可见的输入,都需要对应一个会话事件2。

会话日志的作用体现在三个方面2:

  • 它是模型所见上下文的来源,deriveMessages() 从中投影出模型历史;
  • fork、恢复、transcript、遥测和持久化,都从持久化下来的 settlement 派生;
  • 实时 UI 的增量更新则来自 agent/assistant-stream,它是瞬态的,不会持久化。

2.2 一段真实的日志#

下面这段取自仓库里的快照测试 fixture(snapshots/web/auto-review-denial/session.v3.jsonl),笔者做了截断。{{...}} 是 fixture 中原本就有的占位符:

{"type":"session","version":3,"id":"{{session:1}}","cwd":"{{cwd}}","delegationDepth":0,"agentPreset":"standard"}
{"type":"sandbox/mode","data":{"mode":"danger-full-access"}}
{"type":"turn/start","data":{"turn":1}}
{"type":"step/start","data":{"turn":1,"step":1}}
{"type":"user/message","data":{"content":[{"type":"text","text":"Inspect the protected operation, but do not run it unless authorized."}],"role":"user"}}
{"type":"assistant/message","data":{"turn":1,"step":1,"message":{"role":"assistant","content":[{"type":"tool-call","id":"auto-review-denied-call","name":"mystery","arguments":"{\"secret\":\"hidden-input\"}"}],"source":{"kind":"model","provider":"deepseek-official","model":"deepseek-v4-flash"}}}}
{"type":"tool/call","data":{"turn":1,"step":1,"callId":"auto-review-denied-call","name":"mystery"}}
{"type":"tool/result","data":{"turn":1,"step":1,"message":{"content":[{"type":"tool-result","toolCallId":"auto-review-denied-call","isError":true}]},"error":{"code":"AUTO_REVIEW_DENIED"}}}
{"type":"step/end","data":{"turn":1,"step":1}}

这段日志的出处见6。从中可以看到:

  • 会话头记录了 agent preset 和委派深度;
  • 沙箱模式、审批策略这类“设置”同样是日志事件;
  • 被审核拒绝的工具调用会留下 tool/call 和带错误码的 tool/result,完整可审计。

2.3 会话格式的演进#

要点说明
文件名v0 是 session.jsonl[.zstd];v1 及之后的版本是 session.vN.jsonl[.zstd]2
迁移每个相邻的迁移包只负责一步 vN → vN+1。仓库里目前有 v0-to-v1 到 v3-to-v4 共四个迁移包27
不可变已经提交的 generation 文件,路径绝不会被重命名、替换或删除。写入时会在原文件旁边独占地发布一个新的后继版本2
后端视角

这就是事件溯源(event sourcing)的完整实践:只追加的日志、投影(projection)、版本化的 schema 加上逐步迁移、不可变的历史。与 pi 的树状会话相比,DSH 在格式演进和迁移方面做得更工程化。

2.4 日志驱动的测试:不需要 API Key 的回放#

  • 快照测试(pnpm run test:snapshot):录制好的会话既提供用户输入和模型回放,又作为持久化结果的预期值。test:snapshot:record 用于重新录制,test:snapshot:refresh 用于在回放输入仍然有效时刷新预期值8。
  • Web 浏览器快照测试在 CI 中被强制设为只读的 DSH_SNAPSHOT=replay8。
  • 其他测试层级包括:单元测试(每个注册表都有 HMR 安全测试)、按文件 100% 的覆盖率门禁、带密钥的真实 API e2e 测试(缺少密钥时自动跳过)、性能基准测试8。

3. 工具执行流水线#

下图根据官方的 工具执行流水线 图做了简化3:

allow

ask

deny

允许这一次

拒绝、取消或不可用

allow

deny

assistant 消息里包含工具调用

会话事件 tool/call

(在执行之前就记录下来)

tools/pre-execute(waterfall)

hooks、权限、沙箱

单调守卫

(只能拒绝或弃权,后续无法撤销)

ctx.approval 一次性审批

拿不到审批时默认拒绝

拒绝:不执行工具主体

tools/execute(waterfall)

超时、重试、指标

工具的 execute()

fs/write-intent 等文件守卫

projectContent

tools/post-execute(waterfall)

接受、拦截、替换、追加上下文

finalizeContent

(最后一道只针对内容的不变式)

tools/result 通知

(冻结后的权威结果)

会话事件 tool/result

allow

ask

deny

允许这一次

拒绝、取消或不可用

allow

deny

assistant 消息里包含工具调用

会话事件 tool/call

(在执行之前就记录下来)

tools/pre-execute(waterfall)

hooks、权限、沙箱

单调守卫

(只能拒绝或弃权,后续无法撤销)

ctx.approval 一次性审批

拿不到审批时默认拒绝

拒绝:不执行工具主体

tools/execute(waterfall)

超时、重试、指标

工具的 execute()

fs/write-intent 等文件守卫

projectContent

tools/post-execute(waterfall)

接受、拦截、替换、追加上下文

finalizeContent

(最后一道只针对内容的不变式)

tools/result 通知

(冻结后的权威结果)

会话事件 tool/result

扩展点用来做什么
tools/pre-execute可扩展的允许、拒绝、询问策略。例如 Claude Code 或 Codex 的 hooks 桥接就在这里接入
ctx.tools.guard()最终的单调拒绝,后面的监听器无法撤销
tools/execute为工具分发加上截止时间、重试、指标采集
tools/post-execute替换展示内容或返回值,拦截结果,追加模型可见的上下文
tools/result只读地观察不可变的规范化结果

以上取自 cookbook 的“执行策略与观测”一节9。

为什么要分这么多层

官方的设计意图是:hooks 可以跨越不同的工具系列,而不必让工具与某个策略服务耦合3。工具作者只管实现 execute;权限、审批、沙箱、超时、审计都由流水线上的其他插件负责,而这些插件全部可以替换。


4. Fail-closed 沙箱#

模式效果
read-only禁止写入(/dev/null 这类必要的输出目标除外)
workspace-write允许写入工作区根目录,以及后端定义的临时区域
danger-full-access不做隔离,消费方直接启动原始命令

以上取自 dsh-sandbox 的 README4。

  • 本地后端:bwrap、通过 npm 分发的 landlock-run 启动器、macOS Seatbelt,以及 Windows ACL 受限令牌。启动前会做功能探测,失败时关闭而不是降级7。
  • 强制执行完整度会逐次调用报告:full 表示后端管辖了这个模式承诺的所有文件操作;partial 表示只管辖了一部分。Windows ACL 和较旧的 Landlock ABI 属于 partial4。
  • 被拒绝后的升权:受限调用被拒绝时会返回 [sandbox: file access denied under <mode> mode]。模型可以带上 sandbox_permissions(足以放行的最窄的更宽模式)和 justification(理由),原样重试同一个调用一次,由审批服务向人征求同意4。
  • Fail-closed:没有任何后端能强制执行所请求的模式时,调用以 SANDBOX_UNAVAILABLE 失败,而不会在不受限的状态下运行4。

沙箱本身也是用配置组合出来的,下面是 README 里的组合示例4:

- id: sandbox
name: '@deepseek-ai/dsh-sandbox-local' # 按平台选择的后端提供方(ctx.sandbox)
- id: sandbox-policy
name: '@deepseek-ai/dsh-sandbox-policy' # 部署默认的模式,以及 workspace-write 的根目录
config:
mode: workspace-write
workspaceRoot: !!js process.cwd()
- id: bash
name: '@deepseek-ai/dsh-bash-sandbox' # ctx.shell 背后的受限执行器
沙箱不是万能的

官方的安全说明强调:沙箱、审批提示和权限控制可以降低风险,但不保证隔离;即使限制得到正确执行,也无法保护本项目被允许访问的那些资源。不要把 DSH 当作不可信工作负载唯一的安全控制措施10。


5. 动手:写一个工具插件(✅ 已实际运行)#

5.1 官方约定的最小形态#

下面是官方 cookbook 工具编写参考 中的最小示例9,它对应仓库源码的 API:

// 片段(no-check):仓库源码的 API。npm 上的 rc 包类型可能略有不同,见第 8 节
import { readFile } from 'node:fs/promises'
import type { Context } from '@deepseek-ai/cordis'
import { defineTool } from '@deepseek-ai/dsh-tools'
export const name = 'my-tool'
export const inject = ['tools']
export function apply(ctx: Context) {
ctx.tools.register(defineTool({
name: 'read_file',
description: 'Read a file from disk.', // 模型看到的说明
parameters: {
path: { type: 'string', required: true, description: 'Absolute path' },
limit: { type: 'number' }, // 默认是可选参数
},
output: {
schema: { type: 'string' },
render: (_args, value) => [{ type: 'text', text: value }],
},
async execute(args, exec) {
// args 的类型由 schema 推导出来:{ path: string; limit?: number }
return readFile(args.path, { encoding: 'utf8', signal: exec.signal })
},
}))
}

execute() 的几条约定9:

约定说明
参数已经校验过execute 运行之前,defineTool 已经按 schema 校验了模型给出的参数。schema 表达不了的约束(比如非空、正数、跨字段规则)仍需自己检查
返回规范的 JSON 值execute 只返回值本身,由 output.render 把它渲染成模型能看到的内容。不要让调用方从自然语言里解析 id 或字段
抛出异常即 isError基础设施故障时直接抛出异常,注册表会统一转换成错误结果
遵守 exec.signal信号触发时要取消正在进行的工作
异步通知用 exec.agent.inject(...) 追加持久化的上下文,下一次模型请求就能看到
注册即副作用插件 fiber 被释放时,工具会自动注销;工具的 schema 会自动进入系统提示词的组装
PTC 模式会自动用到你的工具

在 PTC(Programmatic Tool Calling)模式下,每个可见的已注册工具都可以通过 await tools.<name>(args) 调用,不需要额外集成。调用会重新进入正常的执行流水线,返回值是经过策略处理后的规范 JSON 值9。这和 OpenAI 的 Programmatic Tool Calling、pi 的 codemode 是同一种思路。

5.2 完整可运行版:把工具接入真实的 harness 服务#

下面的例子改编自官方 Cordis 教程第 7 章11。它会向 harness 的 tools 服务注册一个工具,然后通过真实的执行流水线调用它,整个过程不需要 API Key,也不调用模型。

// greet-tool.ts(已按 npm 0.0.1-rc 包的 API 调整,并实际运行过)
import type { Context } from '@deepseek-ai/cordis'
import { defineTool } from '@deepseek-ai/dsh-tools'
import type { CallId } from '@deepseek-ai/dsh-llm'
export const name = 'greet-tool'
export const inject = ['tools']
export function apply(ctx: Context) {
ctx.tools.register(defineTool({
name: 'greet',
description: 'Greet the named person.',
parameters: {
name: { type: 'string', required: true, description: 'Who to greet' },
},
output: {
schema: { type: 'string' },
render: (_args, value) => [{ type: 'text', text: value }],
},
async execute(args) {
return `Hello, ${args.name}!`
},
}))
// 代替模型发起一次调用,走真实的执行流水线
void (async () => {
const result = await ctx.tools.execute({
callId: 'demo-1' as CallId, // branded 类型只存在于编译期
name: 'greet',
arguments: { name: 'Cordis' },
signal: new AbortController().signal,
})
console.log('tool replied:', JSON.stringify(result.content))
})()
}
// tool-logger.ts:一个独立的观察插件,通过 tools/result 事件观察所有工具调用
import type { Context } from '@deepseek-ai/cordis'
import type {} from '@deepseek-ai/dsh-tools'
export const name = 'tool-logger'
export const inject = ['tools']
export function apply(ctx: Context) {
ctx.on('tools/result', (exec, result) => {
const text = result.content
.map(block => (block.type === 'text' ? block.text : ''))
.join('')
console.log(`[tool-logger] ${exec.name} -> ${text}`)
})
}
# cordis.yml:dsh-tools 会 inject systemPrompt,所以要同时列出它的提供方
- name: '@deepseek-ai/dsh-system-prompt'
- name: '@deepseek-ai/dsh-tools'
- name: './tool-logger.ts'
- name: './greet-tool.ts'
$ node --import tsx node_modules/@deepseek-ai/cordis/bin.js
[tool-logger] greet -> Hello, Cordis!
tool replied: [{"type":"text","text":"Hello, Cordis!"}]

两点观察:

  • logger 的输出先出现,因为 tools/result 是在结果物化的过程中发出的,早于 execute 返回的 promise 兑现11。
  • 两个插件互不知道对方的存在,它们是通过注册表服务和事件连接起来的11。
从这里到完整的 agent

一个真实的 agent 就是这套组合,再加上 LLM 适配器、agent loop、持久化和应用入口这几个插件11。把 greet-tool.ts 用一个小的 --patch overlay 加进你的 profile,它就成了 agent 可以调用的工具。


6. 新功能应该放在哪:扩展点地图#

你想要机制
添加模型提供方在 ctx.llm 上注册适配器
添加模型可以调用的能力在 ctx.tools 上注册,它的 schema 会自动进入提示词组装
让某个会话拥有不同的能力集合组装一个 agent preset;其中的服务行需要 isolate
添加 shell 执行方式注册一个 ctx.shell 后端
限制所启动的进程使用 ctx.sandbox 后端
拦截请求、工具调用或轮次使用相应的 agent/* 或 tools/* 事件
添加模型可见的上下文调用 agent.inject(),它会落到下一次被接纳的请求中
添加持久化的会话状态扩展 SessionEventMap,从日志中渲染和回放
从外部 webhook 启动会话在 ctx.webhookRuntime 上注册可信规则
在轮次边界 fork 会话ctx.agents.create({ sessionId, seed, meta: { parentSession, seedLength } })
用新的后端存储会话实现 SessionPersistence(create、open、stat、list、export)

以上节选自架构文档的“新行为的归属位置”一节2。


7. 横向看:DSH 的工程纪律#

笔者从官方文档中归纳出的几条“纪律”,后端工程师可以直接借鉴:

  1. 事件即扩展点:持久的事实用会话事件,实时协调用 agent/*,能力策略用 fs/*、tools/*2。
  2. 模型可见即已记录:调试、回放、fork 都建立在同一份日志上2。
  3. fail-closed:沙箱、审批这些地方在“做不到”或“拿不到答复”时,一律拒绝47。
  4. 所有注册都可逆:HMR、热插拔不会留下任何残余(见 5.2 Cordis 与一切皆插件)。
  5. 测试即回放:录制下来的会话就是测试用例,CI 不需要密钥8。

8. 踩坑记录:npm 包落后于仓库源码#

笔者在 2026-10-09 用 npm 上的包运行官方教程,发现以下几处与仓库源码(commit 5badb15)不一致:

现象仓库源码的写法npm rc 包的实际情况解决办法
brandString 找不到教程从 @deepseek-ai/dsh-brand 导入 brandString11dsh-brand 0.0.1-rc.5 的运行时模块是空的(export {})branded 类型只存在于编译期,直接用类型断言 'demo-1' as CallId
ToolCallId 类型不存在import type { ToolCallId } from '@deepseek-ai/dsh-llm'11dsh-llm 0.0.1-rc.1 里的类型名叫 CallId改用 CallId
DeepSeekHarness 不支持 profileREADME 写的是 new DeepSeekHarness({ profile: 'sdk', ... })120.0.1-rc.1 要求显式传入 launch: { command, args }看对应版本 npm 包自带的 README
插件加载失败时什么都不输出教程称会打印错误,并以状态码 1 退出13不挂 @deepseek-ai/cordis-plugin-logger-console 就看不到错误,退出码为 0调试时先挂上 logger-console
通用的排查思路
  1. 插件“没有反应”:先挂上 logger-console,查看错误;再检查是否有插件停在 PENDING,也就是 inject 的服务没人提供。教程第 6 章给了一个列出所有 PENDING fiber 的诊断插件14。
  2. 示例代码报类型错误:先确认你看的文档和你安装的包是不是同一个版本。

小结#

  • DSH 的三大核心机制:事件溯源的会话日志、分层的工具执行流水线、fail-closed 沙箱。它们都是 Cordis 插件,都可以替换。
  • 写工具插件只需要三步:inject: ['tools'],加上 defineTool({...}),再在 cordis.yml 里挂载它。权限、沙箱、审计都交给流水线上的其他插件。
  • 由于项目处于预览阶段,版本差异是最常见的坑。动手前先确认文档和你安装的包是同一个版本。

相关笔记#

参考资料#

注释与出处#

  1. deepseek-ai/deepseek-harness,README.zh.md(commit 5badb15),https://github.com/deepseek-ai/deepseek-harness/blob/5badb15009ae1756c3afe0ae0cef1faafc290ccc/README.zh.md ↩

  2. deepseek-ai/deepseek-harness,docs/architecture.zh.md(轮次流程、会话日志、能力 seam、新行为的归属位置等节),https://github.com/deepseek-ai/deepseek-harness/blob/5badb15009ae1756c3afe0ae0cef1faafc290ccc/docs/architecture.zh.md ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12

  3. deepseek-ai/deepseek-harness,docs/tool-execution-pipeline.zh.md,https://github.com/deepseek-ai/deepseek-harness/blob/5badb15009ae1756c3afe0ae0cef1faafc290ccc/docs/tool-execution-pipeline.zh.md ↩ ↩2 ↩3

  4. deepseek-ai/deepseek-harness,packages/sandbox/sandbox/README.zh.md(模式与强制执行、被拒绝的调用与升权、故障关闭行为等节),https://github.com/deepseek-ai/deepseek-harness/blob/5badb15009ae1756c3afe0ae0cef1faafc290ccc/packages/sandbox/sandbox/README.zh.md ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7

  5. deepseek-ai/deepseek-harness,docs/agent-lifecycle.zh.md,https://github.com/deepseek-ai/deepseek-harness/blob/5badb15009ae1756c3afe0ae0cef1faafc290ccc/docs/agent-lifecycle.zh.md ↩

  6. deepseek-ai/deepseek-harness,snapshots/web/auto-review-denial/session.v3.jsonl,https://github.com/deepseek-ai/deepseek-harness/blob/5badb15009ae1756c3afe0ae0cef1faafc290ccc/snapshots/web/auto-review-denial/session.v3.jsonl ↩

  7. deepseek-ai/deepseek-harness,packages/ 下各包 package.json 的 description 字段(sandbox-local、user-approval、session-format-v* 等),https://github.com/deepseek-ai/deepseek-harness/tree/5badb15009ae1756c3afe0ae0cef1faafc290ccc/packages ↩ ↩2 ↩3

  8. deepseek-ai/deepseek-harness,docs/testing.zh.md,https://github.com/deepseek-ai/deepseek-harness/blob/5badb15009ae1756c3afe0ae0cef1faafc290ccc/docs/testing.zh.md ↩ ↩2 ↩3 ↩4

  9. deepseek-ai/deepseek-harness,docs/cookbook/adding-a-tool.zh.md,https://github.com/deepseek-ai/deepseek-harness/blob/5badb15009ae1756c3afe0ae0cef1faafc290ccc/docs/cookbook/adding-a-tool.zh.md ↩ ↩2 ↩3 ↩4

  10. deepseek-ai/deepseek-harness,SAFETY.zh.md,https://github.com/deepseek-ai/deepseek-harness/blob/5badb15009ae1756c3afe0ae0cef1faafc290ccc/SAFETY.zh.md ↩

  11. deepseek-ai/deepseek-harness,docs/cordis-tutorial/07-into-the-harness.zh.md,https://github.com/deepseek-ai/deepseek-harness/blob/5badb15009ae1756c3afe0ae0cef1faafc290ccc/docs/cordis-tutorial/07-into-the-harness.zh.md ↩ ↩2 ↩3 ↩4 ↩5 ↩6

  12. deepseek-ai/deepseek-harness,packages/sdk/client/README.zh.md,https://github.com/deepseek-ai/deepseek-harness/blob/5badb15009ae1756c3afe0ae0cef1faafc290ccc/packages/sdk/client/README.zh.md ↩

  13. deepseek-ai/deepseek-harness,docs/cordis-tutorial/05-config.zh.md,https://github.com/deepseek-ai/deepseek-harness/blob/5badb15009ae1756c3afe0ae0cef1faafc290ccc/docs/cordis-tutorial/05-config.zh.md ↩

  14. deepseek-ai/deepseek-harness,docs/cordis-tutorial/06-composition-and-hmr.zh.md,https://github.com/deepseek-ai/deepseek-harness/blob/5badb15009ae1756c3afe0ae0cef1faafc290ccc/docs/cordis-tutorial/06-composition-and-hmr.zh.md ↩

输入关键词开始搜索。多个关键词用空格分隔。