第 4 章:工具抽象与执行引擎

工具是思维的延伸;Agent 的能力边界,取决于它能调用什么。


4.1 核心问题

前三章搭起了进程骨架(第1章)、渲染引擎(第2章)、命令路由(第3章)。但我们的 Agent 还什么"实事"都做不了——它只是一个能接受用户输入的空壳。

真正的 Agent 需要行动能力:读写文件、运行命令、搜索代码……这些能力统称工具(Tools)

本章要解决的问题是:

如何设计一套统一的工具接口,让 LLM 能够"调用"任意工具,同时保证类型安全、权限可控、错误可追踪?

这个问题并不简单。LLM 通过 JSON 告诉我们"我想调用哪个工具、传入什么参数",而我们需要:

  1. 验证:JSON 里的参数是否符合工具期望的类型?
  2. 执行:找到对应工具并调用它。
  3. 封装:执行结果(或错误)以统一的格式返回给 LLM。

Claude Code 的解法是:用 Zod 作为输入 schema 的单一来源——它既是运行时的类型守卫,又能通过反射生成 JSON Schema 发送给 LLM,实现了"告诉模型工具长什么样"与"运行时校验模型发回的参数"的双重复用。


4.2 原理讲解

4.2.1 Function Calling:LLM 的"手"

OpenAI 在 2023 年引入了 Function Calling(Anthropic 称为 Tool Use),其本质是一次 API 往返的扩展协议:

用户                  LLM                    我们的代码
 │                     │                        │
 │── 消息 ────────────>│                        │
 │                     │── JSON工具调用 ────────>│
 │                     │                        │── 执行工具
 │                     │<── 工具结果 ───────────│
 │<── 最终回复 ─────── │                        │

LLM 在第一次响应时不直接回答用户,而是输出一个 tool_use 块:

{
  "type": "tool_use",
  "id": "toolu_01abc",
  "name": "Bash",
  "input": {
    "command": "ls -la",
    "description": "List files in current directory"
  }
}

我们的代码解析这个块,执行 ls -la,再把结果用 tool_result 块回传给 LLM,LLM 才最终生成自然语言回复。

这套协议意味着:工具的"输入规格"(即 JSON Schema)必须随系统提示一起发送给 LLM,让模型知道每个工具接受哪些参数。

4.2.2 Zod:一份定义,双重用途

Zod 是 TypeScript 生态中最流行的运行时验证库。其核心思想:

import { z } from 'zod'

// 定义 schema
const BashInputSchema = z.object({
  command: z.string().describe('The command to execute'),
  timeout: z.number().optional().describe('Timeout in milliseconds'),
})

// 用途一:运行时验证(类型守卫)
const result = BashInputSchema.safeParse(unknownInput)
if (!result.success) {
  console.error(result.error.issues)
}

// 用途二:TypeScript 类型推断
type BashInput = z.infer<typeof BashInputSchema>
// 等价于:{ command: string; timeout?: number }

Claude Code 的精妙在于:同一份 Zod schema 同时服务于三个目标

目标 机制
告诉 LLM 工具参数格式 z.toJSONSchema(schema) 生成 JSON Schema 发给 API
验证 LLM 返回的参数 schema.safeParse(input) 运行时类型守卫
TypeScript 类型推断 z.infer<typeof schema> 编译期类型安全

4.2.3 Tool 接口三要素

去掉 Claude Code 生产级代码中的渲染、权限、UI 等横切关注点,工具的核心骨架只有三要素:

┌─────────────────────────────────────────┐
│              Tool 接口                   │
│                                         │
│  name         "Bash" / "Read" / ...     │
│  inputSchema  Zod schema(双重用途)      │
│  call()       异步执行函数               │
└─────────────────────────────────────────┘

名称(name):字符串标识符,LLM 在 tool_use 块的 name 字段中引用它。全局唯一。

输入 schema(inputSchema):Zod schema 对象。负责:① 生成 JSON Schema 给 LLM;② 在 call() 前校验入参。

执行函数(call()):接受已验证的类型安全入参,返回 Promise<ToolResult>

4.2.4 lazySchema:惰性初始化

Claude Code 源码中大量使用了 lazySchema 包装:

const inputSchema = lazySchema(() => z.strictObject({
  command: z.string().describe('The command to execute'),
  // ...
}))

这是因为 Zod 的 .describe() 等链式调用会在模块加载时执行,而部分描述文字依赖运行时配置(如超时限制的动态数值)。lazySchema 将 schema 构造推迟到首次访问,避免启动时的副作用。

4.2.5 buildTool:填充默认值的工厂函数

真实工具除了三要素,还需要大量"样板"方法:isEnabled()isConcurrencySafe()isReadOnly()checkPermissions() 等。

Claude Code 用 buildTool() 工厂函数为这些方法提供安全的默认值:

// src/Tool.ts
const TOOL_DEFAULTS = {
  isEnabled: () => true,
  isConcurrencySafe: () => false,   // 默认不并发安全(保守)
  isReadOnly: () => false,           // 默认会写(保守)
  isDestructive: () => false,
  checkPermissions: (input, ctx) =>
    Promise.resolve({ behavior: 'allow', updatedInput: input }),
  toAutoClassifierInput: () => '',
  userFacingName: () => '',
}

export function buildTool<D extends AnyToolDef>(def: D): BuiltTool<D> {
  return {
    ...TOOL_DEFAULTS,
    userFacingName: () => def.name,  // 用工具名作为默认用户可见名
    ...def,                           // 工具自定义实现优先级最高
  } as BuiltTool<D>
}

设计亮点:所有默认值都是保守的(fail-closed):不假设并发安全、不假设只读——这确保了即使工具实现者忘记声明某个能力,系统也会选择更安全的行为。

4.2.6 错误封装:工具不抛异常

工具执行失败时,正确的做法是返回结构化错误,而不是 throw。原因:LLM 需要读取错误信息来调整后续行为;而异常会打断整个对话循环。

// 错误模式(不推荐)
async call(input) {
  if (!existsSync(input.file_path)) {
    throw new Error('File not found')  // ❌ 打断对话循环
  }
}

// 正确模式
async call(input) {
  if (!existsSync(input.file_path)) {
    return {
      type: 'text',
      text: `File not found: ${input.file_path}`,
    }  // ✅ LLM 可以读取并重试
  }
}

4.2.7 工具注册表:tools.ts

src/tools.ts 是所有工具的装配中心。它根据编译特性标志(feature('...'))和运行时环境变量(process.env.USER_TYPE)动态决定哪些工具参与注册:

// src/tools.ts(精简)
import { BashTool } from './tools/BashTool/BashTool.js'
import { FileReadTool } from './tools/FileReadTool/FileReadTool.js'
import { FileWriteTool } from './tools/FileWriteTool/FileWriteTool.js'

// 编译时特性门控
const SleepTool =
  feature('PROACTIVE') || feature('KAIROS')
    ? require('./tools/SleepTool/SleepTool.js').SleepTool
    : null

// Ant 员工专属工具
const REPLTool =
  process.env.USER_TYPE === 'ant'
    ? require('./tools/REPLTool/REPLTool.js').REPLTool
    : null

这种"按需装配"的设计实现了功能开关(Feature Flag) 的工具级别控制,公开版和内部版可以共用同一套代码库。


4.3 Claude Code 源码细节

4.3.1 Tool 类型:完整接口(src/Tool.ts)

生产级 Tool 类型远比骨架复杂,核心字段如下:

// src/Tool.ts(精简注解版)
export type Tool<
  Input extends AnyObject = AnyObject,
  Output = unknown,
  P extends ToolProgressData = ToolProgressData,
> = {
  readonly name: string

  /** Zod schema,双重用途:JSON Schema 生成 + 运行时校验 */
  readonly inputSchema: Input

  /** 可选:直接提供 JSON Schema(用于 MCP 动态工具) */
  readonly inputJSONSchema?: ToolInputJSONSchema

  /** 核心执行函数 */
  call(
    args: z.infer<Input>,
    context: ToolUseContext,
    canUseTool: CanUseToolFn,
    parentMessage: AssistantMessage,
    onProgress?: ToolCallProgress<P>,
  ): Promise<ToolResult<Output>>

  /** 生成工具描述(异步,可依赖运行时状态) */
  description(
    input: z.infer<Input>,
    options: { isNonInteractiveSession: boolean; toolPermissionContext: ToolPermissionContext; tools: Tools },
  ): Promise<string>

  /** 权限检查:在 call() 前调用 */
  checkPermissions(input: z.infer<Input>, context: ToolUseContext): Promise<PermissionResult>

  /** 是否只读(影响权限提示措辞) */
  isReadOnly(input: z.infer<Input>): boolean

  /** 是否并发安全(影响并发调度) */
  isConcurrencySafe(input: z.infer<Input>): boolean

  /** 是否破坏性操作(删除、覆写、发送) */
  isDestructive?(input: z.infer<Input>): boolean

  /** 用户中断时的行为:取消 or 等待 */
  interruptBehavior?(): 'cancel' | 'block'

  /** 工具结果超过此字节数时持久化到磁盘 */
  maxResultSizeChars: number

  // … 以及大量 UI 渲染方法(renderToolUseMessage、renderToolResultMessage 等)
}

关键观察call() 的签名比我们想象的复杂。除了 argscontext,还有:

  • canUseTool:运行时权限查询函数,用于嵌套工具调用(如 AgentTool 内部再调用其他工具)
  • parentMessage:触发本次工具调用的 AssistantMessage,用于关联上下文
  • onProgress:流式进度回调,BashTool 用它实时推送命令输出

4.3.2 BashTool 的 Schema 定义(src/tools/BashTool/BashTool.tsx)

const fullInputSchema = lazySchema(() => z.strictObject({
  command: z.string().describe('The command to execute'),
  timeout: semanticNumber(z.number().optional())
    .describe(`Optional timeout in milliseconds (max ${getMaxTimeoutMs()})`),
  description: z.string().optional()
    .describe(`Clear, concise description of what this command does...`),
  run_in_background: semanticBoolean(z.boolean().optional())
    .describe('Set to true to run this command in the background.'),
  // 内部字段,不暴露给 LLM(通过 .omit() 从对外 schema 中剔除)
  _simulatedSedEdit: z.object({
    filePath: z.string(),
    newContent: z.string()
  }).optional(),
}))

// 对外 schema 剔除内部字段
const inputSchema = lazySchema(() =>
  fullInputSchema().omit({ _simulatedSedEdit: true })
)

安全设计_simulatedSedEdit 是 sed 编辑预览通过的内部传递字段,绝不能暴露给 LLM——否则模型可以绕过权限检查直接写任意文件。omit() 确保这个字段从 JSON Schema 中消失。

4.3.3 ToolUseContext:工具的运行时上下文

每次工具调用都会收到一个 ToolUseContext,它是工具与外部世界交互的通道:

// src/Tool.ts(精简)
export type ToolUseContext = {
  options: {
    commands: Command[]
    debug: boolean
    mainLoopModel: string
    tools: Tools            // 当前可用工具列表
    verbose: boolean
    mcpClients: MCPServerConnection[]
    maxBudgetUsd?: number   // 预算上限(用于成本控制)
    // ...
  }
  abortController: AbortController  // 取消信号
  readFileState: FileStateCache      // 文件读取缓存
  getAppState(): AppState
  setAppState(f: (prev: AppState) => AppState): void
}

ToolUseContext依赖注入的载体——工具不直接访问全局状态,而是通过 context 获取所需资源,这使得工具在测试中可以轻松 mock。

4.3.4 ToolResult:统一的返回格式

// src/Tool.ts
export type ToolResult<T> = {
  type: 'tool_result'
  content: T
  // 当 isError 为 true 时,content 是错误描述文本
  isError?: boolean
}

isError: true 时,LLM 会收到错误信息并尝试修正;isError: false(默认)时,内容被视为成功结果。


4.4 最小化产出物

代码骨架位于 ../chapters/04/src/,参考实现位于 ../chapters/04/solution/

本章要实现什么

../chapters/04/src/tools.ts 中完成工具引擎。

接口规范(已提供,不要修改):

export type ToolResult =
  | { success: true; output: string }
  | { success: false; error: string }

export type ToolContext = { cwd: string; verbose: boolean }

export interface Tool<Input extends z.ZodTypeAny = z.ZodTypeAny> {
  name: string
  description: string
  inputSchema: Input
  call(input: z.infer<Input>, ctx: ToolContext): Promise<ToolResult>
}

export const TOOL_REGISTRY: Map<string, Tool<any>>
export function callTool(name: string, rawInput: unknown, ctx: ToolContext): Promise<ToolResult>
export function getToolSchemas(): Array<{ name: string; description: string; input_schema: {...} }>

你需要实现

  1. BashTool.call():用 execSync 执行命令,捕获错误返回 { success: false, error }(不 throw)
  2. ReadFileTool.call():读取文件,支持 start_line/end_line 行范围,每行加 " N │ " 前缀
  3. WriteFileTool.call():写入文件,失败时返回错误(不 throw)
  4. TOOL_REGISTRY:用 Map 注册三个工具
  5. callTool():查找工具 → safeParse 校验 → 执行,全程不抛异常
  6. getToolSchemas():遍历注册表,从 Zod schema 提取 JSON Schema

关键约束

  • 工具执行失败必须返回 { success: false, error } 而不是 throw
  • callTool() 本身也不能抛异常

验收

cd docs/chapters/04
npm install
npm test

卡住时查看 ../chapters/04/solution/tools.ts


4.5 从骨架到生产:Claude Code 做了什么

本章产出物是最小可用的工具引擎。Claude Code 的生产代码在此基础上叠加了以下层次:

关注点 本章(最小版) Claude Code(生产版)
输入校验 schema.safeParse() 同上 + validateInput() 钩子(工具自定义额外校验)
权限控制 checkPermissions() → 权限对话框 → alwaysAllow 规则
执行取消 AbortController 贯穿整个调用链
进度回调 onProgress() 流式推送(BashTool 实时输出)
结果存储 内存返回 超过 maxResultSizeChars 时持久化到磁盘
并发控制 isConcurrencySafe() 决定是否可并行调用
UI 渲染 renderToolUseMessage() + renderToolResultMessage()
工具搜索 全量注册 shouldDefer + ToolSearchTool 按需加载
MCP 集成 inputJSONSchema 支持非 Zod 的外部工具

每一层都在第5-16章中逐步引入。


4.6 本章小结

本章构建了工具执行引擎的核心骨架:

  1. Zod schema 是工具的"合同":一份 schema 定义同时承担 JSON Schema 生成(告知 LLM)和运行时校验(验证 LLM 响应)两项职责。

  2. Tool 接口三要素name(标识符)、inputSchema(合同)、call()(实现)。Claude Code 的完整 Tool 类型在此基础上扩展了权限、UI、进度等切面。

  3. buildTool() 思想:用工厂函数提供保守的安全默认值,工具实现者只需关注业务逻辑,不必关心基础设施样板。

  4. 工具不抛异常:所有失败以结构化数据(ToolResult)返回,确保对话循环的健壮性——这是让 LLM 能够从错误中自我恢复的关键设计。

  5. 注册表 + 调度callTool() 通过注册表查找工具、验证输入、执行调用,实现了调度逻辑与工具实现的完全解耦。

  6. tools.ts 独立导出:所有类型、工具实例、注册表、callTool()getToolSchemas() 均从 tools.ts 导出,供第 5-16 章直接 import 复用,不需要重复实现。

下一章将给这个工具引擎插上"大脑":接入 Claude API,处理流式响应,让 LLM 真正开始调用这些工具。


下一章

第 5 章:LLM API 客户端(流式)