第 7 章:工具调用循环(Agentic Loop)

自主性的本质:给 LLM 一个工具,它会用你没预料到的方式完成任务。


7.1 核心问题

前六章打下了所有基础:

  • 第 4 章:工具的抽象与执行(Tool 接口、buildTool()
  • 第 5 章:流式 LLM 调用(callModel()
  • 第 6 章:多轮消息历史管理

但这三块还是独立的零件。现在要把它们焊接在一起:让 LLM 自主决定何时调用工具、调用哪个工具、用结果继续思考,直到完成任务。

这就是 Agentic Loop —— Agent 的心跳。

单轮对话是:用户问 → LLM 答
Agentic Loop 是:用户问 → LLM 决策 → 调用工具 → 看结果 → 再决策 → 调用工具 → ... → 给出最终答案

一个典型的 "列出项目中所有 TODO" 请求,背后可能发生:

用户:找出项目里所有未完成的 TODO
  → LLM:我来用 grep 搜索
  → 调用 Bash("grep -r 'TODO' src/")
  → 结果:找到 12 处,其中 3 处在 auth/ 目录
  → LLM:让我读一下具体文件看看
  → 调用 Read("src/auth/login.ts")
  → 结果:文件内容
  → LLM:好的,我整理一下
  → 最终回复:[结构化的 TODO 清单]

每一次工具调用都是一次独立的 API 往返。Agentic Loop 的核心职责就是:协调这些往返,管理状态,处理错误,直到 LLM 说"我完成了"。


7.2 原理讲解

7.2.1 Loop 的驱动信号:needsFollowUp

一个常见误解:Agentic Loop 通过检测 stop_reason === "tool_use" 来决定是否继续。

Claude Code 的注释明确写道:

// Note: stop_reason === 'tool_use' is unreliable -- it's not always set correctly.
// Set during streaming whenever a tool_use block arrives — the sole loop-exit signal.
let needsFollowUp = false

真正的驱动信号是 needsFollowUp ——一个布尔值,在 流式响应中每发现一个 tool_use 块就设为 true

// 流式处理 assistant 消息时:
if (msgToolUseBlocks.length > 0) {
  toolUseBlocks.push(...msgToolUseBlocks)
  needsFollowUp = true  // ← Loop 继续的信号
}

Loop 的结构是:

while (true) {
  调用 API(流式)
  收集 assistantMessages 和 toolUseBlocks
  
  if (!needsFollowUp) {
    // LLM 没有要调用的工具,任务完成(或中途退出)
    break  // return { reason: 'completed' }
  }
  
  // 执行所有 toolUseBlocks → 得到 toolResults
  // 把 assistantMessages + toolResults 追加到 messages
  // 继续下一轮
}

7.2.2 并行工具调用

同一轮 LLM 响应可以包含多个 tool_use 块,例如:

{
  "role": "assistant",
  "content": [
    { "type": "text", "text": "我同时读这三个文件:" },
    { "type": "tool_use", "id": "tu_a", "name": "Read", "input": {"file_path": "a.ts"} },
    { "type": "tool_use", "id": "tu_b", "name": "Read", "input": {"file_path": "b.ts"} },
    { "type": "tool_use", "id": "tu_c", "name": "Bash", "input": {"command": "ls src/"} }
  ]
}

Claude Code 通过 runTools() 并行执行所有工具(使用 Promise.all 语义),然后将全部结果打包为一条 user 消息:

{
  "role": "user",
  "content": [
    { "type": "tool_result", "tool_use_id": "tu_a", "content": "..." },
    { "type": "tool_result", "tool_use_id": "tu_b", "content": "..." },
    { "type": "tool_result", "tool_use_id": "tu_c", "content": "..." }
  ]
}

重要:每个 tool_use 必须有对应的 tool_result,哪怕工具执行失败。失败的结果通过 is_error: true 标记,内容是错误消息。这避免了 API 因 tool_use/tool_result 不配对而拒绝请求。

7.2.3 流式工具执行(Streaming Tool Execution)

Claude Code 有一个关键优化:StreamingToolExecutor

传统做法是:等 LLM 全部回复完毕后,再执行工具。
StreamingToolExecutor 的做法是:边接收流边并行启动工具执行——

流式接收:  [text_delta][text_delta][tool_use block 1 complete!][text_delta][tool_use block 2 complete!]
                                          ↓                                        ↓
工具执行:                          立即启动 tool_1                        立即启动 tool_2

工具在 LLM 还在"打字"的同时就开始执行了。当流结束时,工具可能已经完成,大幅降低延迟。

// 每收到一个完整的 tool_use 块,立即注册执行
streamingToolExecutor.addTool(toolBlock, message)

// 流结束后,收集所有(已完成或还在跑的)结果
for await (const update of streamingToolExecutor.getRemainingResults()) {
  yield update.message  // 工具结果消息
}

7.2.4 循环终止条件

Agentic Loop 有多种退出路径,每种都有对应的 reason

退出原因 reason 触发条件
任务完成 'completed' LLM 回复没有 tool_useneedsFollowUp = false
用户中断 'aborted_streaming' / 'aborted_tools' AbortController 触发
达到轮次上限 'max_turns' turnCount > maxTurns
上下文过长 'prompt_too_long' 压缩失败后仍超长
模型错误 'model_error' API 抛出异常
Hook 阻止 'hook_stopped' / 'stop_hook_prevented' 生命周期 Hook 拦截
上下文阻塞上限 'blocking_limit' Token 超过硬阻塞阈值(禁用自动压缩时)

maxTurns 是防止无限循环的"安全网"——即使 LLM 一直想调用工具,也最多迭代 maxTurns 次。

7.2.5 错误恢复策略

Agentic Loop 有多层容错机制,这是它区别于简单 while(true) 的地方:

错误类型 1:max_output_tokens(输出被截断)

LLM 在说话说到一半时被截断了(写代码写了一半)。Loop 自动注入恢复提示:

const recoveryMessage = createUserMessage({
  content: `Output token limit hit. Resume directly — no apology, no recap of what you were doing.`
           + ` Pick up mid-thought if that is where the cut happened.`,
  isMeta: true,  // 不显示给用户
})
// 继续迭代,最多重试 MAX_OUTPUT_TOKENS_RECOVERY_LIMIT = 3 次

错误类型 2:prompt_too_long(上下文超长)

按优先级尝试三种恢复:

  1. Context Collapse drain(CONTEXT_COLLAPSE 特性):先提交已暂存的压缩
  2. Reactive Compact:触发一次完整的对话历史压缩
  3. 若均失败:展示错误,退出循环

错误类型 3:工具执行失败

工具执行失败不触发 Loop 退出,而是将错误作为 tool_result(标记 is_error: true)反馈给 LLM,让 LLM 自行决定如何处理(重试、换方法或告知用户)。

错误类型 4:模型降级(Model Fallback)

当首选模型过载时,自动切换到降级模型重试当前请求,并清理孤立的 tool_use 块(防止旧 ID 泄露到重试请求中)。

7.2.6 Loop 的完整状态机

┌─────────────────────────────────────────────────────────────────┐
│                         queryLoop(params)                        │
│                                                                  │
│   state = { messages, toolUseContext, turnCount: 1, ... }       │
│                                                                  │
│   while (true) {                                                 │
│     ┌──────────────────────────────────────────────────────┐    │
│     │  PREPARE                                             │    │
│     │  1. getMicrocompact(messages)  → 清理过时工具输出    │    │
│     │  2. getAutoCompact(messages)   → 超限时压缩历史      │    │
│     │  3. buildSystemPrompt()        → 注入上下文          │    │
│     └──────────────────┬───────────────────────────────────┘    │
│                        │                                        │
│     ┌──────────────────▼───────────────────────────────────┐    │
│     │  STREAM API                                          │    │
│     │  for await (message of callModel(messages)) {        │    │
│     │    if tool_use block → toolUseBlocks.push()          │    │
│     │                      → needsFollowUp = true          │    │
│     │                      → streamingToolExecutor.add()   │    │
│     │  }                                                   │    │
│     └──────────────────┬───────────────────────────────────┘    │
│                        │                                        │
│     ┌──────────────────▼───────────────────────────────────┐    │
│     │  DECIDE                                              │    │
│     │  if (!needsFollowUp) {                               │    │
│     │    ├─ 错误恢复?→ state=recovery; continue           │    │
│     │    ├─ Stop hooks?→ 可能 continue 或 return          │    │
│     │    └─ return { reason: 'completed' }  ◄─── 正常退出  │    │
│     │  }                                                   │    │
│     └──────────────────┬───────────────────────────────────┘    │
│                        │ needsFollowUp = true                   │
│     ┌──────────────────▼───────────────────────────────────┐    │
│     │  EXECUTE TOOLS                                       │    │
│     │  runTools(toolUseBlocks) → toolResults               │    │
│     │    ├─ 并行执行所有 tool_use                          │    │
│     │    ├─ 失败 → is_error: true(不退出循环)            │    │
│     │    └─ 中断 → return { reason: 'aborted_tools' }      │    │
│     └──────────────────┬───────────────────────────────────┘    │
│                        │                                        │
│     ┌──────────────────▼───────────────────────────────────┐    │
│     │  UPDATE STATE                                        │    │
│     │  messages = [...messages, ...assistantMsgs,          │    │
│     │              ...toolResults, ...attachments]         │    │
│     │  turnCount++                                         │    │
│     │  if (turnCount > maxTurns) → return 'max_turns'      │    │
│     └──────────────────┬───────────────────────────────────┘    │
│                        │ continue (下一轮)                       │
│                        └─────────────────────────────────────┘  │
└─────────────────────────────────────────────────────────────────┘

7.3 Claude Code 源码细节

7.3.1 query() vs queryLoop() vs QueryEngine

Claude Code 有两套 Agentic Loop 入口,职责分工清晰:

QueryEngine.submitMessage()    ← SDK / 无头模式(程序调用)
  │  处理:会话初始化、System Prompt、slash 命令预处理
  └─► query(params)
        └─► queryLoop(params)   ← 核心循环(REPL 和 SDK 共用)
  • QueryEnginesrc/QueryEngine.ts):高层封装,管理会话生命周期(mutableMessagesreadFileState、权限追踪)。SDK 调用者(如 VS Code 插件、CI 系统)使用这一层。
  • query()src/query.ts):中间层,将 queryLoop 包装为 AsyncGenerator,处理 slash 命令生命周期通知。
  • queryLoop():无状态的核心循环——纯粹的"调用 → 工具 → 调用"状态机。每一轮通过 state = { ... }; continue 更新状态,而不是修改变量。

7.3.2 State 对象:Loop 的可变心脏

queryLoop 通过一个不可变 State 对象在迭代间传递状态:

type State = {
  messages: Message[]            // 当前消息历史
  toolUseContext: ToolUseContext  // 工具执行上下文(含权限、缓存等)
  autoCompactTracking: AutoCompactTrackingState | undefined
  maxOutputTokensRecoveryCount: number  // max_output_tokens 已重试次数
  hasAttemptedReactiveCompact: boolean  // 防止压缩无限循环
  maxOutputTokensOverride: number | undefined
  pendingToolUseSummary: Promise<...> | undefined  // 后台生成工具摘要
  stopHookActive: boolean | undefined
  turnCount: number
  transition: Continue | undefined  // 记录上一轮的转换原因(用于测试断言)
}

"继续迭代"的惯用写法:

// 例:max_output_tokens 恢复
state = {
  messages: [...messagesForQuery, ...assistantMessages, recoveryMessage],
  toolUseContext,
  autoCompactTracking: tracking,
  maxOutputTokensRecoveryCount: maxOutputTokensRecoveryCount + 1,
  // ...其他字段
  transition: { reason: 'max_output_tokens_recovery', attempt: count + 1 },
}
continue  // 下一轮迭代

这种写法的好处:transition.reason 使测试代码可以精确断言"哪条恢复路径被触发了",而无需检查消息内容。

7.3.3 yieldMissingToolResultBlocks():孤立 tool_use 的清理

当流式响应中断(用户 Ctrl+C、网络故障)时,已发出的 tool_use 块可能没有对应的 tool_result

function* yieldMissingToolResultBlocks(
  assistantMessages: AssistantMessage[],
  errorMessage: string,
) {
  for (const assistantMessage of assistantMessages) {
    const toolUseBlocks = /* 从 assistantMessage 提取 tool_use 块 */
    for (const toolUse of toolUseBlocks) {
      yield createUserMessage({
        content: [{
          type: 'tool_result',
          content: errorMessage,     // 错误描述作为结果内容
          is_error: true,
          tool_use_id: toolUse.id,   // 配对原来的 tool_use
        }],
      })
    }
  }
}

这确保了消息历史始终保持 API 所要求的完整性:每个 tool_use 一定有配对的 tool_result

7.3.4 countToolCalls()maxTurns 保护

// 在每轮 tool 执行后(更新 state 之前)
const nextTurnCount = turnCount + 1
if (maxTurns && nextTurnCount > maxTurns) {
  yield createAttachmentMessage({
    type: 'max_turns_reached',
    maxTurns,
    turnCount: nextTurnCount,
  })
  return { reason: 'max_turns' }
}

maxTurns 是防止 Agent 失控的最后防线。Claude Code 默认值为 undefined(无限制),但 SDK 调用者可以设置,CI 场景通常设为 10~20


7.4 最小化产出物

代码骨架位于 ../chapters/07/src/,参考实现位于 ../chapters/07/solution/前置条件:需要 ANTHROPIC_API_KEY 环境变量(npm start 需要)。

本章要实现什么

../chapters/07/src/loop.ts 中完成 Agentic Loop。

接口规范(已提供,不要修改):

export type LoopResult = {
  reason: 'completed' | 'max_turns' | 'error' | 'aborted'
  finalMessages: Anthropic.MessageParam[]
  turnCount: number
  totalInputTokens: number
  totalOutputTokens: number
}

export function toAnthropicTools(): Anthropic.Tool[]
export async function executeAllTools(
  toolUseBlocks: Anthropic.ToolUseBlock[],
  onToolStart?: (name: string, input: unknown) => void,
  onToolEnd?: (name: string, result: { content: string; isError?: boolean }) => void,
): Promise<Anthropic.ToolResultBlockParam[]>
export async function agenticLoop(
  initialMessages: Anthropic.MessageParam[],
  systemPrompt?: string,
  maxTurns?: number,
  opts?: { onText?; onToolStart?; onToolEnd?; signal? }
): Promise<LoopResult>

你需要实现

  1. toAnthropicTools():调用第 4 章的 getToolSchemas(),转为 Anthropic.Tool[] 格式
  2. executeAllTools()Promise.all 并行执行所有 tool_use 块,每个调用第 4 章的 TOOL_REGISTRY
  3. agenticLoop()
    • 用第 6 章的 ConversationManager 管理消息历史
    • while(true) 循环:调用第 5 章的 streamQuery(带工具列表)
    • 无 tool_use → return { reason: 'completed' }
    • 有 tool_use → executeAllTools → 追加 tool_result → 继续
    • 超过 maxTurns → return { reason: 'max_turns' }

关键约束

  • 必须 import 并复用第 4 章的 TOOL_REGISTRY、第 5 章的 streamQuery、第 6 章的 ConversationManager
  • 工具执行失败不退出循环,以 is_error: truetool_result 反馈给 LLM

验收

cd docs/chapters/07
npm install
npm test

卡住时查看 ../chapters/07/solution/loop.ts


7.5 本章小结

本章完成了 Coding Agent 最核心的一跳——让 LLM 自主驱动工具调用循环

  1. needsFollowUp 是真正的 Loop 驱动信号,而不是 stop_reasonstop_reason === "tool_use" 并不可靠,Claude Code 在流式接收时检测 tool_use 块并设置 needsFollowUp = true

  2. 并行工具执行:同一轮 LLM 响应中的多个 tool_use 块被并行执行(Promise.all),减少等待时间。每个 tool_use 必须有配对的 tool_result,失败也不例外。

  3. Streaming Tool Execution:Claude Code 的 StreamingToolExecutor 在 LLM 流式输出时就开始执行工具,进一步降低端到端延迟。

  4. 多层错误恢复max_output_tokens → 自动注入恢复提示重试;prompt_too_long → 依次尝试 context collapse、reactive compact;工具失败 → 作为 is_error: truetool_result 反馈给 LLM 自行处理。

  5. maxTurns 是安全网:防止 LLM 陷入工具调用无限循环。生产环境中这个值至关重要,设太小会中断合理任务,设太大会带来 API 成本失控风险。

  6. State 对象模式:Claude Code 通过不可变 State + state = {...}; continue 的惯用写法管理迭代状态,使每条恢复路径的触发原因(transition.reason)都可精确断言。

  7. 真正的组合loop.ts 通过 import 复用第 4 章的 TOOL_REGISTRY、第 5 章的 streamQuery、第 6 章的 ConversationManager,本章只新增了 Agentic Loop 的调度逻辑,没有重复实现任何已有功能。

至此,前七章共同构成了一个完整的 Coding Agent 骨架:进程 → 终端 UI → CLI 路由 → 工具 → LLM → 消息历史 → Agentic Loop。下一章将在这个骨架上加装"安全阀"——权限与安全仲裁系统。


下一章

第 8 章:权限与安全仲裁