# agent-loop **Repository Path**: laityy/agent-loop ## Basic Information - **Project Name**: agent-loop - **Description**: No description available - **Primary Language**: Unknown - **License**: Not specified - **Default Branch**: master - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 1 - **Created**: 2026-06-15 - **Last Updated**: 2026-09-13 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # Agent Loop from Scratch 手写一个 agent 循环,不依赖任何 agent 框架(LangChain、Vercel AI SDK 的 agent 封装等),只用 Anthropic Python SDK 的底层 `messages.create(stream=True)` API 实现完整的 **think → act → observe** 循环。 ## 为什么 Agent 框架封装了核心循环,用起来方便但遮蔽了关键细节: - 工具调用怎么从流式响应里解析出来的? - tool call 的 JSON 参数是分块到达的,怎么拼接? - context 超过 token 限制时怎么压缩? - API 报错了怎么重试? - 工具执行报错了怎么反馈给模型? 这个项目把每一层都展开,500 行代码把这些问题说清楚。 ## 快速开始 ```bash # 安装依赖 pip install -r requirements.txt # 运行测试 python -m pytest tests/ -v # 运行示例(需要 Anthropic API Key) ANTHROPIC_API_KEY=sk-ant-... python examples/calculator.py ``` ## 架构 ``` 四个核心模块:(agent + 推理行动记忆) - Agent (agent.py) — 编排层:think→act→observe 主循环,串联所有模块 - LLMClient (llm.py) — 推理引擎:流式调用、事件解析、指数退避重试 - ToolRegistry (tools.py) — 行动系统:工具注册、Schema 生成、异步执行、异常转换 - ContextManager (context.py) — 记忆系统:消息历史、Token 计数、超限自动压缩 ┌─────────────────────────────────────────────────────┐ │ Agent (agent.py) │ │ think → act → observe 循环,连接所有模块 │ └──────────┬──────────────────┬───────────────────────┘ │ │ ┌──────▼──────┐ ┌──────▼──────┐ │ LLMClient │ │ ContextMgr │ │ (llm.py) │ │ (context.py)│ │ │ │ │ │ 流式调用 │ │ 消息历史 │ │ 事件解析 │ │ Token 计数 │ │ 错误重试 │ │ 自动压缩 │ └──────┬──────┘ └─────────────┘ │ ┌──────▼──────┐ │ ToolRegistry│ │ (tools.py) │ │ │ │ 工具注册 │ │ Schema 生成 │ │ 异步执行 │ └─────────────┘ ``` ## 模块说明 | 文件 | 职责 | 行数 | |------|------|------| | `src/types.py` | 15 种 dataclass:Message, Tool, ToolCall, ToolResult, LLMEvent(8 种流事件) | ~170 | | `src/llm.py` | 推理引擎:流式调用 Anthropic SDK,解析原始事件,指数退避重试 | ~110 | | `src/tools.py` | 行动系统:工具注册、Schema 生成、异步执行、异常转换 | ~50 | | `src/context.py` | 记忆系统:消息历史管理,Token 计数,超限自动压缩 | ~90 | | `src/agent.py` | 编排层:think(流式收集+拼接 tool call)→ act(执行工具)→ observe(写回结果) | ~130 | ## 核心循环 ```python # agent.py — 完整循环逻辑 async def run(self, user_message: str) -> AgentRunResult: self._ctx.add_user_message(user_message.strip()) for iteration in range(1, self.max_iterations + 1): # 上下文超限 → 压缩 if self._ctx.should_compact(): self._ctx.compact() # THINK: 调用 LLM,收集文本和 tool call full_text, raw_tool_calls = await self._think() # 没有 tool call → 对话结束 if not raw_tool_calls: return AgentRunResult( final_text=full_text, state=AgentState.FINISHED, ... ) # ACT: 解析 tool call,逐个执行 parsed_calls = self._parse_tool_calls(raw_tool_calls) self._ctx.add_assistant_with_tool_calls(parsed_calls) results = [] for tc in parsed_calls: result = await self._tools.execute(tc.name, tc.input) result.tool_call_id = tc.id results.append(result) # OBSERVE: 工具结果写回上下文,继续循环 self._ctx.add_tool_results(results) return AgentRunResult(state=AgentState.ERROR, ...) ``` ## 深入理解 `_think()` — agent 循环的核心 `_think()` 是 agent 循环中最重要的方法。理解它,就理解了流式 LLM 交互的本质。 ### 背景:流式 tool call 是怎么传输的 当 LLM 决定调用工具时,它不会一次性发回完整的 JSON。Anthropic 的流式 API 会把参数拆成多个分块事件: ``` content_block_start → "我要开始一个叫 'calculator' 的工具调用" content_block_delta → '{"a": ' content_block_delta → '3, "b"' content_block_delta → ': 5, "o' content_block_delta → 'peration": "add"}' content_block_stop → "这个工具调用结束了" ``` 每个 delta 可能只有几字节到几十字节。JSON 参数被碎片化传输,必须由接收方自行拼接。这就是框架帮你做而你在这份代码里看到全貌的东西。 ### 状态机 用单一可变变量 `current_tool` 追踪当前正在装配的 tool call,流事件驱动状态转换: ``` ToolCallStart IDLE ──────────────────────────────→ ASSEMBLING (current_tool=None) (current_tool={id, name, json_str}) ToolCallDelta append partial_json ASSEMBLING ───────────────────────→ ASSEMBLING ContentBlockStop save completed tool ASSEMBLING ───────────────────────→ IDLE reset to None ``` **为什么是状态机而不是"先收集再解析"?** 因为 tool call 参数可能很大(base64 图片、大段代码)。全部缓存到内存再处理会在大数据量时崩溃。流式处理保证恒定的内存开销。 ### 边缘情况 | 场景 | 处理方式 | |------|---------| | tool call id 在 delta 中才到达 | 检查 `_tid`,回填 `current_tool["id"]` | | 流在没有 ContentBlockStop 时结束 | 方法末尾的 cleanup 代码保存最后一个 tool | | 没有注册任何工具 | `schemas=None`,LLM 不会发起 tool call,流只含文本 | | LLM 中途返回错误 | 错误被追加到文本中,agent 可决定重试或告知用户 | | 并行 tool call | ToolCallStart 到达时先保存前一个,再开始新装配 | ### 完整代码(含注释) ```python async def _think(self) -> tuple[str, list[dict]]: full_text_parts: list[str] = [] tool_calls_raw: list[dict] = [] # 当前正在装配的 tool call。None 表示不在 tool call 中。 current_tool: Optional[dict] = None # 传入工具 schema。None(非空列表)表示"不要用工具"。 schemas = self._tools.to_anthropic_schemas() if self._tools.tool_names() else None async for event in self._llm.stream( messages=self._ctx.get_messages(), tools=schemas, ): match event: # --- 文本内容 --- case TextDelta(text=text): # 助手回复的纯文本片段。如果没有后续 tool call,这就是最终回答。 full_text_parts.append(text) # --- Tool call 生命周期 --- case ToolCallStart(tool_call_id=tid, tool_name=tname): # 新的 tool call 开始了。如果之前有正在装配的,先保存它。 if current_tool and current_tool.get("id"): tool_calls_raw.append(current_tool) # 新建装配对象:id、工具名、空 JSON 缓冲区。 current_tool = {"id": tid, "name": tname, "input_json_str": ""} case ToolCallDelta(partial_json=pj): # tool call 参数的 JSON 片段。追加到装配中的 JSON 缓冲区。 if current_tool: current_tool["input_json_str"] += pj case ContentBlockStop(): # 一个内容块(文本或 tool_call)完全结束。 # 对 tool call 来说,这是可靠的"完成"信号。 if current_tool and current_tool.get("id"): tool_calls_raw.append(current_tool) current_tool = None # --- 错误处理 --- case ErrorEvent(error=err): # 流级别的错误(限流、服务端错误)。作为文本捕获。 full_text_parts.append(f"\n[LLM Error: {err}]") # 兜底:如果流在没有 ContentBlockStop 的情况下结束了,不丢最后一个 tool if current_tool and current_tool.get("id"): tool_calls_raw.append(current_tool) return "".join(full_text_parts), tool_calls_raw ``` ## 关键设计决策 **为什么 tool call 的 JSON 要手动拼接?** Anthropic 的流式 API 把 tool call 参数以 `input_json_delta` 事件分块发送(每块几字节到几十字节)。框架帮你拼好了,这里你看到完整的拼接过程。 **为什么用 `match-case`?** Python 3.10+ 的结构模式匹配是处理流事件的天然工具——每个事件类型一个分支,漏了哪种类型编译器不会提醒但代码审查一望便知。 **重试策略?** 只有连接超时、限流这类瞬态错误才重试,2^attempt 秒退避。HTTP 4xx(鉴权失败、参数错误)直接报错不重试。 **压缩策略?** 最简单的方案:保留 system prompt + 最后 N 条消息,中间的历史替换为 `[Context from earlier conversation: ...]` 摘要。真正生产环境会用一个单独的 LLM 调用来做摘要,但原理相同。 **工具执行的 tool_call_id 为什么是空?** 当前设计是顺序执行(一次一个 tool call)。并行 tool call 需要 LLM 客户端在解析事件时维护 index→id 映射,这是一个有意的简化。 ## 移植到 TypeScript Python 版是参考实现,结构可以直接映射到 TypeScript: | Python | TypeScript | |--------|------------| | `@dataclass` | `interface` / `type` | | `match-case` | `switch` / discriminated union | | `async for` | `for await...of` | | `AsyncIterator[T]` | `AsyncIterable` | | `tiktoken` | `js-tiktoken` / `gpt-tokenizer` | | `anthropic` SDK | `@anthropic-ai/sdk` | 核心逻辑不变——流式事件解析、tool call JSON 拼接、指数退避重试逻辑完全一样。 ## 许可证 MIT