- 文档
- 教程
- 人工智能
- 大模型
- AI Agent
【免费下载链接】12-factor-agents
What are the principles we can use to build LLM-powered software that is actually good enough to put in the hands of production customers?
导读
本文基于 12-factor-agents 开源仓库中的 Workshop 章节 08-api-endpoints,讲解如何为前一章基于 BAML 与 TypeScript 构建的"工具循环 Agent"添加一个 Express HTTP 服务,将其从命令行交互升级为可通过POST /thread调用的生产可用 API。读完本文,你将掌握:如何用express.json()解析请求体、如何把用户消息构造成 Agent 的Thread事件并通过agentLoop驱动循环、如何用curl验证返回的 agentic trace,以及如何通过PORT环境变量控制服务端口——这正是 12-Factor 方法论中"进程可配置、服务可通过网络暴露"的一次最小实践。
前置背景:我们正在把什么暴露成 API
在进入本章代码之前,先明确当前 Agent 的形态。上一章(07-context-window)之后,核心的 Agent 循环位于 src/agent.ts:
Thread类持有events: Event[],每个事件形如{ type, data },例如user_input、tool_call、tool_response、human_response;serializeForLLM()将事件序列化为 XML 风格文本,作为 BAML 函数DetermineNextStep的输入上下文;agentLoop(thread)是一个while (true)循环:调用 BAML 生成的b.DetermineNextStep(thread.serializeForLLM())让模型决定下一步,将决策以tool_call推入事件流,若是add/subtract/multiply/divide则通过handleNextStep执行本地计算并把结果以tool_response写回线程,直到模型输出done_for_now或request_more_information才返回。
在此之前,该 Agent 的入口是命令行 src/cli.ts:从process.argv.slice(2)取参数、new Thread([{ type: "user_input", data: message }])、调用agentLoop,并用readline在request_more_information时向人类追问。本章的任务就是:把这段"命令行内循环"迁移到 HTTP 服务里,让任何客户端都能通过网络发起一次对话。
第一步:关闭 BAML 日志并安装依赖
本章为了聚焦 HTTP 层,会先关闭 BAML 的日志输出(若想观察模型决策细节,可自行开启):
export BAML_LOG=off安装 Express 及其类型定义,同时安装 supertest(后续章节做 HTTP 测试时会用到):
npm install express && npm install --save-dev @types/express supertest从 package.json 可以看到,项目运行 TypeScript 的方式是tsx("dev": "tsx src/index.ts"),并依赖baml(BAML 客户端运行时)。Express 与@types/express会作为新依赖加入,tsx可以直接执行带 ES 模块语法的server.ts而无需先编译。
第二步:编写服务器实现
将现成的服务器实现复制到源码目录:
cp ./walkthrough/08-server.ts src/server.ts完整的实现如下(与 walkthrough/08-server.ts 一致):
// ./walkthrough/08-server.ts import express from 'express'; import { Thread, agentLoop } from '../src/agent'; const app = express(); app.use(express.json()); app.set('json spaces', 2); // POST /thread - Start new thread app.post('/thread', async (req, res) => { const thread = new Thread([{ type: "user_input", data: req.body.message }]); const result = await agentLoop(thread); res.json(result); }); // GET /thread/:id - Get thread status app.get('/thread/:id', (req, res) => { // optional - add state res.status(404).json({ error: "Not implemented yet" }); }); const port = process.env.PORT || 3000; app.listen(port, () => { console.log(`Server running on port ${port}`); }); export { app };逐段拆解这段代码:
app.use(express.json()):把请求体按 JSON 解析并挂到req.body,这是POST /thread能读取message字段的前提;app.set('json spaces', 2):让res.json(...)输出的响应带缩进,便于人眼阅读 agentic trace;POST /thread:接收{"message": "..."},构造一个仅含单个user_input事件的Thread,交给agentLoop跑完整个工具循环,再把完整的线程(含每一步tool_call/tool_response)作为 JSON 返回。对比 CLI 版本的 src/cli.ts,这里把"输入来自键盘、输出到终端"替换为"输入来自请求体、输出到响应体",而 Agent 核心循环本身完全复用;GET /thread/:id:本章留作占位(404 Not implemented),下一章 09-state-management 会利用ThreadStore让该端点真正返回已存线程。这也体现了 12-Factor 的渐进式设计——先打通最小链路,再补状态;process.env.PORT || 3000:端口来自环境变量,符合 12-Factor 中"配置存于环境"的原则,本地无配置时默认 3000;export { app }:把 Express 实例导出,为后续引入 supertest 等测试框架或把 app 与 listen 分离做准备。
第三步:启动服务
用tsx直接运行(无需先tsc编译):
npx tsx src/server.ts启动成功后控制台输出:
Server running on port 3000关于运行环境需要注意:Agent 的决策依赖 BAML 生成的客户端 baml_client,该客户端默认以async模式生成(见generators.baml的default_client_mode async),并调用 clients.baml 中定义的CustomGPT4o(provider openai,模型gpt-4o,密钥来自env.OPENAI_API_KEY)。因此运行前需要确保OPENAI_API_KEY已设置、BAML 代码已按 baml_src 生成到客户端目录。
第四步:用 curl 验证完整 Agent 调用链
在另一个终端发起一次"计算任务":
curl -X POST http://localhost:3000/thread \ -H "Content-Type: application/json" \ -d '{"message":"can you add 3 and 4"}'请求会走如下链路:
- Express 解析 JSON 得到
message; new Thread([{ type: "user_input", data: "can you add 3 and 4" }])构造初始事件;agentLoop调用b.DetermineNextStep,模型依据 agent.baml 中的提示词判断出intent: "add"、参数a: 3, b: 4;- src/agent.ts 的
handleNextStep命中case "add"执行a + b,把tool_response推入事件; - 循环再次调用模型,得到
done_for_now,agentLoop返回线程。
响应中会包含完整的 agentic trace(即事件序列),并以done_for_now收尾:
{"intent":"done_for_now","message":"The sum of 3 and 4 is 7."}注意:该响应是"最后一个事件的 data",而实际res.json(result)返回的是整个Thread对象(含events数组),其中每一步的tool_call(模型决策,如add及其参数)与tool_response(本地计算结果7)都会逐条列出,这正是可观测的 agentic trace。
局限与下一章的衔接
本章的 API 是无状态的:每次POST /thread都新建一个空线程,多轮对话、澄清追问(request_more_information)都无法跨请求继续——GET /thread/:id的404占位也暗示了这一点。若要支撑生产场景的"多轮对话 + 异步澄清",需要为线程引入状态存储。仓库中紧随其后的 09-state-management 正好演示了如何用ThreadStore(内存 Map +crypto.randomUUID()生成线程 ID)实现POST /thread返回thread_id与response_url、实现GET /thread/:id读取线程、实现POST /thread/:id/response追加human_response并继续agentLoop。读完本章后,可以顺着这一节把无状态端点升级为可续跑的对话 API。
小结
本章用约 30 行代码完成了 Agent 从 CLI 到 HTTP API 的跃迁,核心要点可归纳为:
- 复用而非重写:
Thread与agentLoop完全复用,HTTP 层只负责"把请求体变成user_input事件、把返回的线程变成 JSON 响应"; - 最小可用端点:
POST /thread是唯一真实实现,GET /thread/:id预留接口,体现了渐进式开发节奏; - 配置外置:端口通过
PORT环境变量控制,符合 12-Factor 配置管理原则; - 可观测性:响应携带完整 agentic trace,方便调试模型决策与工具执行过程。
至此,你已经掌握了用 Express 暴露 Agent 能力的最小路径。下一章引入状态存储后,这套端点将具备真正的多轮对话能力。
- 文档
- 教程
- 人工智能
- 大模型
- AI Agent
【免费下载链接】12-factor-agents
What are the principles we can use to build LLM-powered software that is actually good enough to put in the hands of production customers?
相关推荐
为 12-Factor Agent 添加计算器工具:用 BAML 结构化输出驱动 Agent 的下一步决策
为 12 Factor Agent 添加计算器工具:用 BAML 结构化输出驱动 Agent 的下一步决策 本文以开源仓库 12 factor agents 工
文档教程人工智能大模型AI Agent12-factor-agents 实战第 1 章:用 BAML 搭建你的第一个 CLI Agent 循环
12 factor agents 实战第 1 章:用 BAML 搭建你的第一个 CLI Agent 循环 导读 本文对应仓库 12 factor agents
文档教程人工智能大模型AI Agentx402-express 实战:用 Express 中间件为 API 端点搭建加密货币付费墙
x402 express 实战:用 Express 中间件为 API 端点搭建加密货币付费墙 本文以 x402 项目仓库中的 Express 示例服务器( e2
桌面应用文档
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考