☰
12-Factor Agents 实战:用 Express 为 Agent 添加 HTTP API 端点
2026/10/1 2:39:57 网站建设 项目流程
  • 文档
  • 教程
  • 人工智能
  • 大模型
  • AI Agent

【免费下载链接】12-factor-agents

What are the principles we can use to build LLM-powered software that is actually good enough to put in the hands of production customers?

项目地址:https://gitcode.com/GitHub_Trending/12/12-factor-agents
点击查看免费下载

导读

本文基于 12-factor-agents 开源仓库中的 Workshop 章节 08-api-endpoints,讲解如何为前一章基于 BAML 与 TypeScript 构建的"工具循环 Agent"添加一个 Express HTTP 服务,将其从命令行交互升级为可通过POST /thread调用的生产可用 API。读完本文,你将掌握:如何用express.json()解析请求体、如何把用户消息构造成 Agent 的Thread事件并通过agentLoop驱动循环、如何用curl验证返回的 agentic trace,以及如何通过PORT环境变量控制服务端口——这正是 12-Factor 方法论中"进程可配置、服务可通过网络暴露"的一次最小实践。

前置背景:我们正在把什么暴露成 API

在进入本章代码之前,先明确当前 Agent 的形态。上一章(07-context-window)之后,核心的 Agent 循环位于 src/agent.ts:

  • Thread类持有events: Event[],每个事件形如{ type, data },例如user_input、tool_call、tool_response、human_response;
  • serializeForLLM()将事件序列化为 XML 风格文本,作为 BAML 函数DetermineNextStep的输入上下文;
  • agentLoop(thread)是一个while (true)循环:调用 BAML 生成的b.DetermineNextStep(thread.serializeForLLM())让模型决定下一步,将决策以tool_call推入事件流,若是add/subtract/multiply/divide则通过handleNextStep执行本地计算并把结果以tool_response写回线程,直到模型输出done_for_now或request_more_information才返回。

在此之前,该 Agent 的入口是命令行 src/cli.ts:从process.argv.slice(2)取参数、new Thread([{ type: "user_input", data: message }])、调用agentLoop,并用readline在request_more_information时向人类追问。本章的任务就是:把这段"命令行内循环"迁移到 HTTP 服务里,让任何客户端都能通过网络发起一次对话。

第一步:关闭 BAML 日志并安装依赖

本章为了聚焦 HTTP 层,会先关闭 BAML 的日志输出(若想观察模型决策细节,可自行开启):

export BAML_LOG=off

安装 Express 及其类型定义,同时安装 supertest(后续章节做 HTTP 测试时会用到):

npm install express && npm install --save-dev @types/express supertest

从 package.json 可以看到,项目运行 TypeScript 的方式是tsx("dev": "tsx src/index.ts"),并依赖baml(BAML 客户端运行时)。Express 与@types/express会作为新依赖加入,tsx可以直接执行带 ES 模块语法的server.ts而无需先编译。

第二步:编写服务器实现

将现成的服务器实现复制到源码目录:

cp ./walkthrough/08-server.ts src/server.ts

完整的实现如下(与 walkthrough/08-server.ts 一致):

// ./walkthrough/08-server.ts import express from 'express'; import { Thread, agentLoop } from '../src/agent'; const app = express(); app.use(express.json()); app.set('json spaces', 2); // POST /thread - Start new thread app.post('/thread', async (req, res) => { const thread = new Thread([{ type: "user_input", data: req.body.message }]); const result = await agentLoop(thread); res.json(result); }); // GET /thread/:id - Get thread status app.get('/thread/:id', (req, res) => { // optional - add state res.status(404).json({ error: "Not implemented yet" }); }); const port = process.env.PORT || 3000; app.listen(port, () => { console.log(`Server running on port ${port}`); }); export { app };

逐段拆解这段代码:

  • app.use(express.json()):把请求体按 JSON 解析并挂到req.body,这是POST /thread能读取message字段的前提;
  • app.set('json spaces', 2):让res.json(...)输出的响应带缩进,便于人眼阅读 agentic trace;
  • POST /thread:接收{"message": "..."},构造一个仅含单个user_input事件的Thread,交给agentLoop跑完整个工具循环,再把完整的线程(含每一步tool_call/tool_response)作为 JSON 返回。对比 CLI 版本的 src/cli.ts,这里把"输入来自键盘、输出到终端"替换为"输入来自请求体、输出到响应体",而 Agent 核心循环本身完全复用;
  • GET /thread/:id:本章留作占位(404 Not implemented),下一章 09-state-management 会利用ThreadStore让该端点真正返回已存线程。这也体现了 12-Factor 的渐进式设计——先打通最小链路,再补状态;
  • process.env.PORT || 3000:端口来自环境变量,符合 12-Factor 中"配置存于环境"的原则,本地无配置时默认 3000;
  • export { app }:把 Express 实例导出,为后续引入 supertest 等测试框架或把 app 与 listen 分离做准备。

第三步:启动服务

用tsx直接运行(无需先tsc编译):

npx tsx src/server.ts

启动成功后控制台输出:

Server running on port 3000

关于运行环境需要注意:Agent 的决策依赖 BAML 生成的客户端 baml_client,该客户端默认以async模式生成(见generators.baml的default_client_mode async),并调用 clients.baml 中定义的CustomGPT4o(provider openai,模型gpt-4o,密钥来自env.OPENAI_API_KEY)。因此运行前需要确保OPENAI_API_KEY已设置、BAML 代码已按 baml_src 生成到客户端目录。

第四步:用 curl 验证完整 Agent 调用链

在另一个终端发起一次"计算任务":

curl -X POST http://localhost:3000/thread \ -H "Content-Type: application/json" \ -d '{"message":"can you add 3 and 4"}'

请求会走如下链路:

  1. Express 解析 JSON 得到message;
  2. new Thread([{ type: "user_input", data: "can you add 3 and 4" }])构造初始事件;
  3. agentLoop调用b.DetermineNextStep,模型依据 agent.baml 中的提示词判断出intent: "add"、参数a: 3, b: 4;
  4. src/agent.ts 的handleNextStep命中case "add"执行a + b,把tool_response推入事件;
  5. 循环再次调用模型,得到done_for_now,agentLoop返回线程。

响应中会包含完整的 agentic trace(即事件序列),并以done_for_now收尾:

{"intent":"done_for_now","message":"The sum of 3 and 4 is 7."}

注意:该响应是"最后一个事件的 data",而实际res.json(result)返回的是整个Thread对象(含events数组),其中每一步的tool_call(模型决策,如add及其参数)与tool_response(本地计算结果7)都会逐条列出,这正是可观测的 agentic trace。

局限与下一章的衔接

本章的 API 是无状态的:每次POST /thread都新建一个空线程,多轮对话、澄清追问(request_more_information)都无法跨请求继续——GET /thread/:id的404占位也暗示了这一点。若要支撑生产场景的"多轮对话 + 异步澄清",需要为线程引入状态存储。仓库中紧随其后的 09-state-management 正好演示了如何用ThreadStore(内存 Map +crypto.randomUUID()生成线程 ID)实现POST /thread返回thread_id与response_url、实现GET /thread/:id读取线程、实现POST /thread/:id/response追加human_response并继续agentLoop。读完本章后,可以顺着这一节把无状态端点升级为可续跑的对话 API。

小结

本章用约 30 行代码完成了 Agent 从 CLI 到 HTTP API 的跃迁,核心要点可归纳为:

  • 复用而非重写:Thread与agentLoop完全复用,HTTP 层只负责"把请求体变成user_input事件、把返回的线程变成 JSON 响应";
  • 最小可用端点:POST /thread是唯一真实实现,GET /thread/:id预留接口,体现了渐进式开发节奏;
  • 配置外置:端口通过PORT环境变量控制,符合 12-Factor 配置管理原则;
  • 可观测性:响应携带完整 agentic trace,方便调试模型决策与工具执行过程。

至此,你已经掌握了用 Express 暴露 Agent 能力的最小路径。下一章引入状态存储后,这套端点将具备真正的多轮对话能力。

  • 文档
  • 教程
  • 人工智能
  • 大模型
  • AI Agent

【免费下载链接】12-factor-agents

What are the principles we can use to build LLM-powered software that is actually good enough to put in the hands of production customers?

项目地址:https://gitcode.com/GitHub_Trending/12/12-factor-agents
点击查看免费下载
上一篇:Daft LeRobot 视频解码公开数据集验证:逐行解码到按分片批量解码的 4-13 倍提速
下一篇:Happy Island Designer 岛屿设计工具从零到高手实战手册

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询