1. 从零构建复杂 AI Agent,为什么 Golang 是那个被低估的选项
如果你正在搜索“Golang 开发 AI Agent”“MCP 配置实战”“A2A 多 Agent 协作”,大概率已经翻过不少 Python 教程,却发现落到工程化时总差一口气:并发上不去、部署链路太长、类型安全靠自觉。这篇内容就是围绕这个场景展开的——用 Golang 从零搭一个能跑起来的复杂 AI Agent,核心覆盖两件事:MCP 工具接入和 A2A 多 Agent 协作。
先说清楚它适合谁。如果你已经写过简单的 LLM 调用,想让 Agent 真正调用外部工具、并且多个 Agent 之间能互相派活,那这篇的配置骨架可以直接抄。如果你只是想让模型聊聊天,那用不上这么重的架构。复杂 AI Agent 的“复杂”不在于模型多强,而在于三件事:工具能标准化接入、任务能拆解分发、执行过程能观测和干预。Golang 在这三件事上有天然优势——goroutine 处理流式响应、强类型约束工具参数、编译期就能发现接口不匹配。
我试过用脚本语言快速拼一个 Agent,前期很爽,等到要接三四个工具、再加一个协作 Agent 时,参数对不上、状态丢失、并发写冲突全冒出来了。所以这篇不走“先跑通再说”的路子,而是先把 config.toml 和 settings.json 两个骨架立住,再往里填 MCP 和 A2A 的逻辑。下面所有配置和代码都是可复制的,你跟着改路径和 Key 就能跑。
2. TaoToken 前置:统一 Key 与 API 通道怎么接
在写任何 Agent 代码之前,先把模型调用通道固定下来。复杂 Agent 会频繁请求模型,如果每个 Agent 各自维护一套 Key 和 BaseURL,后面排查问题会非常痛苦。这里用 TaoToken 做统一入口,官网地址是 https://taotoken.net/?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= ,API 端点是 https://taotoken.net/api 。
操作路径很直接:先到控制台创建 API Key,地址是 https://taotoken.net/console?utm_source=taotoken_aicg_blog_end&utm_content=console&utm_campaign=rewrite ,生成后复制保存。然后在项目根目录建一个.env或者直接写进 config.toml。我建议用环境变量注入,避免 Key 进版本库。
export TAOTOKEN_API_KEY="sk-你的key" export TAOTOKEN_BASE_URL="https://taotoken.net/api"如果你要验证模型是否通,可以直接用模型对话页面发一条测试消息: https://taotoken.net/models?utm_source=taotoken_aicg_blog_end&utm_content=models&utm_campaign=rewrite 。这一步别跳过,因为后面 Agent 启动失败时,你要能快速判断是模型通道问题还是代码问题。Key 的管理页面在 https://taotoken.net/api-keys?utm_source=taotoken_aicg_blog_end&utm_content=api-keys&utm_campaign=rewrite ,接入文档在 https://taotoken.net/doc?utm_source=taotoken_aicg_blog_end&utm_content=doc&utm_campaign=rewrite ,遇到 401 或 404 先翻文档比瞎改代码快。
注意:BaseURL 末尾不要多加
/v1,具体以接入文档为准。不同 SDK 对路径拼接的处理不一样,写死之前先用 curl 验证一次。
3. 可复制配置:config.toml 与 settings.json 骨架
复杂 Agent 的配置分两层:一层是运行时配置(模型、超时、并发数),放 config.toml;一层是 Agent 与工具的描述(AgentCard、MCP Server 列表),放 settings.json。这样拆的好处是,改模型参数不用动 Agent 定义,加工具不用重编译。
先看 config.toml:
[model] provider = "taotoken" base_url = "https://taotoken.net/api" api_key_env = "TAOTOKEN_API_KEY" default_model = "claude-sonnet" timeout_seconds = 60 max_retries = 3 [agent] name = "supervisor" max_concurrent_tasks = 8 checkpoint_store = "redis" checkpoint_addr = "127.0.0.1:6379" [mcp.amap] transport = "stdio" command = "npx" args = ["-y", "@amap/mcp-server"] enabled = true [mcp.tavily] transport = "sse" url = "http://127.0.0.1:8081/sse" enabled = true [a2a] listen_addr = "0.0.0.0:10000" agent_card_path = "./settings.json"再看 settings.json,这里定义 AgentCard 和下游 Agent 地址:
{ "name": "deepresearch", "description": "深度搜索 Agent,支持多轮检索与总结", "version": "1.0.0", "capabilities": { "streaming": true, "pushNotifications": false }, "skills": [ { "id": "deep_search", "name": "深度搜索", "description": "对复杂问题多角度检索并生成报告" } ], "downstream_agents": [ { "name": "urlreader", "server_url": "http://127.0.0.1:10001" }, { "name": "lbshelper", "server_url": "http://127.0.0.1:10002" } ] }这两个文件的关系是:config.toml 告诉程序“用什么模型、连哪些 MCP Server”,settings.json 告诉其他 Agent“我是谁、我能干什么、我找谁协作”。A2A 协议里 AgentCard 就是靠 settings.json 生成的,别的 Agent 通过它来发现你的能力。
提示:MCP 的 stdio 传输方式适合本地命令行工具,sse 适合已经跑成 HTTP 服务的工具。如果你不确定用哪种,先看工具官方文档给的启动方式。
4. MCP 工具接入:从配置到可调用
MCP 的核心价值是把工具调用标准化。在 Golang 里,你需要做三件事:读取 config.toml 里的 MCP 配置、建立连接、把工具列表注册给 ChatModel。
先写一个连接 MCP Server 的函数:
package mcp import ( "context" "fmt" "os/exec" "github.com/eino-contrib/mcp" ) func ConnectMCP(ctx context.Context, cfg Config) ([]mcp.Tool, []*schema.ToolInfo, error) { var client *mcp.Client var err error switch cfg.Transport { case "stdio": cmd := exec.Command(cfg.Command, cfg.Args...) client, err = mcp.NewStdioClient(ctx, cmd) case "sse": client, err = mcp.NewSSEClient(ctx, cfg.URL) default: return nil, nil, fmt.Errorf("unsupported transport: %s", cfg.Transport) } if err != nil { return nil, nil, fmt.Errorf("connect mcp failed: %w", err) } tools, err := client.ListTools(ctx) if err != nil { return nil, nil, fmt.Errorf("list tools failed: %w", err) } var toolInfos []*schema.ToolInfo for _, t := range tools { toolInfos = append(toolInfos, &schema.ToolInfo{ Name: t.Name, Desc: t.Description, ParamsOneOf: schema.NewParamsOneOfByParams(t.InputSchema), }) } return tools, toolInfos, nil }然后在构建运行图时,把工具绑定到 ChatModel:
chatModel, err := openai.NewChatModel(ctx, &openai.ChatModelConfig{ APIKey: os.Getenv("TAOTOKEN_API_KEY"), BaseURL: "https://taotoken.net/api", Model: "claude-sonnet", }) if err != nil { log.Fatal(err) } tools, toolInfos, err := mcp.ConnectMCP(ctx, cfg.MCP.Amap) if err != nil { log.Fatal(err) } if err = chatModel.BindTools(toolInfos); err != nil { log.Fatal(err) }这里的关键点是BindTools。绑定之后,模型在推理时如果判断需要调用工具,会返回 ToolCalls,你的 ToolNode 再根据 ToolCalls 去执行实际调用。MCP 帮你把“工具描述”和“工具执行”解耦了——模型只需要知道工具的名字和参数格式,不需要知道底层是 HTTP 还是命令行。
如果你要接多个 MCP Server,把它们的 toolInfos 合并成一个切片再 BindTools 就行。但要注意工具名冲突,两个 Server 都有search工具时,后绑定的会覆盖前面的。解决办法是在注册时加前缀,比如amap_search、tavily_search。
5. A2A 多 Agent 协作:Server 与 Client 配置
A2A 解决的是 Agent 之间怎么互相调用。一个 Agent 把自己的能力通过 AgentCard 暴露出去,另一个 Agent 通过 A2A Client 发任务、收流式结果。
先看 Server 端。核心是实现 TaskProcessor 接口:
type TaskProcessor interface { Process(ctx context.Context, taskID string, initialMsg protocol.Message, handle TaskHandle) error }你的 Agent 逻辑就写在 Process 里。比如深度搜索 Agent:
func (a *DeepResearchAgent) Process(ctx context.Context, taskID string, msg protocol.Message, handle TaskHandle) error { handle.UpdateStatus(protocol.TaskStateWorking, "开始分析问题") // 调用运行图执行 stream, err := a.graph.Stream(ctx, map[string]any{ "query": msg.Parts[0].Text, }) if err != nil { return err } for { chunk, err := stream.Recv() if err == io.EOF { break } if err != nil { return err } handle.AddArtifact(protocol.NewTextPart(chunk.Content)) } handle.UpdateStatus(protocol.TaskStateCompleted, "完成") return nil }然后创建 A2A Server 并注册到 tRPC:
agentCard := loadAgentCard("./settings.json") taskManager, _ := redistaskmanager.NewRedisTaskManager(redisCli, agent) srv, _ := server.NewA2AServer(agentCard, taskManager) s := trpc.NewServer() thttp.RegisterNoProtocolServiceMux( s.Service("trpc.a2a.deepresearch.A2AServerHandler"), srv.Handler(), ) s.Serve()Client 端更简单,三步:创建客户端、发任务、收流。
a2aClient, err := a2aclient.NewA2AClient("http://127.0.0.1:10000", a2aclient.WithTimeout(10*time.Minute)) if err != nil { log.Fatal(err) } taskChan, err := a2aClient.StreamTask(ctx, protocol.SendTaskParams{ ID: uuid.New().String(), Message: protocol.Message{ Role: "user", Parts: []protocol.Part{protocol.NewTextPart("帮我规划深圳到厦门5天行程")}, }, }) if err != nil { log.Fatal(err) } for v := range taskChan { switch event := v.(type) { case protocol.TaskStatusUpdateEvent: fmt.Printf("状态: %s, 消息: %s\n", event.Status.State, event.Status.Message) case protocol.TaskArtifactUpdateEvent: fmt.Printf("结果片段: %s\n", event.Artifact.Parts[0].Text) } }Supervisor 模式的意图识别 Agent 就是靠这个 Client 把任务分发给下游专家 Agent 的。它先用 Function Call 判断用户意图,然后在工具调用节点里通过 A2A Client 把任务转给对应的 Agent,流式接收结果再返回给用户。
6. 验证请求与成功结果
配置写完,启动顺序很重要。先起 MCP Server(如果是独立进程),再起下游 Agent,最后起 Supervisor。
验证分三步:
第一步,单独测 MCP 工具。写一个最小 main 函数,只做 ConnectMCP 和 ListTools,打印工具列表。如果这一步失败,检查 command 路径和 args 是否正确。
go run ./cmd/mcp-test # 期望输出: connected, tools: [search, route, weather]第二步,单独测 A2A Server。启动后访问 AgentCard 端点:
curl http://127.0.0.1:10000/.well-known/agent.json # 期望输出: {"name":"deepresearch","skills":[...]}第三步,端到端测。用 A2A Client 发一个任务,观察流式输出:
go run ./cmd/a2a-client --task "深圳到厦门5天怎么玩" # 期望输出: # 状态: working, 消息: 开始分析问题 # 结果片段: 正在检索厦门景点... # 结果片段: 推荐行程如下... # 状态: completed, 消息: 完成成功标志是:状态从 working 流转到 completed,中间有 artifact 片段输出。如果卡在 working 不动,大概率是下游 Agent 没起来或者 URL 配错了。
7. 本篇常见错排查
错误一:connect mcp failed: exec: "npx": executable file not found
这是 stdio 传输找不到命令。Golang 的 exec 不会走 shell 的 PATH 解析,你需要写绝对路径,或者确保 npx 在系统 PATH 里。用which npx确认路径,然后改成/usr/local/bin/npx。
错误二:401 Unauthorized或invalid api key
检查 TAOTOKEN_API_KEY 环境变量是否真的注入到进程里了。用os.Getenv打印一下,如果是空字符串,说明 export 没生效或者启动方式没继承环境变量。另外确认 BaseURL 是https://taotoken.net/api,不要带多余路径。
错误三:A2A 任务一直卡在 submitted 状态
通常是 TaskManager 的 Redis 连接失败,任务状态写不进去。检查 config.toml 里的 checkpoint_addr 是否正确,Redis 是否在跑。用redis-cli ping确认。
错误四:工具调用返回tool not found
模型返回的 ToolCall 名字和你注册的 ToolInfo 名字对不上。打印一下 BindTools 之前的 toolInfos 列表,确认名字一致。如果用了多个 MCP Server,注意前缀。
错误五:流式输出中断,报context deadline exceeded
超时设太短。复杂任务比如深度搜索,可能需要几分钟。把 config.toml 里的 timeout_seconds 调到 300,A2A Client 的 WithTimeout 也相应调大。
错误六:panic: interface conversion: interface {} is nil
State 没有初始化。检查WithGenLocalState是否返回了非 nil 的 state 对象,以及 Checkpoint 恢复时 state 是否正确反序列化。
8. 继续深入的方向
跑通上面这套骨架之后,你可以往几个方向加东西。一是加 Callbacks 做可观测,把每次模型调用和工具调用的输入输出打到日志或者 Langfuse,这样排查问题不用靠猜。二是加 Checkpoint 做 Human in the Loop,在工具调用前暂停等用户确认,适合涉及写操作的场景。三是把 Connector 层补上,让 Agent 能接入 OpenAI 兼容的客户端,这样调试的时候可以直接在桌面应用里对话。
如果你要长期跑编码类 Agent,可以了解 Coding Plan 的接入方式: https://taotoken.net/coding-plan?utm_source=taotoken_aicg_blog_end&utm_content=coding-plan&utm_campaign=rewrite 。Claude Code 相关的 Anthropic 通道配置在 https://taotoken.net/claude-code?utm_source=taotoken_aicg_blog_end&utm_content=claudecode&utm_campaign=rewrite 。遇到接入问题先翻文档: https://taotoken.net/doc?utm_source=taotoken_aicg_blog_end&utm_content=doc&utm_campaign=rewrite ,Key 管理在 https://taotoken.net/api-keys?utm_source=taotoken_aicg_blog_end&utm_content=api-keys&utm_campaign=rewrite 。
最后说一个实际踩过的坑:MCP 工具的参数 schema 如果是嵌套结构,Golang 这边反序列化容易丢字段。解决办法是在 ToolInfo 里把 InputSchema 原样传进去,不要自己手动拼 ParamsOneOf。模型返回的 arguments 是 JSON 字符串,直接透传给 MCP Client 执行,中间不要做二次解析。这样能避免大部分参数对不上的问题。