最近做的企微机器人,单轮问答没问题,但多轮对话经常断片——客户说"那个产品多少钱",机器人不知道"那个"指什么。问题出在上下文管理。这篇重点讲多轮对话背后调了哪些企微接口——消息接收、历史消息拉取、消息发送,上下文怎么在这些接口调用之间传递。
底层用的是Eyun 平台开放的企微 API,统一 POST+JSON,鉴权用 App Token 加 appid,响应封套{code, data, detail, message, time},code 为 0 成功。
消息接收:Webhook 拿当前消息
客户发消息通过 Webhook 回调到后端。回调体里有fromUin(客户)、toUin(机器人)、content(消息内容)、conversationId(会话 ID)。会话 ID 是上下文的 key——单聊是客户 uin,群聊是 roomId。
@app.route("/wx-api/webhook/", methods=["POST"]) def webhook(): payload = request.json["data"] conv_id = payload["conversationId"] from_uin = payload["fromUin"] content = payload["content"] appid = payload["appid"] # 拿历史消息做上下文 history = get_history(appid, conv_id, limit=10) reply_text = generate_reply(content, history) send_reply(appid, from_uin, reply_text) return "ok"conversationId必须是 JSON number 类型,不能传字符串,这是踩过的坑。
拉历史消息:上下文的原料
多轮对话的上下文来自历史消息。调消息模块拉取会话历史:
def get_history(appid, conversation_id, limit=10): resp = requests.post( f"{BASE}/wx-api/api/message/getHistoryMessageList", headers=HEADERS, json={"appid": appid, "conversationId": conversation_id, "limit": limit} ) msgs = resp.json()["data"]["list"] return [{"from": m["fromUin"], "content": m["content"], "time": m["time"]} for m in msgs]返回的消息按时间倒序,要反转成正序再喂给模型。limit控制上下文长度,太长 token 爆。接口字段和分页参数在Eyun 开发文档。
上下文窗口:只留最近 N 轮
历史消息全给模型 token 爆炸,要滑动窗口——只留最近 N 轮(一来一回算一轮):
def build_context_window(history, rounds=5): # history 是正序 recent = history[-rounds*2:] # 每轮2条 return "\n".join([f"{'客户' if m['from']!='bot' else '机器人'}:{m['content']}" for m in recent])窗口外的关键信息(客户提到的产品、订单号)提取出来存对话状态,不丢。
长对话摘要:窗口外的处理
对话超过窗口怎么办?不能简单截断,要摘要压缩。用 AI 把旧消息生成摘要,存数据库:
def compress_history(history): old = history[:-20] # 窗口外的 if len(old) < 5: return "" text = "\n".join([m["content"] for m in old]) summary = llm.summarize(text) return summary上下文 = 摘要 + 近期窗口消息。这样既不丢关键信息,又不爆 token。
指代消解:从历史里找"那个"
客户说"那个多少钱",要从历史消息里找最近提到的实体。从拉到的历史里提取:
def resolve_reference(current_msg, history): # 从历史里找最近提到的产品/订单 for m in reversed(history): entities = extract_entities(m["content"]) # 产品名、订单号 if entities: return entities[0] return None指代消解在应用层做,企微接口只负责提供历史消息。历史拉得全,消解才准。
话题切换:重新拉历史
客户换话题了("对了,售后电话多少"),旧上下文不该带。检测到话题切换时,上下文只留当前消息,历史重新开始累积。
话题切换靠意图识别——新消息意图和当前对话意图差异大,判定切换。切换后get_history的 limit 设小(只留 1-2 条),避免旧话题干扰。
回复发送:调 sendText 回客户
生成回复后调消息模块发回去:
def send_reply(appid, to_uin, text): resp = requests.post( f"{BASE}/wx-api/api/message/sendText", headers=HEADERS, json={"appid": appid, "to": to_uin, "content": text} ) return resp.json()["code"] == 0to是客户 uin,就是回调里的fromUin。回复也会进历史消息,下次get_history能拉到,形成闭环。
持久化:对话不能丢
对话状态(当前意图、已收集槽位、摘要)存数据库,按 conversationId 存。服务重启、客户隔几天回来,都能恢复:
def load_dialog_state(conversation_id): return db.get(f"dialog:{conversation_id}") def save_dialog_state(conversation_id, state): db.set(f"dialog:{conversation_id}", state, ex=30*86400) # 30天过期不持久化,服务重启客户对话全断。30 天无活动清空,控制存储成本。
写在最后
多轮对话上下文管理这套东西,本质是用好企微的消息接口——Webhook 收当前消息、getHistoryMessageList 拉历史做上下文、sendText 发回复。上下文管理(窗口、摘要、指代、切换)在应用层做,接口负责提供数据。接口路径、参数、会话 ID 规范在Eyun 开发文档里。把接口调对、上下文管好,多轮对话就不跑偏。凭证和接入地址在Eyun 企业微信 API 平台开通。