☰
semantic-router Vela 1.0 文本模型 parity 验证:在 CPU 与 ROCm 上对齐旧路由路径的完整记录
2026/10/12 1:21:24 网站建设 项目流程
  • 后端
  • API网关
  • 模型推理服务
  • AI Agent

【免费下载链接】semantic-router

An open, programmable decision layer for models and compute.

项目地址:https://gitcode.com/gh_mirrors/sem/semantic-router
点击查看免费下载

导读

本文基于开源仓库 semantic-router 的记录文档,完整讲解 Vela 1.0 十款文本模型(Domain、Guard、Safety、Shield、FactCheck、Feedback、Modality、Hazard、PII、Halu)在模型运行时(runtime)中与旧路由器(legacy router path)的逐输入等价性验证方案。文章覆盖验证方法论、CPU 与 ROCm(AMD MI325X)两套硬件的完整结果表、Shield 模型 ONNX 图的异常发现、近似 profile(batching、max_speed)的取舍,以及可复现的完整命令行,读者可以据此理解并复现一次"替换前先对齐"的模型运行时迁移验收流程。

一、验证目标:task_heads 家族与它替换的旧路由路径

Vela 1.0 的task_heads家族在src/model-runtime/vllm_srun/registry/tables/vela1.py中被定义为family="task_heads",包含十款文本模型(Domain、Guard、Safety、Shield、FactCheck、Feedback、Modality、Hazard、PII、Halu),外加 Vela Embedding、Vela Reranker 与 Qwen3-Embedding 三个内置。每款文本模型都是一个 FP32 的 ModernBERT 任务 checkpoint(约 3.07 亿参数,loaded_parameters在 307,531,778~307,557,155 之间),表中的条目同时钉死了:

  • 仓库与修订号(revision),如 Domainf6354f54、PII6d3300c4、Haluca875312;
  • 每个文件(config.json、model.safetensors、tokenizer.json等)的SHA-256 摘要,即"该家族加载什么、下载什么"都由摘要定义;
  • 模型身份(model_sha256)与参数量;
  • ROCm 就绪性检查所用的黄金答案,存放在src/model-runtime/vllm_srun/registry/golden_answers_vela1.json。

parity 验证(即本记录文档)的目标很明确:runtime 在同样的输入上、在 CPU 与 ROCm 上,达到设计文档 section 17 规定的对齐门槛,以证明它可以替换旧路由路径。旧路由路径即路由器自身的原生门面(native facade),位于旧提交61aa7eb2d的pkg/modelruntime/native(即/api/v1/diagnostics/models/*背后的代码);runtime 侧则是Runtime.call在 HTTP 服务器中的调用方式(src/model-runtime/vllm_srun/runtime.py中Runtime.call按(surface, body, size)提供服务,小请求在事件循环内联规划)。所有原始结果(每一条分歧、上下文、阈值)记录在src/model-runtime/docs/records/vela1-parity.json。

二、两侧部署与实验设置

Legacy 侧:legacy_parity.py legacy

Legacy 侧由src/model-runtime/tools/legacy_parity.py驱动:一个编译进pkg/modelruntime/native的 Go 测试(zz_legacy_parity_dump_test.go,由脚本写入旧提交的精确镜像副本),通过门面自己的任务构造器加载每个模型并记录每个答案与延迟。任务构造器覆盖了门面的全部形态:

  • sequence:序列模型,softmax over labels;
  • sequence_windows:带滑窗的序列模型(Guard);
  • operating_point:Hazard 的独立 sigmoid + 打包的工作点(operating point);
  • tokens:BIO token 跨度(PII 的截断形态);
  • token_windows:带滑窗的 token 跨度(PII);
  • grounded:基于给定上下文的答案跨度(Halu)。

对应的RECIPES字典(CPU_JOBS/AMD_JOBS)把每个 job 映射到(模型名, 模式, max_tokens, overflow 策略, 窗口, 执行后端)元组,REVISIONS字典与registry/tables/vela1.py钉住的修订号一一对应。部署有两套:

  • CPU(路由器默认):candle 后端。sequence 模型max_tokens512(截断);Guard 与 PII 用 512/255 的窗口分别覆盖 8,192 / 32,768 token;Hazard 使用打包的工作点(2,048/1,023 窗口覆盖 32,768);Halu 8,192(截断)。绑定(bindings)在golang:1.25-bookworm中构建——因为61aa7eb2d时代的路由器镜像依赖 glibc 2.39,自身无法启动。
  • AMD(config/recipes/vela-amd):ONNX Runtime on MIGraphX(Guard 走 ROCm EP 及其 8K 图),通过镜像内的libonnx_semantic_router.so。该 recipe 为每个模型编译一个固定的 8,192-token session,因此每次请求都是 8K forward。recipe 不含 Shield;Halu 不发布 ONNX 图,因此它没有 legacy ORT 路径。

Runtime 侧:legacy_parity.py runtime --inputs

Runtime 侧在进程内用Runtime服务同样的包(ModelConfig(model=repo, name=job, device=device, profile=profile),result_cache_entries=0),并直接通过表面 API 回答:

  • 相同的部署选项,作为请求的options(max_tokens、overflow、可选的window{tokens, overlap})下发;
  • 每次一个请求、每个输入两次(第一次的答案用于比较,第二次计时),随后每个模型 4 个闭环调用者持续 20 s;
  • exactprofile、FP32;
  • CPU:16 个 EPYC 9575F 核,用taskset绑定(节点 B 64–79;torch 2.10 CPU,oneDNN packed linears)。这批延迟不是验收门槛——vela1-performance.md在同 cgroup 作用域内交错计时两侧;
  • ROCm:一块 MI325X,运行在路由器的 ROCm 镜像栈内(Python 3.12.14、PyTorch 2.12.0+rocm7.2、HIP 7.2.53211、Triton 3.7.0、fla-core 0.5.2),以 uid 65532 运行,启用 encoder 图与融合 rotary kernel。

输入语料

  • 每个 sequence / token job 有547 条文本:e2e 测试数据(e2e/testcases/testdata)中的 531 条提示词 + 六个边界用例 + 十条长文档(8–384 条提示词拼接,约 0.5–48K 字符)。边界用例包括:混合文字与 emoji、联系方式、字面<bos>/<eos>/[CLS]字符串、单字符、首尾空白、德语 PII。
  • Halu 使用64 个 grounded 三元组:60 个上下文为 3–40 条提示词,4 个为 400 条。
  • runtime 运行回答 legacy 运行记录的输入(--inputs,即.jobs.json),保证两侧输入完全一致。输入摘要:7f23ef84(CPU 部署)、8d5a8a55(AMD 部署)、f7475ac8(Shield on ORT,同样的 547 条文本)。

验收门槛(Bars,设计 section 17)

阈值直接固化在legacy_parity.py的常量中:

门槛CPUROCm
标签一致率100%(近并列除外,即 top-two 差值 < 1e-3 视为并列)≥ 99.5%
最大概率差 maxΔp≤ 1e-3≤ 0.02
相同跨度集合至少 99.5% 的输入≥ 98%

跨度集合比较的是标签与码点偏移(code-point offsets,legacy 侧报告 UTF-8 字节,compare阶段转换为码点);跨度概率只在集合一致处比较。compare_job的实现逻辑与此对应:对 sequence/scores 类 job 计算逐项 |Δp| 的最大值并判断标签(含近并列容忍),对 token/grounded 类 job 比较跨度集合。

三、任务矩阵(Jobs)

Job模型HeadCPU 部署AMD 部署
domain, safety, factcheck, feedback, modalitysequence 模型softmax over labels512 处截断超过 8,192 拒绝
shieldShieldsoftmax over labels512 处截断recipe 之外;从其 ONNX 图走 ORT MIGraphX,仅计时
guardGuardsoftmax over labels512/255 窗口覆盖 8,192,逐标签取 max超过 8,192 拒绝(ROCm EP)
hazardHazard独立 sigmoid + 打包工作点2,048/1,023 窗口覆盖 32,768相同
piiPIIBIO token 跨度512/255 窗口覆盖 32,768超过 8,192 拒绝
pii_truncatePIIBIO token 跨度512 处截断(门面的部分结果)—
haluHalugrounded 答案跨度("User request: …", answer)对,8,192 处截断—(无 ONNX 图)

注意AMD_JOBS中 Guard 的 head 是onnx/model_rocm_8k.onnx(ROCm EP 的 8K 图),Shield 行带注释:它不在 recipe 中,其包的 ONNX 图回答方式与 checkpoint 不同(详见下文),因此该 job 只计时 ORT。KIND字典把六种模式归并为 sequence / scores / token / grounded 四类比较语义。

四、CPU 结果

下表为两侧各自单独运行时的延迟(legacy 提前数小时、runtime 在节点 B 64–79 与交错 A/B 并列)。vela1-performance.md在同一批核心上交错对比两侧,并给出吞吐量。

Job比较数(双方均拒绝)一致率Max abs Δp门槛p50 legacy → runtime (ms)p95 legacy → runtime (ms)
domain547 (0)100.00%1.5e-04pass40.38 → 8.21213.65 → 25.19
guard545 (2)100.00%2.9e-05pass37.28 → 7.98210.52 → 23.61
safety547 (0)100.00%3.5e-05pass38.05 → 8.07215.50 → 24.64
shield547 (0)100.00%4.6e-05pass34.35 → 7.56218.32 → 24.60
factcheck547 (0)100.00%2.2e-05pass35.58 → 7.69232.02 → 23.46
feedback547 (0)100.00%9.7e-05pass36.27 → 7.44209.35 → 23.05
modality547 (0)100.00%5.8e-06pass32.57 → 7.71213.46 → 23.68
hazard547 (0)100.00%9.6e-06pass33.93 → 9.56213.88 → 28.67
pii547 (0)100.00%9.2e-05pass34.59 → 10.47211.88 → 30.11
pii_truncate547 (0)100.00%7.2e-05pass82.45 → 10.06244.36 → 27.16
halu64 (0)100.00%4.4e-05pass1424.37 → 110.3838738.26 → 2808.69

要点:全部 11 个 job 通过门槛,最大概率差仅 1.5e-04(domain);延迟上 runtime 全面领先,最悬殊的是 halu——p50 从 1,424 ms 降到 110 ms,p95 从 38.7 s 降到 2.8 s。性能侧的完整交错对比(含 95% 置信区间与吞吐量)见src/model-runtime/docs/records/vela1-performance.md。

五、ROCm 结果

5.1 对 AMD recipe(ONNX Runtime MIGraphX / ROCm EP,同一 GPU)

Job比较数(双方均拒绝)一致率Max abs Δp门槛p50 legacy → runtime (ms)p95 legacy → runtime (ms)
domain545 (2)100.00%1.8e-04pass154.06 → 1.88155.91 → 3.51
guard545 (2)100.00%4.7e-05pass245.81 → 1.88258.07 → 3.53
safety545 (2)100.00%3.5e-06pass155.26 → 1.82164.00 → 3.51
factcheck545 (2)100.00%1.1e-04pass158.24 → 1.82163.86 → 3.48
feedback545 (2)100.00%1.1e-04pass157.93 → 1.89164.11 → 3.55
modality545 (2)100.00%5.5e-06pass158.56 → 1.83167.20 → 3.50
hazard547 (0)100.00%2.5e-06pass13.38 → 1.8414.13 → 3.53
pii545 (2)98.90%2.4e-05pass134.03 → 1.93136.05 → 3.70

除 pii 的跨度一致率 98.90%(仍高于 ROCm 的 98% 跨度门槛)外全部 100%;legacy ORT MIGraphX 每个请求都是固定的 8K forward,所以 p50 在 13–246 ms 之间,而 runtime 单请求 p50 仅约 1.8–1.9 ms。

5.2 重大发现:Shield 的 ONNX 图不是它的 checkpoint

runtime 对 Shield 打包的onnx/model.onnx(ORT MIGraphX,547 个输入:545 个参与比较、2 个双方拒绝)未达门槛:标签一致率 97.80%、max |Δp| 0.81;339 个输入差值超过 0.02,12 个标签翻转。归因是图本身,而非 runtime:

  • legacy ORT 对 legacy candle,在同一批输入上翻转 14 个标签,max |Δp| 0.85、中位数 |Δp| 0.026;
  • 同一条 MIGraphX 路径在其余七个模型上与 runtime 一致到 ≤ 1.8e-4;
  • candle 与 runtime 都服务model.safetensors:它们在 CPU 上一致(max |Δp| 4.6e-5)、在 ROCm 上一致(100%,max |Δp| 4.7e-5,见下)。

结论:发布的 ONNX 导出回答起来像一个不同的模型——这很可能是vela-amdrecipe 从未在 ORT 上部署 Shield 的原因。该行仅以计时身份保留在vela1-performance.md中。registry/tables/vela1.py中 classify 类包声明了 exit 图(Embedding/Reranker 有onnx/model.onnx),而文本 classify 包没有 exit 图,所以 runtime 的onnxruntimeengine 也无法服务这张图。

5.3 对 CPU legacy 参考(legacy candle on CPU、runtime on MI325X)

Job比较数(双方均拒绝)一致率Max abs Δp门槛p50 legacy → runtime (ms)p95 legacy → runtime (ms)
domain547 (0)100.00%3.6e-04pass40.38 → 1.89213.65 → 3.53
guard545 (2)100.00%5.0e-05pass37.28 → 1.90210.52 → 3.56
safety547 (0)100.00%3.5e-05pass38.05 → 1.79215.50 → 3.39
shield547 (0)100.00%4.7e-05pass34.35 → 1.82218.32 → 3.48
factcheck547 (0)100.00%7.1e-05pass35.58 → 1.82232.02 → 3.49
feedback547 (0)100.00%9.2e-05pass36.27 → 1.88209.35 → 3.64
modality547 (0)100.00%8.9e-06pass32.57 → 1.84213.46 → 3.52
hazard547 (0)100.00%1.0e-05pass33.93 → 1.83213.88 → 3.50
pii547 (0)100.00%1.5e-04pass34.59 → 1.95211.88 → 3.74
pii_truncate547 (0)100.00%6.3e-05pass82.45 → 1.92244.36 → 3.66
halu64 (0)100.00%7.2e-05pass1424.37 → 7.7238738.26 → 56.15

runtime 在 MI325X 上把 CPU 延迟从几十毫秒压缩到 1.8–2.0 ms,halu 的 p95 从 38.7 s 降到 56 ms——这正是"从路由器 CPU 默认迁移到 GPU"的收益量化。

六、分歧分析

  • CPU:无分歧。11 个 job 的每个标签与每个跨度集合完全相同,最大概率差 1.5e-04。
  • ROCm 对 AMD recipe:PII 跨度在 545 个输入中的 6 个上不同(p0103、p0246、l01、l04、l06、l07)。每个案例中 runtime 都多报一个单码点跨度,而 ONNX Runtime MIGraphX 路径丢弃了它:电话号码的一位数字,或街道地址/组织名称的一个字符,即位于其标签边界的 token。对照 CPU 参考,runtime 的 ROCm 跨度在每个输入上都相同,因此差异来自 legacy GPU 路径的数值。
  • ROCm 对 CPU 参考:无分歧。
  • 拒绝(rejections):两篇超过 8,192 token 的文档,在"部署要求拒绝"的地方被双方一致拒绝:AMD recipe,以及 CPU 上 Guard 的 8,192-token 窗口上限。pii_truncate比较的是门面的部分结果(截断扫描及其跨度)与 runtime 的input.truncated答案。

七、路由器 ROCm 镜像栈与 release 镜像的对比

自a580be6b9起,路由器的 ROCm 镜像使用 release 镜像的 PyTorch 与 ROCm 库(详见src/model-runtime/docs/records/rocm-router-image.md)。在已发布(slim)镜像31d00387c中,vela1自己的 ROCm runner 将 AMD recipe 的全部 4,376 个答案与 release 镜像逐字节一致(8 个 job × 547;legacy_parity.py compare --baseline runtime一致率 1.0、max |Δ| 0.0),并且就绪性检查对全部 13 个task_heads内置(十款文本模型 + Vela Embedding + Vela Reranker + Qwen3-Embedding)的黄金答案与提交文件逐值相等。

在此之前,镜像安装的是官方 PyTorch wheel。此前的 ROCm 记录运行在包的 release 镜像中(那里的黄金答案被记录):

  • 路由器镜像栈:PyTorch 2.12.0+rocm7.2、HIP 7.2.53211、Triton 3.7.0、fla-core 0.5.2;
  • release 镜像:PyTorch 2.12.0+git6bbd260(源码构建)、HIP 7.2.53211、Triton 3.7.1、fla-core 0.5.2。

同一份代码下,路由器镜像栈给出的标签与跨度与 release 镜像一致:多数答案逐字节相同,其余仅在最后一位上不同:

Job输入答案标签与跨度逐字节相同Max abs Δp
domainAMD recipe547547464 (84.8%)3.5e-06
guardAMD recipe547547471 (86.1%)1.3e-05
safetyAMD recipe547547474 (86.7%)1.3e-06
factcheckAMD recipe547547481 (87.9%)3.1e-06
feedbackAMD recipe547547465 (85.0%)9.0e-06
modalityAMD recipe547547468 (85.6%)1.2e-07
hazardAMD recipe547547462 (84.5%)1.1e-06
piiAMD recipe547547478 (87.4%)4.3e-05
allAMD recipe4,3764,3763,763 (86.0%)4.3e-05
domainCPU recipe547547462 (84.5%)6.1e-06
guardCPU recipe547547469 (85.7%)3.8e-05
safetyCPU recipe547547469 (85.7%)1.3e-06
shieldCPU recipe547547471 (86.1%)5.7e-07
factcheckCPU recipe547547483 (88.3%)3.1e-06
feedbackCPU recipe547547463 (84.6%)5.6e-06
modalityCPU recipe547547465 (85.0%)1.2e-07
hazardCPU recipe547547462 (84.5%)1.1e-06
piiCPU recipe547547476 (87.0%)2.1e-04
pii_truncateCPU recipe547547476 (87.0%)3.4e-05
haluCPU recipe646420 (31.2%)3.3e-05
allCPU recipe5,5345,5344,716 (85.2%)2.1e-04

差异集中在较长的输入上:不超过 200 字符的输入,每个 job 的 450 个中有 446–447 个逐字节相同;不同的答案其输入长度中位数为 379–446 字符,Halu 的 grounded 输入则长达数千字符——两个构建的 kernel 对某些更大形状的舍入方式不同。

  • 黄金答案:就绪性检查的 ROCm 参考(src/model-runtime/vllm_srun/registry/golden_answers_vela1.json)在此栈上对全部 13 个task_heads内置逐字节相同;三个全新进程给出它们(第一个进程带空 Triton 缓存)。release 镜像也逐字节给出,因此没有任何黄金答案被重新记录。
  • 跨进程:该栈上不同进程逐字节一致回答:AMD-recipe 输入的 4,376 个答案在七个进程中一致(vela1-performance.md的五轮 A/B +efb5ec4d7与35ff3a8a5的 parity 运行);CPU-recipe 输入的 5,534 个答案与 Shield 的 547 个在三个进程中一致(第三个进程带空 Triton 缓存)。
  • 镜像成品:从Dockerfile.extproc在be7366c49构建的路由器 ROCm 镜像,除了该栈的 pin 外还带causal-conv1d1.7.0(这些模型并不运行它)、HTTP 服务器的uvloop与httptools、Pillow,以及 Python 3.12.15(代替 3.12.14)。它给出每个黄金值(三个进程,第一个带空 Triton 缓存)与三个 parity 面板的每个答案,与本栈逐字节一致。
  • 质量门槛:Decision Index 不评分任何 Vela 1.0 模型,因此没有 Index 差值可测;Vela 1.0 的参考就是 legacy 答案,在设计 section 17 的门槛之下,此栈在每个 job 上都通过(上文)。

八、近似 profile:batching 与 max_speed(BF16)的取舍

batching

batching合并并发请求;单请求与exact行为相同、答案一致(35ff3a8a5,ROCm,AMD-recipe 输入:逐请求每个值都相同)。它的收益与代价、以及exact自身如何在不改变答案的前提下合并并发(packed linears、rowwise对齐的 GeGLU 与 score sigmoid、无 padding 行的 unmasked grids、每 grid 宽度一张 rotary 表、单长度 grid)的细节见vela1-performance.md。

max_speed(BF16 副本)

max_speed的降精度副本(设计 section 5.4)在任何同意之前、以整份语料对exact记录。一致率统计标签或相同跨度集合。

MI325X 上的 BF16 副本(head 保持 FP32),对同一批输入上的exact(973842d9e):

Job输入与 exact 一致率Max abs Δp地板(99%)
domainAMD-recipe99.82%1.1e-01pass
guardAMD-recipe100.00%6.3e-02pass
safetyAMD-recipe99.82%2.0e-02pass
factcheckAMD-recipe99.82%2.3e-01pass
feedbackAMD-recipe100.00%3.0e-01pass
modalityAMD-recipe100.00%9.5e-02pass
hazardAMD-recipe100.00%1.4e-02pass
piiAMD-recipe95.23%4.9e-02fail
domainrouter CPU100.00%1.1e-01pass
guardrouter CPU100.00%3.5e-01pass
safetyrouter CPU99.82%2.0e-02pass
shieldrouter CPU100.00%2.1e-02pass
factcheckrouter CPU99.82%2.3e-01pass
feedbackrouter CPU100.00%3.0e-01pass
modalityrouter CPU100.00%9.5e-02pass
hazardrouter CPU100.00%1.4e-02pass
piirouter CPU94.70%4.9e-02fail
pii_truncaterouter CPU95.06%4.9e-02fail
halurouter CPU75.00%2.6e-02fail

CPU 上的 BF16 副本(AVX-512 BF16,16 个 EPYC 9575F 核,head FP32),对路由器 CPU 输入上的exact(0a2ced483,均在节点 B 64–79):

Job与 exact 一致率Max abs Δp地板(99%)p50 exact → copy (ms)p95 exact → copy (ms)4 调用者,calls/s
domain100.00%7.6e-02pass7.68 → 12.1022.89 → 33.5484.9 → 102.1
guard100.00%1.3e-01pass7.35 → 11.6722.24 → 31.2152.3 → 48.3
safety100.00%1.2e-02pass7.65 → 12.0123.84 → 31.9584.5 → 108.5
shield100.00%2.1e-02pass7.69 → 11.4423.41 → 30.8983.7 → 107.4
factcheck100.00%2.1e-01pass7.70 → 11.7223.57 → 31.6083.9 → 108.8
feedback100.00%3.0e-01pass7.64 → 11.9823.54 → 32.5884.0 → 106.4
modality100.00%3.7e-02pass7.62 → 11.7223.51 → 31.0384.3 → 107.1
hazard100.00%1.4e-02pass7.75 → 11.9323.51 → 31.1840.4 → 27.6
pii94.52%3.5e-02fail7.91 → 12.1023.52 → 32.5238.2 → 22.8
pii_truncate94.88%3.5e-02fail8.97 → 11.6925.81 → 31.3675.9 → 102.8
halu70.31%6.0e-02fail118.97 → 109.152456.87 → 2154.144.1 → 3.1

结论:该家族不同意任何降精度副本

registry/tables/vela1.py中没有任何 Vela 1.0 文本模型的BuiltinModel.reduced条目;按设计 section 5.4,一个副本需要 ≥ 99% 一致率且更快:

  • GPU:BF16 对 token head 与 grounded head 达不到 99% 地板。它单请求还更慢——forward 受 launch 限制,autocast 又为每个 linear 加了一次 cast。
  • CPU:BF16 保住了 sequence 与 scores head 的每个标签,但 PII 与 Halu 达不到地板:
    • 它单请求比exact慢约 1.5×,因为exact的 linear 已经走 oneDNN 的 packed FP32 kernel(Domain p50 7.68 → 12.10 ms);只有 Halu 的长 grounded 输入更快(p50 119 → 109 ms),而 Halu 恰恰不及格;
    • 它短输入上的更高吞吐来自max_speed最多 2 ms 的请求合并;batching不改任何数值就能获得同样的增益。
  • float32-packed副本会重复exact路径自身的 kernel,而 int8 会破坏 ModernBERT 的激活(见src/model-runtime/docs/records/decision1-performance.md)。

顺带一提,vela1-performance.md在 ROCm 上记录的 profile 对比同样印证:exactDomain p50/p95 1.88/3.51 ms、batching3.98/5.56 ms(但负载下 calls/s 从 413–439 升到 601–612)、max_speedBF16 副本 6.69/11.00 ms 且 PII/Halu 不及格——所以batching等待最多 2 ms 换约 40% 负载增益,而 BF16 副本"更慢且不及格",家族不采纳任何副本。

九、复现命令

src/model-runtime/tools/legacy_parity.py的子命令(corpus/legacy/build-legacy/ab/runtime/compare)在模块 docstring 中有完整定义。关键参数包括:--recipe cpu|amd、--cache(HF 缓存,须持有钉死的快照)、--jobs(逗号分隔的 job 列表,默认 recipe 的全部 job)、--repeats、--concurrency、--seconds、--limit、--inputs(用早期运行的.jobs.json精确回答其输入);legacy 侧额外有--tree(旧提交镜像副本)、--flat、--libs、--compile-cache;runtime 侧有--device cpu|rocm:0、--profile exact、--threads。compare的--device-class cpu|rocm决定使用CPU_THRESHOLDS还是ROCM_THRESHOLDS,--baseline runtime则启用REDUCED_THRESHOLDS(99% 一致率、near ties 计为并列、|Δp| 仅报告不限),用于"降精度副本或 profile 对 exact"的对比。

# legacy(bookworm userland,绑定在 61aa7eb2d 构建;AMD recipe 在 legacy ROCm 镜像内) python3 tools/legacy_parity.py legacy --recipe cpu|amd --tree <legacy tree> \ --cache <hf cache> --flat <dir> --repeats 2 --concurrency 4 --out legacy.jsonl # Shield 走 AMD recipe 的 ORT 路径,用 CPU 运行的输入 python3 tools/legacy_parity.py legacy --recipe amd --jobs shield \ --inputs legacy-cpu.jobs.json ... --out legacy-amd-shield.jsonl # runtime,用某次 legacy 运行的输入(其 job,除非 --jobs 指定其他) python3 tools/legacy_parity.py runtime --recipe cpu|amd --inputs legacy.jobs.json \ --device cpu|rocm:0 --cache <hf cache> --repeats 2 --concurrency 4 --out runtime.jsonl python3 tools/legacy_parity.py compare --legacy legacy.jsonl --runtime runtime.jsonl \ --device-class cpu|rocm --context '{"runtime_commit": "..."}' --record record.json # 副本或 profile 对 exact:--baseline runtime --legacy exact.jsonl

每条结果都是 JSON lines{job, id, result | error, latency_ns};compare生成format: vela1-legacy-parity/1的记录(含thresholds、inputs_sha256、context、逐 job 报告与全局passed),并打印每 job 的 PASS/FAIL、一致率、max |Δ| 与 p50/p95 延迟。

十、相关记录与进一步阅读

  • 本记录的原始 JSON:src/model-runtime/docs/records/vela1-parity.json
  • 性能侧(交错 A/B、吞吐、95% 置信区间):src/model-runtime/docs/records/vela1-performance.md
  • 模型表(修订号、摘要、参数量、黄金答案):src/model-runtime/vllm_srun/registry/tables/vela1.py与src/model-runtime/vllm_srun/registry/golden_answers_vela1.json
  • ROCm 镜像栈说明:src/model-runtime/docs/records/rocm-router-image.md
  • AMD recipe(配置与使用说明):config/recipes/vela-amd/config.yaml与config/recipes/vela-amd/README.md
  • Vela 2.0 与 embedding 的同类 parity 记录:src/model-runtime/docs/records/vela2-parity.md、src/model-runtime/docs/records/embed-parity.md

简而言之:这是一份"替换路由器前先证明等价"的工程档案——十款 Vela 1.0 文本模型在 CPU 与 ROCm 上逐输入对齐旧路径,除 Shield 的 ONNX 图(本身即异常模型)外全部通过 section 17 门槛;BF16 降精度副本因 PII/Halu 达不到 99% 地板而整体否决。这套legacy_parity.py驱动的方法论与命令行,可以直接复用于后续任何模型家族的迁移验收。

  • 后端
  • API网关
  • 模型推理服务
  • AI Agent

【免费下载链接】semantic-router

An open, programmable decision layer for models and compute.

项目地址:https://gitcode.com/gh_mirrors/sem/semantic-router
点击查看免费下载

相关推荐

上一篇:【亲测免费】 SpringBlade 微服务开发平台安装与配置指南
下一篇:SpringBlade 微服务开发平台推荐

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询