1. 这不是“温度监控”那么简单:thermal framework 是内核里最被低估的功耗调度中枢
你翻过 Linux 内核源码树,大概率在drivers/thermal/目录下扫过几眼——目录名很直白,但里面代码的复杂度远超“读个温度、降个频”这种表面理解。我第一次在 ARM64 平台调试一个 SoC 热关机问题时,以为只要改改trip_point阈值就行,结果连续三天卡在thermal_zone_device_update()调用链里,发现它根本不是独立模块,而是像一张网,从硬件传感器探针,串起 CPU idle governor、cpufreq、cpu cooling device、甚至 GPU 频率策略、内存带宽限制器,最后连到用户空间的thermald或power-profiles-daemon。thermal framework 的真实角色,是 Linux 内核功耗子系统里那个沉默的“中央调度员”:它不直接执行降频或关核,但它决定“谁该在什么时候、以什么力度、用哪种方式降温”,所有功耗调控动作都必须向它注册、受它仲裁、按它策略执行。
这个框架之所以常被误读,是因为它的名字太具象——“thermal”,让人本能聚焦在温度本身;而它的设计哲学恰恰相反:温度只是输入信号,功耗才是调控目标,热行为只是功耗约束下的外在表现。比如你在笔记本上跑stress-ng --cpu 8 --timeout 60s,风扇狂转,CPU 频率掉到 800MHz,表面看是“热了所以降频”,但内核实际执行的是:thermal zone 检测到THERMAL_TRIP_ACTIVE触发 → 启动step_wisecooling policy → 查询所有已注册的cooling device(cpufreq、cpu-idle、intel-rapl)→ 根据其cur_state和max_state计算可调空间 → 综合当前thermal_zone的passive_delay和polling_delay→ 最终下发set_cur_state(3)到 cpufreq cooling device → cpufreq driver 才真正调用__cpufreq_driver_target()修改频率。整个链条里,thermal framework 只负责“决策”和“分发”,不碰硬件寄存器,也不管频率算法细节。
这也是为什么标题强调“通用架构梳理”——它不是某个芯片厂商的私有驱动,而是内核为所有 SoC 提供的标准化接口层。无论你是高通骁龙、联发科天玑、全志 H6,还是 Intel Core i7,只要遵循这套架构注册 thermal zone 和 cooling device,就能复用同一套策略引擎、同一套 sysfs 接口、同一套用户空间交互协议。我见过太多嵌入式项目,工程师自己写一套“温度检测+硬编码降频”逻辑,结果在多核异构场景下失效,就是因为绕过了 thermal framework 的状态同步与竞争仲裁机制。它解决的从来不是“怎么读温度”,而是“当 8 个 CPU 核、2 个 GPU cluster、1 个 NPU 共享同一块散热片时,如何让它们不互相抢资源、不重复降频、不漏判热点”。这正是它成为功耗子系统核心枢纽的根本原因。
2. 架构拆解:五层结构,三层抽象,一个统一调度器
thermal framework 的代码看似松散,实则严格遵循分层抽象原则。我把它拆成五个逻辑层,每层解决一类问题,且层间依赖单向清晰——上层只调用下层接口,绝不反向渗透。这种设计保证了可移植性:SoC 厂商只需实现最底层的硬件适配,上层策略、用户接口、驱动集成全部复用内核标准代码。
2.1 第一层:硬件感知层(Hardware Sensing Layer)
这是整个框架的地基,负责把物理世界的温度信号数字化。它不直接操作传感器,而是通过struct thermal_zone_device_ops定义统一回调接口:
struct thermal_zone_device_ops { int (*get_temp)(struct thermal_zone_device *, int *); int (*set_trips)(struct thermal_zone_device *, int, int); int (*notify)(struct thermal_zone_device *, int); };get_temp():必须实现,返回当前温度(单位:毫摄氏度)。注意,这里返回的是原始值,不做任何滤波或校准——校准由上层策略或用户空间完成。set_trips():可选,用于动态设置 trip point 阈值。很多 SoC 的 thermal sensor 支持硬件比较器,触发中断后自动上报,此时此函数可为空。notify():可选,当硬件中断触发 trip 事件时调用,用于快速响应(如立即关断某路电源)。
我实测过全志 H6 的sun8i_ths驱动,它的get_temp()实际做了三件事:读取 ADC 原始值 → 查表转换为温度 → 应用芯片厂提供的二阶补偿公式。而 Intel 的x86_pkg_temp_thermal驱动则直接读 MSR 寄存器,省去 ADC 步骤。但对外暴露的get_temp()接口完全一致,上层无需关心差异。
提示:硬件层最大的坑是温度单位。内核强制要求单位为millidegree Celsius(m°C),即 25°C 必须返回 25000。曾有个项目因驱动返回 25(误以为是 °C),导致所有 trip point 判定失效——
trip=80000对应 80°C,但驱动返回 25,内核认为才 0.025°C,永远不触发降温。
2.2 第二层:热区抽象层(Thermal Zone Abstraction)
struct thermal_zone_device是核心数据结构,代表一个物理热域(如 “cpu-thermal”、“gpu-thermal”、“battery-thermal”)。它封装了:
- 温度采集源(指向 hardware sensing layer)
- Trip point 配置(数组,每个元素含 temperature、type、hysteresis)
- Cooling device 关联列表(
struct list_head cooling_devices) - 更新策略(
update_mode:THERMAL_DEVICE_ENABLED或THERMAL_DEVICE_DISABLED) - 延迟参数(
passive_delay,polling_delay)
Trip point 是关键设计。内核定义了 5 种类型:
THERMAL_TRIP_CRITICAL:不可逆动作,如强制关机(kernel_power_off())THERMAL_TRIP_HOT:主动降温起点,启动 cooling deviceTHERMAL_TRIP_PASSIVE:被动降温,通常触发 cpufreq 降频THERMAL_TRIP_ACTIVE:主动散热设备启动,如风扇全速THERMAL_TRIP_HOT:与PASSIVE类似,但优先级更高(部分平台用)
注意:CRITICAL和HOT是硬性保护,PASSIVE/ACTIVE是策略调控。我在调试 RK3399 板子时,发现PASSIVEtrip 设置为 70°C,但CRITICAL设为 125°C,中间留出 55°C 的缓冲带——这 55°C 就是 thermal framework 发挥策略调度的空间。
2.3 第三层:冷却设备抽象层(Cooling Device Abstraction)
struct thermal_cooling_device代表一个可调控的功耗单元。它不关心自己是什么,只提供两个核心能力:
get_max_state():返回最大可调级别(如 cpufreq cooling device 返回num_online_cpus() * 10,表示 0~10 级 per CPU)set_cur_state():设置当前级别(如state=5表示将 CPU 频率降至 50%)
内核预置了多种标准 cooling device:
cpufreq_cooling: 绑定到 cpufreq policy,通过cpufreq_frequency_table映射 state 到频率cpu_cooling: 控制 CPU online/offline 状态(state=0 全开,state=1 关 1 核...)power_allocator: 基于 PID 控制器动态分配功耗预算(用于 GPU/NPU)fan_cooling: 控制风扇 PWM 占空比
关键点在于:一个 cooling device 可被多个 thermal zone 共享。例如,cpu-thermalzone 和soc-thermalzone 都能绑定到同一个cpufreq_coolingdevice。framework 会自动聚合请求——如果 zone A 要求 state=3,zone B 要求 state=5,则最终取 max(3,5)=5。这避免了多 zone 竞争导致的过度降频。
2.4 第四层:策略引擎层(Governor Layer)
这才是 thermal framework 的“大脑”。它决定:当温度越过 trip point 后,如何选择 cooling device、如何调整 state、何时再次采样。内核提供 4 种内置 governor:
| Governor | 触发条件 | 调控逻辑 | 适用场景 |
|---|---|---|---|
step_wise | 温度持续高于 trip | 线性步进:每次 +1 state,直到达标 | 通用,默认选项 |
bang_bang | 温度越过 trip 瞬间 | 全力降温:直接设为 max_state,低于 trip 后全关 | 风扇控制,响应快 |
user_space | 用户空间写入cur_state | 完全由用户程序控制 | 自定义策略,如机器学习预测 |
power_allocator | 温度偏差 | PID 控制:output = Kp*e + Ki*∫e dt + Kd*de/dt | 精确功耗分配,GPU/NPU |
我强烈建议新手从step_wise入手。它的逻辑清晰:thermal_zone_device_update()被调用时,遍历所有 trip,对每个PASSIVE/ACTIVEtrip,计算当前温度与 trip 的差值delta = temp - trip_temp,然后根据delta和slope(斜率,单位 m°C/state)计算目标 state:target_state = (delta + slope/2) / slope。slope默认为 10000(即 10°C/state),意味着每升高 10°C,state +1。这个设计让调控平滑,避免抖动。
2.5 第五层:用户空间接口层(Userspace Interface)
全部通过 sysfs 暴露,路径为/sys/class/thermal/thermal_zoneX/。关键文件:
temp: 当前温度(m°C)type: zone 名称(如 "cpu-thermal")trip_point_[0-9]_temp: 各 trip 阈值trip_point_[0-9]_type: trip 类型policy: 当前 governormode: "enabled"/"disabled"emul_temp: 用于模拟测试(写入任意值覆盖temp)
cooling device 接口在/sys/class/thermal/cooling_deviceY/:
cur_state: 当前 statemax_state: 最大 statetype: device 类型("Processor", "Fan")power/allocated: 分配的功耗(仅power_allocator)
注意:
emul_temp是调试神器。不用烧板子,直接echo 90000 > emul_temp就能触发 90°C 逻辑,验证整个链路是否通畅。我调试时必做三步:1.cat temp确认读数正常;2.echo 95000 > emul_temp;3.watch -n 1 'cat cur_state'观察 cooling device 是否响应。三步下来,80% 的配置问题都能定位。
3. 实操详解:从零构建一个可用的 thermal zone(以 ARM64 SoC 为例)
假设你拿到一块新 SoC,文档里只有一行:“Thermal sensor located at APB bus offset 0x1200, 12-bit ADC, 0.5°C/LSB”。下面是我实际在瑞芯微 RK3328 上搭建 thermal zone 的完整过程,步骤可直接复用。
3.1 第一步:编写硬件驱动(drivers/thermal/rockchip_thermal.c)
核心是实现thermal_zone_device_ops:
static int rk3328_thermal_get_temp(struct thermal_zone_device *tz, int *temp) { struct rk3328_thermal_data *data = thermal_zone_device_priv(tz); u32 val; // 1. 使能 ADC 通道(寄存器操作) regmap_write(data->grf, GRF_SOC_CON1, BIT(12)); // 2. 触发 ADC 转换 regmap_write(data->pmu, PMU_ADC_CTRL, 0x1); // 3. 等待完成(轮询,实际项目建议用中断) usleep_range(100, 200); regmap_read(data->pmu, PMU_ADC_DATA, &val); // 4. 转换:12-bit 值,0.5°C/LSB,基准 25°C // 公式:T = 25 + (val - 0x800) * 0.5 *temp = 25000 + ((val & 0xfff) - 0x800) * 500; return 0; } static const struct thermal_zone_device_ops rk3328_tz_ops = { .get_temp = rk3328_thermal_get_temp, };关键细节:
usleep_range(100,200):ADC 转换需要时间,太短读到 0,太长影响 polling 效率。实测 100μs 足够。- 温度公式必须精确。
0x800是 2048,对应 25°C 基准点,500是 0.5°C 转为 m°C(0.5 * 1000)。 regmap是内核推荐的寄存器访问方式,比裸写ioremap更安全。
3.2 第二步:在 DTS 中定义 thermal zone
&tsadc { #thermal-sensor-cells = <1>; rockchip,grf = <&grf>; rockchip,pmu = <&pmu>; /* 定义 thermal zone */ cpu_thermal: cpu-thermal { thermal-sensors = <&tsadc 0>; /* 引用 tsadc 的 channel 0 */ polling-delay-passive = <1000>; /* 1s 被动轮询 */ polling-delay = <2000>; /* 2s 主动轮询 */ thermal-zone { /* Trip points */ temperature = <60000>; /* 60°C */ type = "cpu"; cooling-min-level = <0>; cooling-max-level = <15>; /* 关联 cooling device */ cooling-maps { map0 { cooling-device = <&cpu0_cooling 0 15>; cooling-device = <&cpu1_cooling 0 15>; cooling-device = <&cpu2_cooling 0 15>; cooling-device = <&cpu3_cooling 0 15>; }; }; }; }; };重点解析:
thermal-sensors = <&tsadc 0>:指定使用tsadc的第 0 个通道。&tsadc是前面定义的 ADC controller node。polling-delay-passive和polling-delay:前者用于PASSIVEtrip 触发后的高频轮询(如降频),后者是常态轮询间隔。数值单位是毫秒。cooling-maps:声明哪些 cooling device 属于此 zone。<&cpu0_cooling 0 15>表示绑定cpu0_coolingdevice,其 state 范围是 0~15。
3.3 第三步:注册 cooling device(cpufreq)
cpufreq cooling device 由drivers/thermal/cpu_cooling.c提供,只需在 cpufreq driver 中注册:
// 在 cpufreq driver init 函数中 struct thermal_cooling_device *cdev; cdev = of_cpufreq_cooling_register(policy); if (IS_ERR(cdev)) { pr_err("Failed to register cpufreq cooling device\n"); return PTR_ERR(cdev); }of_cpufreq_cooling_register()会自动解析 DTS 中cooling-maps的绑定关系,并创建cpu0_cooling等 device。它内部做了三件事:
- 读取
cpufreq_frequency_table,确定 frequency levels 数量(即max_state) - 将
state映射到frequency_table[state] - 注册
set_cur_state()回调,调用__cpufreq_driver_target()
3.4 第四步:验证与调试
编译烧录后,检查 sysfs:
# 查看 thermal zone 是否创建 ls /sys/class/thermal/ # 应看到 thermal_zone0(cpu_thermal) # 读取温度 cat /sys/class/thermal/thermal_zone0/temp # 输出类似 42350(42.35°C) # 查看 trip point cat /sys/class/thermal/thermal_zone0/trip_point_0_temp # 应为 60000 # 查看绑定的 cooling device ls /sys/class/thermal/thermal_zone0/cdev[0-9]* # 应看到 cdev0 -> ../cooling_device0(cpu0_cooling) # 强制触发降温(模拟高温) echo 70000 > /sys/class/thermal/thermal_zone0/emul_temp # 观察 cooling device state 变化 watch -n 1 'cat /sys/class/thermal/cooling_device0/cur_state' # state 应从 0 逐步升到 3、5...如果cur_state不变,按顺序排查:
dmesg | grep thermal:看是否有thermal zone registered或failed to register错误cat /sys/class/thermal/thermal_zone0/mode:确认是enabledcat /sys/class/thermal/thermal_zone0/policy:确认是step_wisecat /sys/class/thermal/cooling_device0/max_state:确认 cooling device 已注册
4. 深度避坑指南:那些文档里不会写的实战陷阱
我在 7 个不同 SoC 平台上部署 thermal framework,踩过足够多的坑,总结出 5 个最致命、最隐蔽的问题,每个都附带真实案例和解决方案。
4.1 陷阱一:Trip Point 的 hysteresis(迟滞)缺失导致抖动
现象:温度在 70°C 附近时,cur_state在 0 和 5 之间疯狂跳变,风扇“哒哒哒”响个不停。
原因:hysteresis参数未设置。内核默认hysteresis=0,意味着温度 ≥70°C 时触发PASSIVE,一旦降到 69.999°C 就立刻退出,没有缓冲带。
修复:在 DTS 中为 trip point 添加hysteresis:
thermal-zone { temperature = <70000>; type = "cpu"; hysteresis = <2000>; /* 2°C 迟滞 */ ... };原理:当温度 ≥70°C,触发 trip;只有温度 ≤68°C(70-2)时,才清除 trip。这 2°C 的窗口就是防抖空间。实测hysteresis=1000(1°C)对大多数 SoC 足够,2000更稳妥。
4.2 陷阱二:Cooling device state 映射错误引发“降频无效”
现象:cur_state显示为 10,但cpupower frequency-info显示频率仍是最大值。
原因:cpufreq_cooling的 state 映射表与实际frequency_table不匹配。常见于自定义 cpufreq table 时,table[i].frequency未按降序排列,或table[i].driver_data未正确设置。
诊断:查看cpufreqtable:
cat /sys/devices/system/cpu/cpufreq/policy0/scaling_available_frequencies # 输出:1200000 1000000 800000 600000 # 但你的 table 可能是:600000 800000 1000000 1200000(升序!)修复:确保frequency_table严格降序,且table[i].index = i:
static struct cpufreq_frequency_table rk3328_freq_table[] = { { .frequency = 1200000 }, // state 0 { .frequency = 1000000 }, // state 1 { .frequency = 800000 }, // state 2 { .frequency = 600000 }, // state 3 { .frequency = CPUFREQ_TABLE_END }, };4.3 陷阱三:Passive delay 过长导致“热失控”
现象:stress-ng跑 30 秒后,板子烫手关机,dmesg显示CRITICALtrip 被触发。
原因:polling-delay-passive设置过大(如 5000ms),而 SoC 温升速率快(1°C/s)。温度从 65°C 升到 85°C 只需 20 秒,但 thermal framework 每 5 秒才检查一次,错过最佳调控时机。
修复:根据 SoC 热特性设置合理 delay:
- 低功耗 SoC(Allwinner H3):
polling-delay-passive = <500>(500ms) - 高性能 SoC(RK3399):
polling-delay-passive = <200>(200ms) - x86 笔记本:
<100>(100ms)
实测法则:delay_ms < (trip_temp - current_temp) / (dT/dt)。例如当前 60°C,trip 70°C,温升率 2°C/s,则delay < 10°C / 2°C/s = 5s,设为 2s 更安全。
4.4 陷阱四:Multi-zone 竞争导致“过度降频”
现象:GPU 负载高时,CPU 频率被拉到最低,即使 CPU 自身温度很低。
原因:gpu-thermalzone 和cpu-thermalzone 共享同一个cpufreq_coolingdevice,且gpu-thermal的 trip 更激进(如 65°C),它要求state=10,而cpu-thermal只需state=2,最终取max(2,10)=10。
修复:为不同 zone 分配独立 cooling device,或使用power_allocatorgovernor 实现功耗隔离:
gpu_thermal: gpu-thermal { thermal-sensors = <&tsadc 1>; ... cooling-maps { map0 { cooling-device = <&gpu_cooling 0 15>; /* 独立 device */ }; }; };4.5 陷阱五:User space governor 的权限与同步问题
现象:thermald进程运行,但cur_state不更新。
原因:thermald需要CAP_SYS_ADMIN权限写cur_state,且必须先echo user_space > /sys/class/thermal/thermal_zone0/policy切换 governor。
验证:
# 检查权限 sudo getcap /usr/sbin/thermald # 应输出:/usr/sbin/thermald cap_sys_admin+ep # 检查 governor cat /sys/class/thermal/thermal_zone0/policy # 若为 step_wise,需先切换 echo user_space > /sys/class/thermal/thermal_zone0/policy更可靠的做法:在thermald配置文件/etc/thermald/thermal-conf.xml中,明确指定 zone:
<configuration> <thermal-zones> <thermal-zone type="cpu"> <kernels> <kernel type="pid"/> </kernels> <cooling-devices> <cooling-device type="Processor"/> </cooling-devices> </thermal-zone> </thermal-zones> </configuration>5. 架构演进与未来方向:从 thermal 到 power-aware scheduling
thermal framework 的设计初衷是热保护,但随着异构计算(big.LITTLE、CPU+GPU+NPU)普及,它正悄然演变为功耗感知调度(Power-Aware Scheduling)的核心基础设施。这不是我的猜测,而是内核社区正在发生的事实。
5.1 EAS(Energy Aware Scheduler)的深度集成
Android 12+ 的 EAS 调度器不再只看 CPU 负载,而是实时查询 thermal zone 的passive_delay和当前cur_state,动态调整util_avg(平均利用率)的权重。当cpu-thermalzone 的cur_state > 5,EAS 会认为“此 CPU cluster 功耗受限”,主动将新任务调度到 cooler 的 cluster 上。这要求 thermal framework 提供低延迟、高精度的cur_state查询接口——thermal_zone_get_cur_state()的响应时间必须 < 100μs,否则影响调度实时性。
5.2 Power Allocator Governor 的工业级应用
power_allocator不再是实验特性。在 NVIDIA Jetson Orin 上,它被用于 GPU 功耗闭环控制:thermal_zone读取 GPU junction temperature →power_allocator计算所需功耗 budget → 下发到nvidia,gpu-powercooling device → GPU driver 依据 budget 调整 shader clock 和 memory bandwidth。整个环路延迟 < 5ms,比传统step_wise精确 10 倍。
5.3 用户空间策略的爆发式增长
user_spacegovernor 正催生新生态:
power-profiles-daemon:根据电池模式(Performance/Balanced/Power Saver)动态修改 trip pointsthermald:集成机器学习模型,基于历史负载预测温度峰值,提前降频custom thermal daemon:游戏本厂商用它实现“性能模式”一键解锁全部功耗墙
这意味着,thermal framework 的未来,不再是内核里的一个“子系统”,而是连接硬件、内核、用户空间的功耗策略总线(Power Policy Bus)。你写的每一行 DTS,每一个trip_point,都在为这个总线注入策略基因。
我个人在实际项目中的体会是:不要把 thermal framework 当作“温度监控模块”来用,而要把它当作“功耗策略执行器”。它的价值不在于多精准地读出 0.1°C,而在于能否让你用最简洁的 DTS 描述,表达出“当 GPU 温度超过 75°C 时,将 CPU 频率限制在 1.2GHz,同时提升风扇至 60% 占空比”这样复杂的跨域协同策略。这才是通用架构的真正力量——它把硬件细节封装起来,把策略逻辑解放出来,让功耗管理从“工程师手动调参”走向“系统自动决策”。