Linux thermal framework:功耗调度核心架构解析 1. 这不是“温度监控”那么简单thermal framework 是内核里最被低估的功耗调度中枢你翻过 Linux 内核源码树大概率在drivers/thermal/目录下扫过几眼——目录名很直白但里面代码的复杂度远超“读个温度、降个频”这种表面理解。我第一次在 ARM64 平台调试一个 SoC 热关机问题时以为只要改改trip_point阈值就行结果连续三天卡在thermal_zone_device_update()调用链里发现它根本不是独立模块而是像一张网从硬件传感器探针串起 CPU idle governor、cpufreq、cpu cooling device、甚至 GPU 频率策略、内存带宽限制器最后连到用户空间的thermald或power-profiles-daemon。thermal framework 的真实角色是 Linux 内核功耗子系统里那个沉默的“中央调度员”它不直接执行降频或关核但它决定“谁该在什么时候、以什么力度、用哪种方式降温”所有功耗调控动作都必须向它注册、受它仲裁、按它策略执行。这个框架之所以常被误读是因为它的名字太具象——“thermal”让人本能聚焦在温度本身而它的设计哲学恰恰相反温度只是输入信号功耗才是调控目标热行为只是功耗约束下的外在表现。比如你在笔记本上跑stress-ng --cpu 8 --timeout 60s风扇狂转CPU 频率掉到 800MHz表面看是“热了所以降频”但内核实际执行的是thermal zone 检测到THERMAL_TRIP_ACTIVE触发 → 启动step_wisecooling policy → 查询所有已注册的cooling devicecpufreq、cpu-idle、intel-rapl→ 根据其cur_state和max_state计算可调空间 → 综合当前thermal_zone的passive_delay和polling_delay→ 最终下发set_cur_state(3)到 cpufreq cooling device → cpufreq driver 才真正调用__cpufreq_driver_target()修改频率。整个链条里thermal framework 只负责“决策”和“分发”不碰硬件寄存器也不管频率算法细节。这也是为什么标题强调“通用架构梳理”——它不是某个芯片厂商的私有驱动而是内核为所有 SoC 提供的标准化接口层。无论你是高通骁龙、联发科天玑、全志 H6还是 Intel Core i7只要遵循这套架构注册 thermal zone 和 cooling device就能复用同一套策略引擎、同一套 sysfs 接口、同一套用户空间交互协议。我见过太多嵌入式项目工程师自己写一套“温度检测硬编码降频”逻辑结果在多核异构场景下失效就是因为绕过了 thermal framework 的状态同步与竞争仲裁机制。它解决的从来不是“怎么读温度”而是“当 8 个 CPU 核、2 个 GPU cluster、1 个 NPU 共享同一块散热片时如何让它们不互相抢资源、不重复降频、不漏判热点”。这正是它成为功耗子系统核心枢纽的根本原因。2. 架构拆解五层结构三层抽象一个统一调度器thermal framework 的代码看似松散实则严格遵循分层抽象原则。我把它拆成五个逻辑层每层解决一类问题且层间依赖单向清晰——上层只调用下层接口绝不反向渗透。这种设计保证了可移植性SoC 厂商只需实现最底层的硬件适配上层策略、用户接口、驱动集成全部复用内核标准代码。2.1 第一层硬件感知层Hardware Sensing Layer这是整个框架的地基负责把物理世界的温度信号数字化。它不直接操作传感器而是通过struct thermal_zone_device_ops定义统一回调接口struct thermal_zone_device_ops { int (*get_temp)(struct thermal_zone_device *, int *); int (*set_trips)(struct thermal_zone_device *, int, int); int (*notify)(struct thermal_zone_device *, int); };get_temp()必须实现返回当前温度单位毫摄氏度。注意这里返回的是原始值不做任何滤波或校准——校准由上层策略或用户空间完成。set_trips()可选用于动态设置 trip point 阈值。很多 SoC 的 thermal sensor 支持硬件比较器触发中断后自动上报此时此函数可为空。notify()可选当硬件中断触发 trip 事件时调用用于快速响应如立即关断某路电源。我实测过全志 H6 的sun8i_ths驱动它的get_temp()实际做了三件事读取 ADC 原始值 → 查表转换为温度 → 应用芯片厂提供的二阶补偿公式。而 Intel 的x86_pkg_temp_thermal驱动则直接读 MSR 寄存器省去 ADC 步骤。但对外暴露的get_temp()接口完全一致上层无需关心差异。提示硬件层最大的坑是温度单位。内核强制要求单位为millidegree Celsiusm°C即 25°C 必须返回 25000。曾有个项目因驱动返回 25误以为是 °C导致所有 trip point 判定失效——trip80000对应 80°C但驱动返回 25内核认为才 0.025°C永远不触发降温。2.2 第二层热区抽象层Thermal Zone Abstractionstruct thermal_zone_device是核心数据结构代表一个物理热域如 “cpu-thermal”、“gpu-thermal”、“battery-thermal”。它封装了温度采集源指向 hardware sensing layerTrip point 配置数组每个元素含 temperature、type、hysteresisCooling device 关联列表struct list_head cooling_devices更新策略update_modeTHERMAL_DEVICE_ENABLED或THERMAL_DEVICE_DISABLED延迟参数passive_delay,polling_delayTrip point 是关键设计。内核定义了 5 种类型THERMAL_TRIP_CRITICAL不可逆动作如强制关机kernel_power_off()THERMAL_TRIP_HOT主动降温起点启动 cooling deviceTHERMAL_TRIP_PASSIVE被动降温通常触发 cpufreq 降频THERMAL_TRIP_ACTIVE主动散热设备启动如风扇全速THERMAL_TRIP_HOT与PASSIVE类似但优先级更高部分平台用注意CRITICAL和HOT是硬性保护PASSIVE/ACTIVE是策略调控。我在调试 RK3399 板子时发现PASSIVEtrip 设置为 70°C但CRITICAL设为 125°C中间留出 55°C 的缓冲带——这 55°C 就是 thermal framework 发挥策略调度的空间。2.3 第三层冷却设备抽象层Cooling Device Abstractionstruct thermal_cooling_device代表一个可调控的功耗单元。它不关心自己是什么只提供两个核心能力get_max_state()返回最大可调级别如 cpufreq cooling device 返回num_online_cpus() * 10表示 0~10 级 per CPUset_cur_state()设置当前级别如state5表示将 CPU 频率降至 50%内核预置了多种标准 cooling devicecpufreq_cooling: 绑定到 cpufreq policy通过cpufreq_frequency_table映射 state 到频率cpu_cooling: 控制 CPU online/offline 状态state0 全开state1 关 1 核...power_allocator: 基于 PID 控制器动态分配功耗预算用于 GPU/NPUfan_cooling: 控制风扇 PWM 占空比关键点在于一个 cooling device 可被多个 thermal zone 共享。例如cpu-thermalzone 和soc-thermalzone 都能绑定到同一个cpufreq_coolingdevice。framework 会自动聚合请求——如果 zone A 要求 state3zone B 要求 state5则最终取 max(3,5)5。这避免了多 zone 竞争导致的过度降频。2.4 第四层策略引擎层Governor Layer这才是 thermal framework 的“大脑”。它决定当温度越过 trip point 后如何选择 cooling device、如何调整 state、何时再次采样。内核提供 4 种内置 governorGovernor触发条件调控逻辑适用场景step_wise温度持续高于 trip线性步进每次 1 state直到达标通用默认选项bang_bang温度越过 trip 瞬间全力降温直接设为 max_state低于 trip 后全关风扇控制响应快user_space用户空间写入cur_state完全由用户程序控制自定义策略如机器学习预测power_allocator温度偏差PID 控制output Kp*e Ki*∫e dt Kd*de/dt精确功耗分配GPU/NPU我强烈建议新手从step_wise入手。它的逻辑清晰thermal_zone_device_update()被调用时遍历所有 trip对每个PASSIVE/ACTIVEtrip计算当前温度与 trip 的差值delta temp - trip_temp然后根据delta和slope斜率单位 m°C/state计算目标 statetarget_state (delta slope/2) / slope。slope默认为 10000即 10°C/state意味着每升高 10°Cstate 1。这个设计让调控平滑避免抖动。2.5 第五层用户空间接口层Userspace Interface全部通过 sysfs 暴露路径为/sys/class/thermal/thermal_zoneX/。关键文件temp: 当前温度m°Ctype: zone 名称如 cpu-thermaltrip_point_[0-9]_temp: 各 trip 阈值trip_point_[0-9]_type: trip 类型policy: 当前 governormode: enabled/disabledemul_temp: 用于模拟测试写入任意值覆盖tempcooling device 接口在/sys/class/thermal/cooling_deviceY/cur_state: 当前 statemax_state: 最大 statetype: device 类型Processor, Fanpower/allocated: 分配的功耗仅power_allocator注意emul_temp是调试神器。不用烧板子直接echo 90000 emul_temp就能触发 90°C 逻辑验证整个链路是否通畅。我调试时必做三步1.cat temp确认读数正常2.echo 95000 emul_temp3.watch -n 1 cat cur_state观察 cooling device 是否响应。三步下来80% 的配置问题都能定位。3. 实操详解从零构建一个可用的 thermal zone以 ARM64 SoC 为例假设你拿到一块新 SoC文档里只有一行“Thermal sensor located at APB bus offset 0x1200, 12-bit ADC, 0.5°C/LSB”。下面是我实际在瑞芯微 RK3328 上搭建 thermal zone 的完整过程步骤可直接复用。3.1 第一步编写硬件驱动drivers/thermal/rockchip_thermal.c核心是实现thermal_zone_device_opsstatic int rk3328_thermal_get_temp(struct thermal_zone_device *tz, int *temp) { struct rk3328_thermal_data *data thermal_zone_device_priv(tz); u32 val; // 1. 使能 ADC 通道寄存器操作 regmap_write(data-grf, GRF_SOC_CON1, BIT(12)); // 2. 触发 ADC 转换 regmap_write(data-pmu, PMU_ADC_CTRL, 0x1); // 3. 等待完成轮询实际项目建议用中断 usleep_range(100, 200); regmap_read(data-pmu, PMU_ADC_DATA, val); // 4. 转换12-bit 值0.5°C/LSB基准 25°C // 公式T 25 (val - 0x800) * 0.5 *temp 25000 ((val 0xfff) - 0x800) * 500; return 0; } static const struct thermal_zone_device_ops rk3328_tz_ops { .get_temp rk3328_thermal_get_temp, };关键细节usleep_range(100,200)ADC 转换需要时间太短读到 0太长影响 polling 效率。实测 100μs 足够。温度公式必须精确。0x800是 2048对应 25°C 基准点500是 0.5°C 转为 m°C0.5 * 1000。regmap是内核推荐的寄存器访问方式比裸写ioremap更安全。3.2 第二步在 DTS 中定义 thermal zonetsadc { #thermal-sensor-cells 1; rockchip,grf grf; rockchip,pmu pmu; /* 定义 thermal zone */ cpu_thermal: cpu-thermal { thermal-sensors tsadc 0; /* 引用 tsadc 的 channel 0 */ polling-delay-passive 1000; /* 1s 被动轮询 */ polling-delay 2000; /* 2s 主动轮询 */ thermal-zone { /* Trip points */ temperature 60000; /* 60°C */ type cpu; cooling-min-level 0; cooling-max-level 15; /* 关联 cooling device */ cooling-maps { map0 { cooling-device cpu0_cooling 0 15; cooling-device cpu1_cooling 0 15; cooling-device cpu2_cooling 0 15; cooling-device cpu3_cooling 0 15; }; }; }; }; };重点解析thermal-sensors tsadc 0指定使用tsadc的第 0 个通道。tsadc是前面定义的 ADC controller node。polling-delay-passive和polling-delay前者用于PASSIVEtrip 触发后的高频轮询如降频后者是常态轮询间隔。数值单位是毫秒。cooling-maps声明哪些 cooling device 属于此 zone。cpu0_cooling 0 15表示绑定cpu0_coolingdevice其 state 范围是 0~15。3.3 第三步注册 cooling devicecpufreqcpufreq cooling device 由drivers/thermal/cpu_cooling.c提供只需在 cpufreq driver 中注册// 在 cpufreq driver init 函数中 struct thermal_cooling_device *cdev; cdev of_cpufreq_cooling_register(policy); if (IS_ERR(cdev)) { pr_err(Failed to register cpufreq cooling device\n); return PTR_ERR(cdev); }of_cpufreq_cooling_register()会自动解析 DTS 中cooling-maps的绑定关系并创建cpu0_cooling等 device。它内部做了三件事读取cpufreq_frequency_table确定 frequency levels 数量即max_state将state映射到frequency_table[state]注册set_cur_state()回调调用__cpufreq_driver_target()3.4 第四步验证与调试编译烧录后检查 sysfs# 查看 thermal zone 是否创建 ls /sys/class/thermal/ # 应看到 thermal_zone0cpu_thermal # 读取温度 cat /sys/class/thermal/thermal_zone0/temp # 输出类似 4235042.35°C # 查看 trip point cat /sys/class/thermal/thermal_zone0/trip_point_0_temp # 应为 60000 # 查看绑定的 cooling device ls /sys/class/thermal/thermal_zone0/cdev[0-9]* # 应看到 cdev0 - ../cooling_device0cpu0_cooling # 强制触发降温模拟高温 echo 70000 /sys/class/thermal/thermal_zone0/emul_temp # 观察 cooling device state 变化 watch -n 1 cat /sys/class/thermal/cooling_device0/cur_state # state 应从 0 逐步升到 3、5...如果cur_state不变按顺序排查dmesg | grep thermal看是否有thermal zone registered或failed to register错误cat /sys/class/thermal/thermal_zone0/mode确认是enabledcat /sys/class/thermal/thermal_zone0/policy确认是step_wisecat /sys/class/thermal/cooling_device0/max_state确认 cooling device 已注册4. 深度避坑指南那些文档里不会写的实战陷阱我在 7 个不同 SoC 平台上部署 thermal framework踩过足够多的坑总结出 5 个最致命、最隐蔽的问题每个都附带真实案例和解决方案。4.1 陷阱一Trip Point 的 hysteresis迟滞缺失导致抖动现象温度在 70°C 附近时cur_state在 0 和 5 之间疯狂跳变风扇“哒哒哒”响个不停。原因hysteresis参数未设置。内核默认hysteresis0意味着温度 ≥70°C 时触发PASSIVE一旦降到 69.999°C 就立刻退出没有缓冲带。修复在 DTS 中为 trip point 添加hysteresisthermal-zone { temperature 70000; type cpu; hysteresis 2000; /* 2°C 迟滞 */ ... };原理当温度 ≥70°C触发 trip只有温度 ≤68°C70-2时才清除 trip。这 2°C 的窗口就是防抖空间。实测hysteresis10001°C对大多数 SoC 足够2000更稳妥。4.2 陷阱二Cooling device state 映射错误引发“降频无效”现象cur_state显示为 10但cpupower frequency-info显示频率仍是最大值。原因cpufreq_cooling的 state 映射表与实际frequency_table不匹配。常见于自定义 cpufreq table 时table[i].frequency未按降序排列或table[i].driver_data未正确设置。诊断查看cpufreqtablecat /sys/devices/system/cpu/cpufreq/policy0/scaling_available_frequencies # 输出1200000 1000000 800000 600000 # 但你的 table 可能是600000 800000 1000000 1200000升序修复确保frequency_table严格降序且table[i].index istatic struct cpufreq_frequency_table rk3328_freq_table[] { { .frequency 1200000 }, // state 0 { .frequency 1000000 }, // state 1 { .frequency 800000 }, // state 2 { .frequency 600000 }, // state 3 { .frequency CPUFREQ_TABLE_END }, };4.3 陷阱三Passive delay 过长导致“热失控”现象stress-ng跑 30 秒后板子烫手关机dmesg显示CRITICALtrip 被触发。原因polling-delay-passive设置过大如 5000ms而 SoC 温升速率快1°C/s。温度从 65°C 升到 85°C 只需 20 秒但 thermal framework 每 5 秒才检查一次错过最佳调控时机。修复根据 SoC 热特性设置合理 delay低功耗 SoCAllwinner H3polling-delay-passive 500500ms高性能 SoCRK3399polling-delay-passive 200200msx86 笔记本100100ms实测法则delay_ms (trip_temp - current_temp) / (dT/dt)。例如当前 60°Ctrip 70°C温升率 2°C/s则delay 10°C / 2°C/s 5s设为 2s 更安全。4.4 陷阱四Multi-zone 竞争导致“过度降频”现象GPU 负载高时CPU 频率被拉到最低即使 CPU 自身温度很低。原因gpu-thermalzone 和cpu-thermalzone 共享同一个cpufreq_coolingdevice且gpu-thermal的 trip 更激进如 65°C它要求state10而cpu-thermal只需state2最终取max(2,10)10。修复为不同 zone 分配独立 cooling device或使用power_allocatorgovernor 实现功耗隔离gpu_thermal: gpu-thermal { thermal-sensors tsadc 1; ... cooling-maps { map0 { cooling-device gpu_cooling 0 15; /* 独立 device */ }; }; };4.5 陷阱五User space governor 的权限与同步问题现象thermald进程运行但cur_state不更新。原因thermald需要CAP_SYS_ADMIN权限写cur_state且必须先echo user_space /sys/class/thermal/thermal_zone0/policy切换 governor。验证# 检查权限 sudo getcap /usr/sbin/thermald # 应输出/usr/sbin/thermald cap_sys_adminep # 检查 governor cat /sys/class/thermal/thermal_zone0/policy # 若为 step_wise需先切换 echo user_space /sys/class/thermal/thermal_zone0/policy更可靠的做法在thermald配置文件/etc/thermald/thermal-conf.xml中明确指定 zoneconfiguration thermal-zones thermal-zone typecpu kernels kernel typepid/ /kernels cooling-devices cooling-device typeProcessor/ /cooling-devices /thermal-zone /thermal-zones /configuration5. 架构演进与未来方向从 thermal 到 power-aware schedulingthermal framework 的设计初衷是热保护但随着异构计算big.LITTLE、CPUGPUNPU普及它正悄然演变为功耗感知调度Power-Aware Scheduling的核心基础设施。这不是我的猜测而是内核社区正在发生的事实。5.1 EASEnergy Aware Scheduler的深度集成Android 12 的 EAS 调度器不再只看 CPU 负载而是实时查询 thermal zone 的passive_delay和当前cur_state动态调整util_avg平均利用率的权重。当cpu-thermalzone 的cur_state 5EAS 会认为“此 CPU cluster 功耗受限”主动将新任务调度到 cooler 的 cluster 上。这要求 thermal framework 提供低延迟、高精度的cur_state查询接口——thermal_zone_get_cur_state()的响应时间必须 100μs否则影响调度实时性。5.2 Power Allocator Governor 的工业级应用power_allocator不再是实验特性。在 NVIDIA Jetson Orin 上它被用于 GPU 功耗闭环控制thermal_zone读取 GPU junction temperature →power_allocator计算所需功耗 budget → 下发到nvidia,gpu-powercooling device → GPU driver 依据 budget 调整 shader clock 和 memory bandwidth。整个环路延迟 5ms比传统step_wise精确 10 倍。5.3 用户空间策略的爆发式增长user_spacegovernor 正催生新生态power-profiles-daemon根据电池模式Performance/Balanced/Power Saver动态修改 trip pointsthermald集成机器学习模型基于历史负载预测温度峰值提前降频custom thermal daemon游戏本厂商用它实现“性能模式”一键解锁全部功耗墙这意味着thermal framework 的未来不再是内核里的一个“子系统”而是连接硬件、内核、用户空间的功耗策略总线Power Policy Bus。你写的每一行 DTS每一个trip_point都在为这个总线注入策略基因。我个人在实际项目中的体会是不要把 thermal framework 当作“温度监控模块”来用而要把它当作“功耗策略执行器”。它的价值不在于多精准地读出 0.1°C而在于能否让你用最简洁的 DTS 描述表达出“当 GPU 温度超过 75°C 时将 CPU 频率限制在 1.2GHz同时提升风扇至 60% 占空比”这样复杂的跨域协同策略。这才是通用架构的真正力量——它把硬件细节封装起来把策略逻辑解放出来让功耗管理从“工程师手动调参”走向“系统自动决策”。