返回专辑
·Johan·5 分钟阅读

控制频率:RL 与底层 PD 分层

策略 10 Hz 出目标,底层 1 kHz PD 跟踪;delay、action hold 和 sim dt 不一致会让 sim policy 真机振荡。

控制频率:RL 与底层 PD 分层

1. Sim 零 delay,真机 50 ms,policy 学 phase lead 真机振荡

常见架构:RL 输出 target joint pos/vel,底层 PD/阻抗 1 kHz 跟踪;不要 RL 直接输出 motor current,除非非常清楚动力学与饱和。RL step 每 0.1 s,sim 内部 0.001 s 物理步——同一 action 须持 control_period / sim_dt 个 substeps,即 action repeat。上次 sim2real,sim 零 delay 训出的 policy 在真机 sense→act 约 50 ms 时高频振荡——policy 学 phase lead,deploy 变 lag 振荡。Timing 不对齐,换 PPO 也救不了;根因在 MDP 定义(action repeat、delay、PD gain),不在算法 hyperparam。

2. 数字对齐与 timing 表

sim dt=0.002(500 Hz),RL control_freq=10 Hz → 每个 action 重复 50 个 sim step:

python
for _ in range(k):
    obs, r, done, trunc, info = env.step(action)

真机只有 100 Hz 而 sim 按 500 Hz 训,transfer 前要么 sim 降频,要么 real 插值。README 画 timing 表三列:RL step period、PD period、sensor publish rate——onboarding 新人先看表再改代码,别先动 PPO。Control decimation 写进 env spec,不是 deploy 时才发现的隐含假设。

3. PD gain 匹配

PD gain 在 sim/real 差 10× 时,同样 position setpoint 跟踪误差不同,policy 会学 compensating 高频抖动;real 上 PD 不同则抖动变失控。先 match PD,再谈 RL——Low-level PD 跟踪误差大时,RL 学 compensating 抖动,这是 sim2real 振荡的常见根因之一。

Integration test:RL 10 Hz + mock PD 100 Hz 跑 1h,log tracking error RMS;换 50 Hz PD 看 error 是否超训练分布。超则 retrain 或 restrict real PD 带宽。

4. Delay 与执行器滞后

Sim 里加执行器一阶滞后再训,否则 policy 高频抖动在 real 被低通滤掉,表现像没响应。Misalign 1 ms 有时就够 oscillate。Documentation 画 timing diagram:camera exposure midpoint、RL step、PD tick 对齐——camera 曝光中点与 RL step 不对齐,视觉 obs 和 action 因果错位,policy 学 spurious correlation。

RL 输出是 reference trajectory 还是 torque residual 要写清;后者风险大,底层 PID 饱和时 RL 仍以为自己在控。Torque residual 需要 RL 与底层控制器职责边界非常清楚。

5. 常见失败

Sim dt 改大省算力,动力学变「软」,transfer 失败。Eval 频率和 train 不同——debug 时临时改 repeat 等于换 MDP。Action repeat 别隐式藏在 wrapper 里忘了 k 值——文档写清 k,code review 查 wrapper 链。

Sim 里 sense→act 零 delay,真机 50 ms——policy 学 phase lead。Sim 加 random delay 做 robustness test 是 pragmatic 做法,但最好从训练就对齐真实 delay。

6. Sim2real 频率对齐清单

记录:sim physics dt、control decimation、real controller rate、camera exposure 与 RL step 对齐关系。任一项不一致,同一 policy 输出不同物理效果。真机录 torque/position 看是否跟踪 RL setpoint——tracking error 分布应在 sim 训练时见过的范围内。

7. 案例:action repeat 不一致导致 eval 通过、真机抖

Train 时 k=50(10 Hz RL,500 Hz sim),deploy 时工程师改 wrapper 为 k=10「为了快」,policy 每 0.02 s 换 setpoint 而非 0.1 s——MDP 变了,behavior 振荡。修复:deploy 强制 k=50,与 train 一致。Eval 脚本也要 k=50——eval 时改 repeat 等于换 MDP。

8. 验收

  • Log action hold 次数 control_dt / sim_dt。
  • Eval 时 action repeat 与 train 一致。
  • PD gain sim/real 文档化;先 match PD 再训 RL。
  • Timing 表进 README;integration test tracking error RMS 在训练分布内。
  • Sim 加 delay 或执行器滞后后再训,若真机有已知 delay。

9. 与 sim2real 的交点

Control frequency 是 sim2real 最易忽略的 MDP 维度——domain randomization 补不了 delay 和 action repeat 不一致。Transfer 前 checklist:sim physics dt、control decimation k、real controller rate、camera-RL 对齐四项文档一致。Integration test:mock PD 不同 rate 下 tracking error RMS 是否超训练分布——超则 retrain 或 restrict real PD 带宽。Torque residual 输出风险大:PID 饱和时 RL 仍以为在控,spec 里写清 RL 输出是 reference 还是 residual。新人 onboarding 第一周只改 timing 表不改算法——确认 k、PD gain、delay 理解后再动 PPO,多数振荡是 timing 不是 clip。

10. 录包与 transfer checklist

Sim2real transfer 前四项文档一致:sim physics dt、control decimation k、real controller rate、camera-RL 对齐。录包振荡时三列同轴:RL setpoint、PD tracking error、torque——tracking error 大是 PD 不匹配,setpoint 高频抖是 k 或 delay 不对。Integration test:mock PD 不同 rate 下 tracking error RMS 是否超训练分布。Sim 零 delay、真机 50 ms 几乎必振荡——sim 加 delay 或执行器滞后再训。Action repeat 写进 env spec;eval 与 train 的 k 必须一致,debug 时改 k 等于换 MDP。frame_skip(Atari)与 control decimation(robotics)语义相同:一个 RL 决策持多步。PD gain sim/real 差 10× 时 policy 学 compensating 抖动——先 match PD 再训 RL,integration test 录 tracking error RMS 确认在训练分布内。Camera exposure 与 RL step 不对齐时视觉 obs 与 action 因果错位;timing diagram 进 README 是 sim2real onboarding 契约。Sim 加 random delay 或执行器滞后再训,零 delay sim 配 50 ms 真机几乎必振荡。Torque residual 输出要写清职责边界;RL 10 Hz 出 setpoint、PD 1 kHz 跟踪是常见分层,别 RL 直接出 current 除非非常清楚。

11. 收束

Control frequency 是 MDP 的一部分,timing 不对齐换算法无效。README timing 表、action repeat 显式实现、eval 与 train 的 k 一致——sim2real 振荡排查的三条契约。先 match PD 再训 RL;integration test 录 tracking error RMS,超训练分布则 retrain 或 restrict 真机 PD 带宽。Domain randomization 补不了 delay 与 action repeat 不一致——timing 是 MDP 契约,不是 deploy 补丁。Sim 加 delay 或执行器滞后再训,是 sim2real 振荡排查的标准手段之一。新人 onboarding 先看 timing 表再改算法,多数振荡根因在 timing 不在 clip。

← 全部文章

johan's blog