本质矩阵与基础矩阵:别混层
F 在像素平面、E 在归一化相机坐标;有可靠 K 时应估 E 并在归一化坐标上做 RANSAC,混层或漏去畸变会导致高 inlier 的错误位姿。

1. 80% inlier 的位姿仍是错的
单目 VO 初始化,findEssentialMat RANSAC 报告 inlier 率 82%,recoverPose 也返回大量正深度点——但轨迹飘、scale 乱跳。查下来是用 uncorrected 像素点估 E,却把 threshold=1.0 当成像素阈值;归一化平面 1.0 与像素 1.0 差一个 的量级。RANSAC 在错误几何层上照样能「收敛」,silent bug 比 crash 更贵。复盘 bag 时若没有记录 threshold 单位与是否 undistort,几乎无法区分算法失败与契约失败。
我的记法:有标定尽量待在 E 层。用 F 当 E、或漏 undistort,是高 inlier 错误位姿的头号来源。
2. 两层几何
极线约束 , 是 、rank 2。内参 已知时,归一化坐标 ,
有 5 DOF,分解出 4 组 ,需 cheirality(两相机前方)筛选。无 K 或跨相机内参未知时才在像素层用 F。两层之间差的不只是矩阵名字,而是 RANSAC threshold 的 单位与数量级——混层等于用错误的 inlier 定义做几何。
3. OpenCV API 与 threshold 单位
E, mask = cv2.findEssentialMat(
pts1, pts2, K, method=cv2.RANSAC, prob=0.999, threshold=1.0)
_, R, t, mask_pose = cv2.recoverPose(E, pts1, pts2, K, mask=mask)
F, mask_f = cv2.findFundamentalMat(
pts1, pts2, cv2.FM_RANSAC, 3.0, 0.99)findEssentialMat 的 threshold 在 归一化平面(通常 0.5–2.0);findFundamentalMat 在 像素。搞反是最常见 silent bug。有 K 时绕 F 多此一举——除非 cross-camera 内参未知。RANSAC 内点判定常用 Sampson 距离;自己实现八点法时 Hartley normalization 是必做,否则条件数爆炸,inlier 统计也会骗你。
4. recoverPose 与 cheirality
recoverPose 返回满足 cheirality 的 inlier 数。4 解歧义靠三角化:选使 most points 在两相机 且 reproj 最小的那组。我会 persist matching inlier mask,同一 mask 喂 recoverPose 和后续 triangulatePoints,避免几何链前后用不同点集。
纯旋转(三脚架 pan)时 ,E 退化——检测 translation 范数,别硬初始化 VO。单目 E 定 但 scale 任意,需 PnP 或已知 baseline 物体定 scale。纯旋转 clip 在 VO 日志里常表现为「深度全噪但 inlier 不少」,应直接 abort 初始化。
5. 极线可视化 debug
lines2 = cv2.computeCorrespondEpilines(pts1.reshape(-1, 1, 2), 1, F)F 错时 epiline 不穿过对应点——肉眼看比盲调 RANSAC 迭代快。匹配质量仍是上限:几何只能消化 inlier,不能发明对应。ORB ratio test 太松时,先 tighten matching 再动 E threshold。我会抽 20 对 match 画 epiline overlay 进 debug 视频,给非几何同事也能看出「线没穿过点」。
6. VO 初始化 pipeline 契约
两帧匹配 → E RANSAC → recoverPose → 三角化少量点验正深度比例 → 才进 track/BA。任一步 inlier 高但 reproj/depth 异常,应 abort 初始化而不是带错 pose 起跑。launch yaml 写:e_threshold_norm、f_threshold_px、min_parallax_deg、min_triangulate_ratio;换分辨率只改 matching,不改 E threshold 数值含义。Composable 容器里多个 VO 节点共用 yaml 时,更要启动打印有效 threshold 摘要。
7. 平面场景与 E 的边界
平面 dominant、baseline 很小时 方向不稳——E 仍可用但 translation 噪声大。此时 homography 竞争有时更诚实;若仍走 E,monitor 和 triangulate depth 方差,不达标则等下一关键帧。室内滑轨平移少、旋转多时,应优先怀疑 motion 是否接近退化,而不是先换特征检测器。
8. 失败模式
| 情况 | 现象 |
|---|---|
| 匹配外点多 | inlier 低或位姿跳 |
| 平面 + 小 baseline | 不稳 |
| 未 undistort | Sampson threshold 无意义 |
| 纯旋转 | ,depth 炸 |
9. 验收
RANSAC 后 inlier ratio、Sampson error 分布;recoverPose inlier count。下游:三角化正深度比例、triangulatePoints reproj。VO 里画 epipolar line 对齐;threshold 与坐标层写进 yaml,启动 assert 输入分辨率与 K 一致。录 bag 对比「混层」与「E+undistort」同一 clip 的轨迹 drift,差一个数量级时优先修契约。
10. 案例:混层与正确层的轨迹差
同一 indoor clip,像素 threshold 1.0 估 E 的轨迹 30 s 内 drift 到墙里;归一化 threshold 1.0 + undistort 后 inlier 率略低但轨迹可跑满 bag。团队从此把 e_threshold_norm 与 coord_layer: normalized 写进 VO yaml 必填项,code review 见到 findFundamentalMat 有 K 可用时直接问「为什么不用 E」。这类回归靠契约字段防,不靠记忆。
10. 与 bag 复盘的衔接
VO 初始化失败 bag 应存:匹配点数、E/F threshold 单位、是否 undistort、Sampson 误差分位数、recoverPose 四解 cheirality 计数。没有这些字段,复盘只能猜「匹配差」——实际上常是混层。我会把几何层写进 metadata:coord_layer=normalized、matrix=E,与 K yaml 版本同 hash。
有 K 就别长期在像素层硬算 F。混层是高 inlier 错误位姿的温床。
相关
也可以看看
johan's blog