Overview概览
This guide helps you turn an RGB-D image and an object mesh into a 6D pose, then use the same components for detection, localization, and tracking. It explains what each module receives and returns, when to add 2D detection, what the reported scores and timings measure, and which example to run for your scene. Start with one object and its mask; the later examples add unknown instances, video, reconstruction, and robot tasks.
这份文档帮助你从一张 RGB-D 图像和物体网格得到 6D 位姿,再把同一组模块用于检测、定位和跟踪。这里会说明各模块接收什么、输出什么,何时需要加 2D 检测,分数与耗时各自衡量什么,以及你的场景该运行哪个示例。可以先从一件物体和它的 mask 入手,再逐步处理未知实例、视频、重建与机器人任务。
The difficult part is often the starting pose: a 2D box or mask locates an image region but does not determine its 3D rotation, while a small-angle correction can be insufficient when the initial orientation is far away. WAPR is designed to correct large initial rotation errors; SAPR handles the remaining smaller error, and WBPS chooses among the corrected hypotheses. This gives a practical model-based route for objects absent from pose-model training, provided a metric mesh is available at inference time. Occlusion, symmetry, a poor mask, or a wrong object match can still cause failure; the examples show those limits alongside successful cases.
难点往往是起始位姿:2D 框或 mask 能指出图像区域,却不能确定物体的三维旋转;若初始朝向偏差很大,只做小角度修正也可能不够。WAPR 用于修正较大的初始旋转误差,SAPR 处理剩余的小误差,WBPS 从修正后的候选中选出位姿。只要推理时有米制网格,这条基于模型的流程就能用于位姿模型训练时未见过的物体。遮挡、对称、错误的 mask 或物体匹配仍可能导致失败;后面的示例会同时展示效果与边界。