WAPR, wide-angle pose refinement
WAPR

How the modules work together模块如何衔接

Each step passes its result to the next. The 2D front end is optional when regions are supplied. On one frame, rendering, WAPR, SAPR and WBPS batch all supplied instances, whether their categories are the same or different. WBPS keeps a separate hypothesis group for each object.

每个阶段把结果交给下一阶段。已提供区域时,可省略 2D 检测。同一帧上,无论同类还是不同类,渲染、WAPR、SAPR 和 WBPS 都跨实例批量计算;WBPS 保留每个物体独立的候选组。

  1. Region and object区域与物体Supply RGB-D, a metric mesh, and a mask or box. If instances are unknown, 2D proposals and DINOv2[1] CAD matching find candidate regions.提供 RGB-D、米制网格以及 mask 或框。实例未知时,由 2D 候选和 DINOv2[1] CAD 匹配寻找区域。
  2. Pose hypotheses位姿候选Depth seeds translation; view directions and in-plane rotations form a pose group. The default is 4 × 3 = 12 hypotheses for each object.深度给出起始平移;视线方向与面内旋转构成一组位姿。默认每个物体为 4 × 3 = 12 条候选。
  3. WAPR → SAPRWAPR → SAPRMasked or unmasked WAPR makes three wide-angle updates; SAPR makes two smaller-angle updates. Both correct pose, not object identity.带 mask 或不带 mask 的 WAPR 做三次广角更新,SAPR 再做两次小角更新。它们修正位姿,不判定物体类别。
  4. WBPS → 6D poseWBPS → 6D 位姿WBPS scores the corrected group and selects one object-to-camera pose. Its pose score is distinct from the earlier 2D matching score.WBPS 给修正后的一整组位姿评分,选出物体到相机的位姿。该位姿分数与前面的 2D 匹配分数不同。

For video, DINOv2 target-mask tracking uses cross-frame feature matching to update the target mask in later RGB-D frames; pose updates then repeat on sampled frames. A single-frame reconstruction can supply a metric mesh before this flow, while robot examples can consume its estimated pose afterward. The four inference checkpoints—masked WAPR, unmasked WAPR, SAPR, and WBPS—are described on the Weights page; all were trained on this project's SA6D data.

处理视频时,DINOv2 目标掩码跟踪通过跨帧特征匹配,在后续 RGB-D 帧中更新目标 mask;抽样更新帧再重复位姿计算。单帧重建可以在流程之前提供米制网格;机器人示例可以在流程之后使用估计位姿。推理使用带 mask WAPR、不带 mask WAPR、SAPR 和 WBPS 四份权重,具体输入与作用见权重页;它们均在本项目的 SA6D 数据上训练。

References and licenses参考文献与许可

  1. DINOv2 — Apache-2.0. Code and official DINOv2 weights; retain copyright and license.源码与官方 DINOv2 权重;保留版权和许可。 · GitHubGitHub · License/notice 1许可/声明 1
    Oquab et al. DINOv2: Learning Robust Visual Features without Supervision. arXiv:2304.07193, 2023. · Paper论文 ↩ ↩

Model weights and dataset/task assets may have separate terms. WAPR's source license does not replace them.模型权重、数据集与任务资源可能有独立条款;WAPR 源码许可不替代这些许可。