WAPR, wide-angle pose refinement
WAPR

ROBI detection experimentsROBI 检测实验

This experiment compares WAPR, AAE[6], PPF[8] + ICP and Line2D[7] on four ROBI[4] scene-4, view-0 images from the Ensenso camera. The Zigzag RGB-D input and CAD mesh are provided with examples/10_robi_zigzag.py; dataset sources are listed under ROBI downloads.

本实验在 ROBI[4] 场景 4、第 0 帧的四种零件图像上比较 WAPR、AAE[6]、PPF[8] + ICP 与 Line2D[7],输入来自 Ensenso 相机。examples/10_robi_zigzag.py 随包提供 Zigzag 的 RGB-D 输入与 CAD 网格;数据来源见 ROBI 下载。

Nearest-pose comparison最近预测位姿对比

For each annotation with visibility >0.1, select the saved prediction with the smallest symmetry-aligned ADD. Blue indicates ADD <0.1d; red indicates ADD ≥0.1d, where d is the object diameter. Select an object, then click an instance number beneath any panel to inspect its prediction, ADD error and available score.

对可见度 >0.1 的每个真值实例,选择对称对齐后 ADD 最小的已保存预测。蓝色表示 ADD <0.1d,红色表示 ADD ≥0.1d,d 为物体直径。选择物体后,点击任一图下的实例编号,即可查看对应预测、ADD 误差与可用评分。

Loading comparisons…正在加载对比图片…

Nearest saved candidates, colored by ADD agreement.按 ADD 达标情况着色的最近候选。

Counts report covered annotations. Reused predictions are drawn once, colored by their minimum ADD; each instance button uses its own ADD. Zigzag baselines contain only saved subsets, marked in the panels. Published baselines have no confidence scores; file order is not treated as confidence.

面板统计达标的真值覆盖数。复用的预测仅画一次,总览颜色按其最小 ADD 判定;每个实例按钮使用自身的 ADD。Zigzag 基线仅有保存子集,已在面板中标注。公开基线未提供置信度,文件顺序不作为置信度排序。

This display shows candidate coverage matched to reference poses, rather than detection precision or one-to-one recall. Saved poses are displayed without additional NMS or score filtering. WAPR uses 4×4 hypotheses here; the chrome comparison on the project page uses 4×3. Colors and instance buttons report ADD agreement.

这是按真值匹配的候选覆盖展示,不作为检测查准率或一对一召回率。保存位姿不追加 NMS 或分数筛选;这里采用 4×4 候选,宣传页的铬螺丝对照采用 4×3。颜色与实例按钮表示 ADD 达标情况。

ADD uses corresponding vertices from the merged CAD mesh. Zigzag and DIN have no symmetry alignment; gear retains twelve z-axis rotation equivalents and chrome retains continuous z-axis alignment. Full model silhouettes may extend behind occluders and are not visible segmentation boundaries.

ADD 使用合并 CAD 网格中的对应顶点。Zigzag 与 DIN 不作对称对齐;齿轮沿用绕 z 轴的十二个旋转副本,铬螺丝沿用连续 z 轴对齐。完整模型投影可能延伸到遮挡物后方,不等同于可见分割边界。

Chrome-screw pose comparison铬螺丝位姿对照

The following panels reuse the same saved result as the project page: ROBI scene 4, view 0, with 38 annotations passing the visibility condition. This separate WAPR 4×3 run combines GroundingDINO[1]/SAM 2[3] and FastSAM[5] regions, matches CAD features with DINOv2[2] and starts from twelve pose hypotheses per region. Baselines use the published ROBI predictions. The nearest-prediction rule and ADD symmetry treatment are the same as above; no score threshold is applied before reference matching.

下方直接展示宣传页使用的同一份保存结果:ROBI 场景 4、第 0 帧,有 38 个通过可见度条件的标注。该独立 WAPR 4×3 运行结合 GroundingDINO[1]/SAM 2[3] 与 FastSAM[5] 候选区域,通过 DINOv2[2] 匹配 CAD 特征,每个区域从十二个候选姿态开始。基线采用 ROBI 公布的预测,最近预测规则与 ADD 对称处理同上,匹配真值前不按分数阈值筛选。

Loading comparisons…正在加载对比图片…

ROBI pose candidates and their ADD agreement.ROBI 位姿候选及其 ADD 达标情况。

Covered annotations: WAPR 27/38, Line2D 12/38, AAE 5/38, PPF + ICP 2/38. Blue indicates ADD <0.1d, red ADD ≥0.1d; here 0.1d is 2.91 mm. Click an instance to inspect its own error. Coverage is a post-hoc candidate diagnostic, not detection precision or one-to-one recall.

达标覆盖数:WAPR 27/38、Line2D 12/38、AAE 5/38、PPF + ICP 2/38。蓝色为 ADD <0.1d,红色为 ADD ≥0.1d;本例 0.1d 为 2.91 mm。点击实例查看自身误差。覆盖数用于事后候选诊断,不代表检测查准率或一对一召回率。

Runtime reference耗时参考

Separate RTX 5090 measurements use the 4×3 pose recipe and GroundingDINO/SAM2 plus FastSAM proposals. Total medians include 2D detection, batched 6D estimation and GPU silhouette NMS, after three warmups and ten synchronized calls. Loading, onboarding and drawing are excluded. These timings belong to those full workloads, not the selected contours above. Detailed standard/mixed measurements remain in the timing record.

另一组 RTX 5090 计时采用 4×3 位姿估计配置与 GroundingDINO/SAM2、FastSAM 联合候选。舍弃三次预热后同步计时十次;总耗时中位数包含 2D 检测、6D 批量估计和 GPU 轮廓 NMS,不含加载、建库与绘图。这些耗时对应该次完整负载,不是上方所选轮廓的计时。标准/混合方案的详细记录保留在计时文件中。

Object物体WAPR total, sWAPR 总耗时,秒
Zigzag2.48
Gear齿轮2.28
DIN2.30
Chrome screw铬螺丝2.70

Baseline stage timings. On Xeon 8352V, the post-match measurements discard three full-path warmups and time five calls. Line2D NMS, cloud preparation and ICP for 516 candidates take 1.572 s; PPF scene-cloud preparation and ICP for 39 candidates take 0.221 s. Matching and model setup are excluded from these values. The matcher measurements give 393.2 ms for Line2D and 1991.0 ms for PPF, using the same candidate CSVs as the post-match inputs. Line2D NMS uses RTX 4080 SUPER. These measurements describe the tested implementations and stages. AAE’s 2.1 s is an estimate from paper stage times for ten serial candidates, rather than a measured total.

基线阶段耗时。匹配后阶段测量在 Xeon 8352V 上完整预热三次、计时五次:Line2D 对 516 条候选执行 NMS、点云准备与 ICP,中位数为 1.572 s;PPF 对 39 条候选执行场景点云准备与 ICP,中位数为 0.221 s。以上耗时不含匹配和模型准备。匹配阶段测量给出 Line2D 393.2 ms、PPF 1991.0 ms,候选 CSV 与后处理输入一致。Line2D 的 NMS 使用 RTX 4080 SUPER。这些耗时对应所测实现及阶段。AAE 的 2.1 s 是按论文阶段耗时、十候选串行计算得到的估算值,并非完整流程的实测值。

References and licenses参考文献与许可

  1. GroundingDINO — Apache-2.0. Vendored detector source; original notices retained.随包检测源码;保留原始声明。 · GitHubGitHub · License/notice 1许可/声明 1 · License/notice 2许可/声明 2
    Liu et al. Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection. ECCV 2024. · Paper论文 ↩ ↩
  2. DINOv2 — Apache-2.0. Code and official DINOv2 weights; retain copyright and license.源码与官方 DINOv2 权重;保留版权和许可。 · GitHubGitHub · License/notice 1许可/声明 1
    Oquab et al. DINOv2: Learning Robust Visual Features without Supervision. arXiv:2304.07193, 2023. · Paper论文 ↩ ↩
  3. SAM 2 — Apache-2.0; cctorch: BSD-3-Clause. Native source and official checkpoints; cctorch carries an additional BSD notice. Ultralytics-converted files require checking their distributor terms.原生源码与官方权重;cctorch 另附 BSD 声明。Ultralytics 转换文件还需核对分发方条款。 · GitHubGitHub · License/notice 1许可/声明 1 · License/notice 2许可/声明 2
    Ravi et al. SAM 2: Segment Anything in Images and Videos. arXiv:2408.00714, 2024. · Paper论文 ↩ ↩
  4. ROBI data — Not separately verified / 未单独核实. The saved public poses and dataset are credited to ROBI; an independent grant to redistribute them has not been verified.保存的公开位姿和数据均注明 ROBI 来源;其再分发授权尚未独立核实。 · GitHubGitHub
    Yang et al. ROBI: A Multi-View Dataset for Reflective Objects in Robotic Bin-Picking. IROS 2021. · Paper论文 ↩ ↩
  5. FastSAM · GitHubGitHub
    Zhao et al. Fast Segment Anything. arXiv:2306.12156, 2023. · Paper论文 ↩ ↩
  6. AAE · GitHubGitHub
    Sundermeyer et al. Implicit 3D Orientation Learning for 6D Object Detection from RGB Images. ECCV 2018. · Paper论文 ↩ ↩
  7. Line2D · GitHub: ROBI published predictionsGitHub:ROBI 公开预测
    Hinterstoisser et al. Model Based Training, Detection and Pose Estimation of Texture-Less 3D Objects in Heavily Cluttered Scenes. ACCV 2012. · Paper论文 ↩ ↩
  8. PPF · GitHub: ROBI published predictionsGitHub:ROBI 公开预测
    Drost et al. Model Globally, Match Locally: Efficient and Robust 3D Object Recognition. CVPR 2010. · Paper论文 ↩ ↩

Model weights and dataset/task assets may have separate terms. WAPR's source license does not replace them.模型权重、数据集与任务资源可能有独立条款;WAPR 源码许可不替代这些许可。