WAPR, wide-angle pose refinement
WAPR

Localization and detection examples定位与检测示例

下方 IC-BIN[3] 调用展示保存对照的接口,需要 IC-BIN 的 RGB-D、内参和双物体模板库;上一页的 LM-O[2] 输入不能复现这张图。随包可直接执行的 LM-O 教程分别运行 examples/04_6d_localization.py 或 examples/05_bop_6d_detection.py。

The IC-BIN[3] calls below show the interface for the saved comparison. They require IC-BIN RGB-D, camera intrinsics and its two-mesh detector bank; the LM-O[2] inputs from the preceding page cannot reproduce this picture. Run examples/04_6d_localization.py or examples/05_bop_6d_detection.py for the included, fully executable LM-O lessons.

Localization assumes the object categories and their counts are known before reading the image. examples/04_6d_localization.py uses eight LM-O categories, one instance each. The IC-BIN comparison below restricts the search to object 1 and retains 14 poses; object 2 is excluded even though its CAD is in the bank. The first image shows 2D proposals before the count limit, and the second shows the retained poses ranked by score_6d. This recorded figure uses a 0.35 detection floor; the current omitted-argument default is 0.1 and may produce a different proposal set.

定位任务在读图前已给定类别及每类实例数。examples/04_6d_localization.py 使用 LM-O 的八个类别,每类一个实例。下方 IC-BIN 对照只搜索物体 1 并保留 14 条位姿;物体 2 虽在 CAD 库中,却不参与这次定位。第一张图是数量约束前的 2D 候选,第二张图是按 score_6d 排序后保留的位姿。此记录图使用 0.35 的检测阈值;如今省略参数时默认值为 0.1,可能得到不同的候选集合。

Example 04 applies the recipe's category and count variables. Its default is LM-O; the independent IC-BIN illustration here uses its stated category/count choices. The illustration's scores are not console output of the default script.

示例 04 使用配置中的类别与数量变量,默认数据集为 LM-O。本页独立 IC-BIN 插图采用所述的类别与数量设置;插图分数不作为默认脚本的控制台输出。

In [5b]
poses, timing, detections = estimate_frame_many_categories_many_instances(
    estimator, detector, rgb, depth_m, K, pose_meshes,
    obj_ids=obj_ids,
    inst_count=inst_count,
)
print("det_ms", timing["total_ms"], "detections", len(detections), "instances", len(poses), flush=True)

Source: examples/04_6d_localization.py, lines 125–130代码来源:examples/04_6d_localization.py,第 125–130 行

Out [5b]
IC-BIN scene 1 image 26, 34 boxes of object 1 IC-BIN scene 1 image 26, 14 poses kept by score_6d

IC-BIN scene 1, image 26. The first picture is the 2D step: only object 1, 34 boxes, and the number is score_2d. The second picture is the return: 14 poses, and the number is score_6d. The lowest of the 14 is 0.848.IC-BIN 场景 1,第 26 帧。第一张是 2D 这一步:只有物体 1,34 个框,数字是 score_2d。第二张是返回的结果:14 条位姿,数字是 score_6d。这 14 条里最低的是 0.848。

6D pose detection6D 位姿检测

下方 IC-BIN 调用展示保存对照的接口,需要 IC-BIN 的 RGB-D、内参和双物体模板库;上一页的 LM-O 输入不能复现这张图。随包可直接执行的 LM-O 教程分别运行 examples/04_6d_localization.py 或 examples/05_bop_6d_detection.py。

The IC-BIN calls below show the interface for the saved comparison. They require IC-BIN RGB-D, camera intrinsics and its two-mesh detector bank; the LM-O inputs from the preceding page cannot reproduce this picture. Run examples/04_6d_localization.py or examples/05_bop_6d_detection.py for the included, fully executable LM-O lessons.

Detection does not assume an instance count. On the same IC-BIN frame, both CAD categories are searched and all accepted regions proceed to batched pose estimation. The figures below show the 2D candidates followed by their estimated 6D poses. This run uses a 0.35 detection floor to retain the weaker object-2 proposals; the current omitted-argument default is 0.1. A 0.50 floor would remove object 2 on this frame. The 2D gate and optional WBPS pose gate act at different stages. Here, score_6d_min=None retains all estimated poses, including incorrect and duplicate outputs; annotated poses and instance counts are not used to select them.

检测不预设实例数量。同一帧 IC-BIN 图像中,两种 CAD 类别均参与搜索,保留的区域进入批量计算位姿估计。下方依次展示 2D 候选及其对应的 6D 位姿。本次运行采用 0.35 的检测阈值,以保留较弱的物体 2 候选;省略参数时当前默认值为 0.1。若在此帧使用 0.50,物体 2 会被过滤。2D 检测阈值与可选的 WBPS 位姿阈值作用于不同阶段。这里设置 score_6d_min=None,保留所有估计位姿,包括误检与重复输出;筛选不使用标注位姿或实例数量。

Example 05 uses its selected CAD library without a count prior. Its default is LM-O scene 2, image 1. The independent IC-BIN illustration uses scene 1, image 26 and the settings stated here.

示例 05 使用选定 CAD 库,不提供数量先验;默认读取 LM-O 场景 2 第 1 帧。独立 IC-BIN 插图使用场景 1 第 26 帧及本页所述设置。

In [5c]
detector = WAPRDet2D(template, device=device, backend=det_backend)
meshes = {}
for obj_id in obj_ids:
    mesh, diameter_m = load_mesh_m(bop_path, dataset, obj_id)
    if obj_id in object_names:
        mesh.metadata["name"] = str(object_names[obj_id])
    meshes[obj_id] = (mesh, diameter_m)
estimator = WAPREstimator(device=device)
prepared = estimator.prepare_meshes([pair[0] for pair in meshes.values()])
mesh_setup_seconds = estimator.last_mesh_setup_seconds
pose_meshes = {obj_id: (mesh, pair[1]) for (obj_id, pair), mesh in zip(meshes.items(), prepared)}
rgb, depth_m, K = load_bop_rgbd(bop_path, dataset, scene_id, im_id)
print("rgb", tuple(rgb.shape), "depth_m", tuple(depth_m.shape), "K", tuple(K.shape), flush=True)
# This call discovers classes and instances before refining their 6D poses.
# 这次调用先发现类别与实例,再修正各实例的 6D 位姿。
poses, timing, detections = estimate_frame_many_categories_many_instances(estimator, detector, rgb, depth_m, K, pose_meshes)
print("POSE_TIMING", timing, "model_setup_seconds", estimator.model_setup_seconds,
      "initial_mesh_setup_seconds", mesh_setup_seconds, flush=True)

Source: examples/05_bop_6d_detection.py, lines 112–129代码来源:examples/05_bop_6d_detection.py,第 112–129 行

Out [5c]
IC-BIN scene 1 image 26, 2D detection candidates of object 1 and object 2

2D candidates on IC-BIN scene 1, image 26. Each row gives an object id, score_2d and bbox_xywh. Object 1 has 33 candidates and object 2 has 9. Blue and orange rectangles identify the two CAD categories; no per-frame category-presence or instance-count prior is supplied.IC-BIN 场景 1 第 26 帧的 2D 候选。每行给出物体编号、score_2d 与 bbox_xywh;物体 1 有 33 个候选,物体 2 有 9 个。蓝色与橙色矩形分别表示两个 CAD 类别,未提供本帧中物体是否出现或实例数量的先验。

IC-BIN scene 1 image 26, all 42 estimated 6D poses with projected object contours, 3D bounding boxes and score_6d labels

Saved 6D detection outputs: 42 poses estimated in batches, with 12 initial hypotheses per candidate region. Projected object mesh contours and 3D bounding boxes use blue for category 1 and orange for category 2; labels report score_6d, not score_2d. All 42 outputs are shown without a pose-score gate, instance-count cap or additional pose NMS.保存的 6D 检测输出:每个候选区域生成 12 条初始候选姿态,批量估计得到 42 条位姿。物体网格模型的投影轮廓与 3D 包围盒按类别着色,物体 1 为蓝色、物体 2 为橙色;标签显示 score_6d,与上图的 score_2d 区分。图中保留全部 42 条输出,未使用位姿分数阈值、实例数量截断或额外的位姿 NMS。

The displayed poses use OpenGL + TensorRT[1] FP16 on RTX 5090. After three complete warmup calls, ten CUDA-synchronized calls give medians of 276.9 ms for 2D detection, 595.7 ms for the pose stage and 899.2 ms for the complete detection-to-pose call, including output assembly. Loading, CAD/template preparation, warmup, drawing and final pose NMS are excluded; hot mesh preparation/packing/upload counters remain unchanged. Stage medians are not an additive decomposition of the complete-call median. See the saved prediction record.

图中位姿使用 RTX 5090、OpenGL + TensorRT[1] FP16 计算。三次完整预热后,十次 CUDA 同步调用的中位数为:2D 检测 276.9 毫秒,位姿阶段 595.7 毫秒,检测到位姿的连续调用合计 899.2 毫秒,包含结果整理。不含加载、CAD 与模板准备、预热、绘图及最终位姿 NMS;计时中网格准备、打包、上传计数均不增加。各阶段中位数不必相加等于完整调用中位数。详见保存的预测记录。

References and licenses参考文献与许可

  1. tensorrt-cu12 — Other/Proprietary License. Runtime dependency; dependencies and model/data assets retain their own licenses.运行依赖;依赖及模型/数据资源保留各自许可。 · GitHubGitHub · License/notice 1许可/声明 1 ↩ ↩
  2. LM-O data — CC-BY-SA-4.0. Data and model excerpt, not the loader code.数据与模型摘录,不是读取代码。 · Original source原始来源 · GitHub: BOP toolkitGitHub:BOP 工具集
    Brachmann et al. Learning 6D Object Pose Estimation Using 3D Object Coordinates. ECCV 2014. · Paper论文 ↩ ↩
  3. IC-BIN · Official BOP dataset pageBOP 官方数据页
    Doumanoglou et al. Recovering 6D Object Pose and Predicting Next-Best-View in the Crowd. CVPR 2016. · Paper论文 ↩ ↩

Model weights and dataset/task assets may have separate terms. WAPR's source license does not replace them.模型权重、数据集与任务资源可能有独立条款;WAPR 源码许可不替代这些许可。