WAPR, wide-angle pose refinement
WAPR

2D to 6D integrated pipeline2D 到 6D 一体化流程

All stages below use LM-O[3] scene 2, image 307: the rgb, depth_m, K and meter-scale meshes loaded in the preceding sections. Detection and segmentation supply candidate regions; pose estimation recovers the object-to-camera rigid transformation for each region. The WAPR detection route does not use supplied instance counts; the published-candidate comparison below sets explicit per-category caps. Later sections introduce localization with known identities and counts, detection without those priors, and tracking over successive frames, following the distinction between BOP localization and detection tasks.

以下各阶段共用上一节加载的 LM-O[3] 场景 2 第 307 帧,包括 rgb、depth_m、K 和米制 meshes。2D 检测与分割提供候选区域,位姿估计为每个区域求解物体系到相机系的刚体变换。WAPR 检测输入不提供实例数量;下方公布候选的对照单独设置每类数量限制。后续分别介绍已知物体类别与数量的定位、不提供这些先验的检测,以及连续帧中的跟踪;定位与检测的区分参照 BOP 任务定义。

  1. Prepare candidates准备候选RGB detector or published detections → regionsRGB 检测器或已公布检测 → 候选区域obj_id · mask · score_2d
  2. Estimate and score poses位姿估计与评分Regions + depth + K + object meshes候选区域 + 深度 + K + 物体网格模型pose_4x4 · score_6d
  3. Select outputs筛选输出WBPS ordering + GPU silhouette NMSWBPS 排序 + GPU 轮廓 NMSRetained poses保留的位姿

1. Prepare 2D candidates1. 准备 2D 候选

Choose one of the two input sources below. Both produce pose_inputs for the shared 6D stage; they use the same RGB-D frame, camera intrinsics and CAD models.

从下面两种来源中选择一种准备候选。两种方式都生成 pose_inputs,随后进入同一个 6D 步骤,共用同一帧 RGB-D、相机内参和 CAD 模型。

A. WAPR detection and segmentationA. WAPR 检测与分割

The following code uses WAPRDet2D for detection and segmentation: GroundingDINO[1] proposes boxes, SAM segments them, and DINOv2[2] matches the regions to the eight CAD categories. onboard_meshes builds the template bank in GPU memory; set template_path only to save it for a later run. Each returned row contains obj_id, score_2d, a pixel bbox, and a visible COCO RLE mask. This recorded frame uses a 0.35 detection threshold, then selects ten candidates at or above 0.50 for pose estimation. The RTX 5090 detection median is 282.2 ms over 30 calls after 3 warmups; loading, template construction, and drawing are excluded. The figure shows all 48 boxes: 29 blue boxes at or above 0.40 and 19 red boxes below 0.40. See the detector guide for the pipeline and CAD templates and feature encoding.

下面的代码用 WAPRDet2D 完成检测和分割:GroundingDINO[1] 提出包围盒,SAM 分割,DINOv2[2] 将区域与八个 CAD 类别匹配。onboard_meshes 在显存中建立模板库;只有需要保存供后续运行使用时,才设置 template_path。每个返回项含 obj_id、score_2d、像素包围盒 bbox 和可见区域的 COCO RLE mask。这一帧以 0.35 阈值检测,再选择分数不低于 0.50 的十个候选估计位姿。RTX 5090 预热 3 次后测量 30 次,2D 中位数为 282.2 毫秒,不计加载、模板构建和绘图。图中展示全部 48 个框:29 个蓝框不低于 0.40,19 个红框低于 0.40。整体流程见检测器指南,模板生成与编码见CAD 模板与特征编码。

The standalone detection script defaults to image 1 of LM-O scene 2. This page uses image 307 from the same sample pack. Set image_path to test/000002/rgb/000307.png under that root and confidence = 0.35, then filter the returned rows at 0.50 as shown below. Detecting directly at 0.50 uses a different timing protocol.

独立检测脚本默认使用 LM-O 场景 2 的第 1 帧,本页使用同一数据包的第 307 帧。将 image_path 设为数据根目录下的 test/000002/rgb/000307.png,并设置 confidence = 0.35,再按下方代码以 0.50 过滤返回项。直接以 0.50 检测的计时口径不同。

The code fragment is from example 01, whose default RGB is LM-O scene 2, image 1. The independent image below uses image 307, with the thresholds described above; it is not example 01 stdout.

代码片段摘自示例 01,该脚本默认读取 LM-O 场景 2 第 1 帧。下方独立插图使用第 307 帧及上述阈值,不作为示例 01 的标准输出。

In [3]
with Image.open(image_path) as frame:
    rgb = np.asarray(frame.convert('RGB'))
# Detection returns all accepted boxes and masks; no depth or pose is used here.
# 检测返回全部保留的框与掩码;此处不使用深度,也不估计位姿。
detector = WAPRDet2D(template, device=device, backend=backend, dino=dino, grounding=grounding)
torch.cuda.synchronize(device)
setup_seconds = time.perf_counter() - setup_started
# Warm the same image and confidence; reported inference excludes loading.
# 在相同图像与置信度下预热,报告的推理时间不包含模型加载。
warm_started = time.perf_counter()
for _ in range(3):
    detector.detect_many_categories_many_instances(rgb, confidence=confidence)
torch.cuda.synchronize(device)
warmup_seconds = time.perf_counter() - warm_started
instances, timing = detector.detect_many_categories_many_instances(rgb, profile=True, confidence=confidence)

Source: examples/01_one_rgb_detect_segment.py, lines 127–141代码来源:examples/01_one_rgb_detect_segment.py,第 127–141 行

Out [3]
WAPRDet2D boxes on LM-O scene 2 image 307, blue at or above 0.40 and red below; RTX 5090 median 0.282 seconds

Same image, scene 2 image 307. The picture draws boxes at or above 0.35; each number is the detector score_2d. Blue denotes scores at or above 0.40, and red denotes scores below 0.40. The ten lines above pass the 0.50 pose-input filter after detection at 0.35. Returned rows also have masks. The RTX 5090 detection median is 282.2 ms over 30 calls after 3 warmups.同为场景 2 第 307 帧。图中画出分数不低于 0.35 的框;每个数字是检测器的 score_2d。蓝色框的分数不低于 0.40,红色框低于 0.40。上面十行是在 0.35 检测之后通过 0.50 位姿输入阈值的候选。返回项还包含 mask。RTX 5090 预热 3 次后测量 30 次,检测中位数为 282.2 毫秒。

B. Read published 2D candidatesB. 读取已公布的 2D 候选

Alternatively, read the published CNOS[5]-FastSAM[4] detections (method 370) for the same frame. The code below reads the included excerpt samples/bop/lmo/cnos-fastsam_scene2_im307.json. Of its 58 rows, this comparison recipe retains the highest-scoring candidate for each of the eight target categories with score_thr = 0.0 and max_per_class. These caps belong to this comparison recipe. The WAPR input above uses a score threshold without per-category counts. For the full published file, set use_full_published_json = True in examples/03_published_boxes_to_pose.py; it downloads the file into outputs/cache/bop_det/.

另一种方式是读取同一帧已公布的 CNOS[5]-FastSAM[4] 检测(方法 370)。下方代码使用随包提供的摘录 samples/bop/lmo/cnos-fastsam_scene2_im307.json,从 58 条候选中按 score_thr = 0.0 和 max_per_class 为八个目标类别各保留分数最高的一条。每类数量限制是本次对照的设置;上面的 WAPR 输入按分数阈值筛选,不指定每类数量。需要完整公布文件时,在 examples/03_published_boxes_to_pose.py 中设置 use_full_published_json = True,脚本会下载到 outputs/cache/bop_det/。

In [4]
for scene_id, im_id in frames:
    mesh_setup_seconds = 0.0
    rgb, depth_m, K = load_bop_rgbd(bop_path, dataset, scene_id, im_id)
    print(
        "frame", int(scene_id), int(im_id),
        "rgb", tuple(rgb.shape),
        "depth_m", tuple(depth_m.shape),
        "K", tuple(K.shape),
        flush=True,
    )
    records = [
        record for record in detections.get((int(scene_id), int(im_id)), [])
        if int(record["obj_id"]) in wanted
    ]
    kept = keep_top_per_class(records, max_per_class)
    print("kept", len(kept), "of", len(records), "obj_ids", sorted(wanted), flush=True)

Source: examples/03_published_boxes_to_pose.py, lines 142–157代码来源:examples/03_published_boxes_to_pose.py,第 142–157 行

Out [4]
frame 2 307 rgb (480, 640, 3) depth_m (480, 640) K (3, 3)
kept 8 of 58 obj_ids [1, 5, 6, 8, 9, 10, 11, 12]

Captured stdout excerpt · 已保存标准输出摘录 ·

CNOS-FastSAM boxes and masks on LM-O scene 2 image 307

The published CNOS-FastSAM file on the same image, shown beside the WAPRDet2D boxes above. Rectangles are 2D boxes; filled regions are masks; labels give detection scores. One row is retained for each of the eight target categories.同一帧的已公布 CNOS-FastSAM 检测,可与上方的 WAPRDet2D 框对照。矩形是 2D 框,填色区域是 mask,标签给出检测分数。八个目标类别各保留一条检测。

Keeping one row per category does not guarantee eight correct regions. These are predicted candidates; the shared 6D stage below estimates and scores their poses.每类保留一条并不意味着八个区域都正确。这些仍是预测候选,接下来由统一的 6D 步骤估计位姿并评分。

2. 6D pose estimation, scoring and selection2. 6D 位姿估计、评分与筛选

Whichever source you choose, pass its pose_inputs to one estimate_many_categories_many_instances call with RGB, depth, camera intrinsics and CAD models. Decode any RLE masks first; for a candidate with only a box, fill that box as its region mask. Rendering, WAPR/SAPR refinement and WBPS scoring batch all hypotheses, with an independent 12-hypothesis attention group per candidate. Finally, rank by score_6d and suppress overlapping rendered silhouettes within each CAD category. Annotated masks and poses do not enter inference or selection.

无论选用哪一种来源,都将其 pose_inputs 与 RGB、深度、相机内参和 CAD 模型一次交给 estimate_many_categories_many_instances。先解码 RLE mask;只有包围盒的候选则以填充包围盒作为区域 mask。渲染、WAPR/SAPR 位姿修正和 WBPS 评分批量处理全部候选姿态,每个候选的 12 条候选姿态保持独立注意力组。最后按 score_6d 排序,在同一 CAD 类别内进行渲染轮廓重叠抑制。标注 mask 和位姿不进入推理或筛选。

The independent figure and timing record below use input A: ten WAPR candidates produce 120 hypotheses; silhouette suppression removes one duplicate can, leaving nine poses. The RTX 5090 record reports a 6D median of 0.134 s and a continuously timed detector + pose median of 0.419 s over 30 calls after three warmups. Input B instead passes eight CNOS-FastSAM candidates to the same code; its outputs and timings differ. The green reference is added only after prediction and selection.

下方独立可视化与计时记录对应输入 A:十个 WAPR 候选产生 120 条候选姿态,轮廓抑制去掉一个重复 can,保留九个位姿。RTX 5090 预热三次、测量三十次,记录的 6D 中位数为 0.134 秒,检测与位姿连续计时中位数为 0.419 秒。输入 B 则将八个 CNOS-FastSAM 候选送入同一段代码,输出和耗时不同。绿色对照只在预测和筛选完成后加入。

The code and Out excerpt below use example 03 and its eight published CNOS-FastSAM masks (input B). The independent figure uses input A as described above.

下方代码与 Out 摘录对应示例 03 及其八块公布的 CNOS-FastSAM 掩码,即输入 B;独立插图仍采用上文所述的输入 A。

In [5]
pose_instances = instances
shape = (tuple(depth_m.shape), tuple(estimator.renderer.cached_mesh_id(item["mesh"]) for item in pose_instances))
warmup_seconds = 0.0
if pose_instances and shape not in warmed_shapes:
    warmup_seconds = estimator.warmup_pose(rgb, depth_m, K, pose_instances)
    warmed_shapes.add(shape)
# Synchronize the complete batch; exclude mesh setup, warmup and drawing.
# 同步完整批量调用;网格准备、预热和绘图均在预热后的计时外。
torch.cuda.synchronize(estimator.device)
pose_started = time.perf_counter()
poses = estimator.estimate_many_categories_many_instances(rgb, depth_m, K, pose_instances)
torch.cuda.synchronize(estimator.device)
pose_hot_seconds = time.perf_counter() - pose_started
print("POSE_TIMING", {"model_setup_seconds": estimator.model_setup_seconds,
      "mesh_setup_seconds": mesh_setup_seconds, "warmup_seconds": warmup_seconds,
      "pose_hot_seconds": pose_hot_seconds, "instances": len(instances)}, flush=True)
rows = []
for record, pose in zip(instances, poses):
    print(
        {
            "scene_id": int(scene_id),
            "im_id": int(im_id),
            "obj_id": int(record["obj_id"]),
            "score_2d": float(record["score_2d"]),
            "score_6d": float(pose["score_6d"]),
            "t_m": np.asarray(pose["t_m"]).reshape(3).tolist(),
        },
        flush=True,
    )

Source: examples/03_published_boxes_to_pose.py, lines 183–211代码来源:examples/03_published_boxes_to_pose.py,第 183–211 行

Out [5]
{'scene_id': 2, 'im_id': 307, 'obj_id': 1, 'score_2d': 0.5453650951385498, 'score_6d': 0.9718749523162842, 't_m': [0.08519824594259262, 0.05933931842446327, 0.7537625432014465]}
{'scene_id': 2, 'im_id': 307, 'obj_id': 5, 'score_2d': 0.7412353754043579, 'score_6d': 0.9090625047683716, 't_m': [0.03579800948500633, -0.13113363087177277, 0.763190507888794]}

Captured stdout excerpt · 已保存标准输出摘录 ·

Saved RTX 5090 batched 6D poses from predicted WAPRDet2D masks on LM-O scene 2 image 307; median 0.134 seconds

Red is the prediction; green is the offline pose reference. Labels give score_6d; t_m is translation in meters. The two extra driller scores are 0.249 and 0.251. The footer reports the pose batch median, excluding detection and drawing.红色是预测,绿色是事后位姿对照。标签为 score_6d,t_m 是米制平移。额外两条 driller 的分数为 0.249 和 0.251。页脚标出位姿批量计算中位数,不含检测与绘图。

References and licenses参考文献与许可

  1. GroundingDINO — Apache-2.0. Vendored detector source; original notices retained.随包检测源码;保留原始声明。 · GitHubGitHub · License/notice 1许可/声明 1 · License/notice 2许可/声明 2
    Liu et al. Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection. ECCV 2024. · Paper论文 ↩ ↩
  2. DINOv2 — Apache-2.0. Code and official DINOv2 weights; retain copyright and license.源码与官方 DINOv2 权重;保留版权和许可。 · GitHubGitHub · License/notice 1许可/声明 1
    Oquab et al. DINOv2: Learning Robust Visual Features without Supervision. arXiv:2304.07193, 2023. · Paper论文 ↩ ↩
  3. LM-O data — CC-BY-SA-4.0. Data and model excerpt, not the loader code.数据与模型摘录,不是读取代码。 · Original source原始来源 · GitHub: BOP toolkitGitHub:BOP 工具集
    Brachmann et al. Learning 6D Object Pose Estimation Using 3D Object Coordinates. ECCV 2014. · Paper论文 ↩ ↩
  4. FastSAM · GitHubGitHub
    Zhao et al. Fast Segment Anything. arXiv:2306.12156, 2023. · Paper论文 ↩ ↩
  5. CNOS · GitHubGitHub
    Nguyen et al. CNOS: A Strong Baseline for CAD-based Novel Object Segmentation. ICCV Workshops 2023. · Paper论文 ↩ ↩

Model weights and dataset/task assets may have separate terms. WAPR's source license does not replace them.模型权重、数据集与任务资源可能有独立条款;WAPR 源码许可不替代这些许可。