WAPR, wide-angle pose refinement
WAPR

Detector usage and evaluation检测器使用与评估

This page covers model resources, detection calls, output inspection, and reported latency and accuracy for WAPRDet2D.

本页介绍 WAPRDet2D 的模型准备、检测调用、结果检查,以及已有延迟与精度结果。

Model and runtime configuration模型与运行配置

WAPRDet2D uses dino and grounding to select the visual encoder and box detector checkpoints. The published configuration is vitl14 with swinb; SAM 2.1[3]-L and BERT[6] are fixed dependencies. An unsupported model name raises ValueError before any download.

WAPRDet2D 通过 dino 与 grounding 分别选择视觉编码器和框检测器的权重。已公布的配置为 vitl14 与 swinb,SAM 2.1[3]-L 和 BERT[6] 为固定依赖。不支持的模型名称会在下载前触发 ValueError。

Argument参数 Names可取的名字 File under assets/weights/det2d/assets/weights/det2d/ 里的文件
dino vits14, vitb14, vitl14 dinov2_vits14_pretrain.pth, dinov2_vitb14_pretrain.pth, dinov2_vitl14_pretrain.pth
grounding swinb, swint groundingdino_swinb_cogcoor.pth, groundingdino_swint_ogc.pth

ViT-S, ViT-B, and ViT-L use embedding widths 384, 768, and 1024. The template bank has to be built with the same dino. A bank whose width does not match raises ValueError and asks for a rebuild. ViT-g and the register models are not in this table. The patch size stays 14.

ViT-S、ViT-B、ViT-L 的特征宽度是 384、768、1024。模板库必须用同一个 dino 来建。宽度不一致时抛出 ValueError,并要求重建这份库。表里没有 ViT-g,也没有带 register 的模型。图块大小仍是 14。

Constructing WAPRDet2D runs this schedule once. It is not repeated on the next frame.

构造 WAPRDet2D 时按这个顺序做一次。下一帧不会再做。

  1. Check dino and grounding. An unknown name stops here.先核对 dino 和 grounding。名字不对就停在这里。
  2. Place the files in assets/weights/det2d/. A file that is already there is kept. A missing DINOv2[2] weight, GroundingDINO[1] weight, sam2.1_l.pt, or BERT file is downloaded from the official URL into that directory. The GroundingDINO JSON is copied from third_party/GroundingDINO/groundingdino/config/ when assets/weights/det2d/ does not have it. vitl14 with swinb is then checked against wapr/det2d_assets.json. A known hash that does not match raises, and the file on disk is not overwritten.文件放在 assets/weights/det2d/。已存在的文件直接使用;缺失的 DINOv2[2]、GroundingDINO[1]、sam2.1_l.pt 或 BERT 权重从官方地址下载。若权重目录缺少 GroundingDINO JSON,则从 third_party/GroundingDINO/groundingdino/config/ 复制。vitl14 配 swinb 还会对照 wapr/det2d_assets.json;已知哈希不一致时报错,不覆盖磁盘文件。
  3. backend="trt" builds the DINOv2 FP16 engine when this weight, this TensorRT[5], and this GPU do not already have a matching one. vitl14 writes dino_patches_fp16.engine. vits14 and vitb14 write dino_patches_fp16_vits14.engine and dino_patches_fp16_vitb14.engine. A matching engine is loaded and not built again. backend="torch" skips the engine and runs NativeDino. GroundingDINO and SAM stay PyTorch[4] FP32. They are not converted.backend="trt" 时,这份权重、这份 TensorRT[5]、这块 GPU 还没有匹配的引擎,就构建 DINOv2 的 FP16 引擎。vitl14 写出 dino_patches_fp16.engine。vits14 和 vitb14 写出 dino_patches_fp16_vits14.engine 和 dino_patches_fp16_vitb14.engine。已经匹配的引擎直接加载,不再构建。backend="torch" 不建引擎,走 NativeDino。GroundingDINO 和 SAM 保持 PyTorch[4] FP32,不做这步转换。
  4. Load the template bank, then the encoder and GroundingDINO. The query crop uses the TensorRT engine when backend is trt. The bank itself was encoded by NativeDino in PyTorch.加载模板库,再加载编码器和 GroundingDINO。查询裁剪在 backend 为 trt 时用 TensorRT 引擎。库本身是用 PyTorch 的 NativeDino 编码的。

Compatible source trees belong in third_party/. See Detector resources for source and checkpoint files, and Detection setup for the GroundingDINO attention adaptation.

兼容源码放在 third_party/。源码与权重清单见检测器资源;GroundingDINO 注意力适配见检测环境配置。

Stage阶段What it does做什么
Boxes包围盒 GroundingDINO[1], FP32. grounding="swinb" is Swin-B. swint is Swin-T. Text items ., bbox threshold 0.1, bbox NMS 0.7GroundingDINO[1],FP32。grounding="swinb" 是 Swin-B。swint 是 Swin-T。文本 items .,包围盒阈值 0.1,包围盒 NMS 0.7
Segment分割 SAM 2.1[3]-L, FP32, one call. Each GroundingDINO box is segmented, and the mask is clipped to that box. This mask is what DINOv2[2] encodes, and it is the mask returned with the box.SAM 2.1[3]-L,FP32,一次调用。每个 GroundingDINO 包围盒都做分割,mask 再裁在这个包围盒内。DINOv2[2] 编码的就是这块 mask,返回时它和包围盒在同一条里。
Match匹配 DINOv2 on a 224 px crop, against 42 views × 4 in-plane template rotations. dino="vitl14" is ViT-L/14. vits14 and vitb14 are ViT-S/14 and ViT-B/14. trt uses that choice’s FP16 engine.DINOv2 编码 224 像素裁剪,对照 42 个视角 × 4 次面内旋转的模板。dino="vitl14" 是 ViT-L/14。vits14 和 vitb14 是 ViT-S/14 和 ViT-B/14。trt 用这一选择自己的 FP16 引擎。

Run detection检测调用

Detection filtering proceeds in three stages in wapr/det2d.py: box NMS ranks GroundingDINO proposals by objectness at box_nms_threshold = 0.7; score filtering removes candidates below the configured gate (default score_threshold = 0.1); mask_nms_same_category then suppresses visible masks with IoU above 0.5 within each category, retaining higher scores first. Different categories are not compared, and sufficiently separated instances of the same category remain. obj_ids limits eligible categories before selecting each candidate's best identity. The score gate is exposed as confidence or score_min in the detector, and as score_2d_min in the integrated frame interface.

wapr/det2d.py 中的检测过滤分为三个阶段:框 NMS 按 GroundingDINO 的 objectness 排序,采用 box_nms_threshold = 0.7;分数过滤移除低于指定阈值的候选,默认 score_threshold = 0.1;随后 mask_nms_same_category 在同类别内抑制可见掩码 IoU 大于 0.5 的重复候选,优先保留较高分数。不同类别互不比较,同类别中重叠较小的实例仍可保留。obj_ids 在选择各候选的最佳身份之前限定可参与的类别。检测器以 confidence 或 score_min 接收分数阈值,一体化整帧接口则使用 score_2d_min。

Box suppression uses torchvision.ops.nms with CUDA inputs. Mask suppression uses wapr/suppression.py: mask intersections, stable ranking and greedy suppression remain on GPU, and only the final indices return to CPU for the output interface. The shared mask implementation rejects CPU tensors. Post-estimation suppression of rendered pose contours is documented under 6D pose detection.

框抑制采用 CUDA 输入的 torchvision.ops.nms。掩码抑制使用 wapr/suppression.py:掩码交集、稳定排序与贪心抑制均保留在 GPU,仅将最终下标传回 CPU 以生成输出;共享掩码实现拒绝 CPU 张量。估计完成后的渲染位姿轮廓抑制见6D 位姿检测。

detect_many_categories_many_instances takes one uint8 RGB image and returns boxes and visible masks for multiple categories and instances. obj_ids=None searches all bank categories; one object id restricts the category while allowing multiple instances. confidence retains candidates whose best allowed-category score meets the threshold. Omitted or None, it uses score_threshold = 0.1. score_min is an alias; conflicting values raise ValueError.

detect_many_categories_many_instances 接收一张 uint8 RGB 图像,返回多类别、多实例的包围盒和可见区域掩码。obj_ids=None 搜索模板库的全部类别;指定单个物体编号时限定类别,仍允许多个实例。confidence 保留允许类别中最高分达到阈值的候选;省略或传入 None 时使用 score_threshold = 0.1。score_min 是同义参数,两者取值冲突时抛出 ValueError。

onboard_meshes takes {obj_id: mesh} with vertices in meters. Use the same dino for the template bank and detector. An empty cache_path keeps templates on the GPU; set a file path for reuse across runs. Example 01 runs this 2D call without a pose-estimation stage.

onboard_meshes 接收顶点以米计的 {obj_id: mesh}。模板库与检测器应使用同一个 dino。cache_path 留空时模板保存在显存,设置文件路径后可跨进程复用。示例 01 仅执行该 2D 检测调用。

from wapr.det2d import WAPRDet2D, onboard_meshes

dino = "vitl14"
grounding = "swinb"
confidence = 0.1

bank = onboard_meshes(meshes_m, cache_path="", device="cuda:0", dino=dino)
detector = WAPRDet2D(
    bank, device="cuda:0", backend="trt",
    dino=dino, grounding=grounding,
)
instances, timing = detector.detect_many_categories_many_instances(rgb, obj_ids=None, confidence=confidence)

Each instance has obj_id, score_2d, bbox as xywh pixels, and mask as a visible COCO RLE. score_2d is the detector's ranking score. On the pose path, estimate_frame_many_categories_many_instances retains this field and applies score_2d_min as its gate. examples/05_bop_6d_detection.py leaves template_path empty and then calls estimate_frame_many_categories_many_instances.

每条结果包含 obj_id、score_2d、像素 xywh 格式的 bbox,以及可见区域的 COCO RLE mask。score_2d 是检测器的排序分数。位姿流程保留这个字段,并以 score_2d_min 作为检测阈值。examples/05_bop_6d_detection.py 将 template_path 留空,再调用 estimate_frame_many_categories_many_instances。

Inspect and save results结果检查与保存

visualize_2d_detection in wapr/view.py draws the detection bbox in blue and mask with a tint per instance. Labels use names[obj_id] when supplied, otherwise obj N, followed by score_2d. Example 01 saves all detections to JSON and renders only the highest preview_top_k scores; this display limit does not change inference outputs.

wapr/view.py 中的 visualize_2d_detection 绘制蓝色检测框 bbox 和按实例着色的 mask。标签使用传入的 names[obj_id],未提供名称时使用 obj N,随后显示 score_2d。示例 01 将完整检测结果保存为 JSON,仅绘制分数最高的 preview_top_k 条;显示数量不影响推理输出。

from wapr.view import visualize_2d_detection

image_out = visualize_2d_detection(
    rgb, instances,
    image="results/det2d/single_image.jpg",
    show=False,
    names=names,
)

image accepts a destination path, creating parent directories as needed, or True to use recipe.visualize_path (default outputs/vis) and filename (default det2d.jpg). The example uses show=False for a headless server. Set it to True for a local window that closes on a keypress; this requires DISPLAY. The call is independent of recipe.visualize.

image 为文件路径时写入该 JPEG 或 PNG,并创建所需目录;设为 True 时使用 recipe.visualize_path(默认 outputs/vis)与 filename(默认 det2d.jpg)。上例的 show=False 适用于无显示服务器。需要本地窗口时可设为 True,按键关闭窗口;无 DISPLAY 时会报错。该调用独立于 recipe.visualize。

For subsequent pose results, visualize_6d_pose uses pose_4x4, mesh and K to draw red 3D boxes and rendered contours. Optional gt_pose_4x4 is green for comparison. recipe.visualize = True enables pose-image output without opening a window; the 2D visualization call is independent of this switch.

后续位姿结果由 visualize_6d_pose 根据 pose_4x4、mesh 和 K 绘制红色 3D 包围盒与渲染轮廓;可选的 gt_pose_4x4 以绿色显示,用于对照。recipe.visualize = True 启用位姿图像输出,不打开窗口;2D 可视化调用独立于该开关。

Reported latency and accuracy延迟与精度结果

On one RTX 5090 image, the warmed-up GroundingDINO + SAM2 + DINOv2 CAD-matching front end measured 220.7 ms median. This is a single-image 2D measurement; it excludes 6D pose estimation and does not represent dataset-wide latency.

在 RTX 5090 上,GroundingDINO + SAM2 + DINOv2 CAD 匹配前端对一张图像预热后测得的耗时中位数为 220.7 ms。这是单图 2D 测量,不包含 6D 位姿估计,也不代表数据集级延迟。

Existing full-test BOP2D evaluations report the bbox / mask AP below (%). The evaluation uses a score threshold of 0.35; the detector default is 0.1. Full TensorRT AP has been measured for LM-O[7], IC-BIN[10] and TUD-L[9] only.

已有完整测试集的 BOP2D 本地评测结果如下,数值为包围盒 / 掩码 AP(%)。评测采用分数阈值 0.35,检测器默认阈值为 0.1。TensorRT 完整 AP 仅测量了 LM-O[7]、IC-BIN[10] 和 TUD-L[9]。

Dataset数据集DINOv2[2] PyTorch[4] bbox / mask APDINOv2[2] PyTorch[4] 包围盒 / 掩码 APDINOv2 TensorRT[5] bbox / mask APDINOv2 TensorRT[5] 包围盒 / 掩码 AP
LM-O[7]50.704 / 47.67950.706 / 47.692
IC-BIN[10]31.364 / 39.63331.419 / 39.658
TUD-L[9]62.009 / 59.07061.971 / 59.020
YCB-V[11]70.168 / 69.996Not measured未测量
T-LESS[8]47.543 / 45.919Not measured未测量

References and licenses参考文献与许可

  1. GroundingDINO — Apache-2.0. Vendored detector source; original notices retained.随包检测源码;保留原始声明。 · GitHubGitHub · License/notice 1许可/声明 1 · License/notice 2许可/声明 2
    Liu et al. Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection. ECCV 2024. · Paper论文 ↩ ↩ ↩ ↩
  2. DINOv2 — Apache-2.0. Code and official DINOv2 weights; retain copyright and license.源码与官方 DINOv2 权重;保留版权和许可。 · GitHubGitHub · License/notice 1许可/声明 1
    Oquab et al. DINOv2: Learning Robust Visual Features without Supervision. arXiv:2304.07193, 2023. · Paper论文 ↩ ↩ ↩ ↩ ↩ ↩
  3. SAM 2 — Apache-2.0; cctorch: BSD-3-Clause. Native source and official checkpoints; cctorch carries an additional BSD notice. Ultralytics-converted files require checking their distributor terms.原生源码与官方权重;cctorch 另附 BSD 声明。Ultralytics 转换文件还需核对分发方条款。 · GitHubGitHub · License/notice 1许可/声明 1 · License/notice 2许可/声明 2
    Ravi et al. SAM 2: Segment Anything in Images and Videos. arXiv:2408.00714, 2024. · Paper论文 ↩ ↩ ↩ ↩
  4. torch — BSD License. Runtime dependency; dependencies and model/data assets retain their own licenses.运行依赖;依赖及模型/数据资源保留各自许可。 · Original source原始来源 · License/notice 1许可/声明 1 · License/notice 2许可/声明 2 ↩ ↩ ↩ ↩
  5. tensorrt-cu12 — Other/Proprietary License. Runtime dependency; dependencies and model/data assets retain their own licenses.运行依赖;依赖及模型/数据资源保留各自许可。 · GitHubGitHub · License/notice 1许可/声明 1 ↩ ↩ ↩ ↩
  6. BERT base uncased — Apache-2.0. Official selected text model; preserve its model-repository terms.所选官方文本模型;保留其模型仓库条款。 · Original source原始来源 · License/notice 1许可/声明 1
    Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL-HLT 2019. · Paper论文 ↩ ↩
  7. LM-O data — CC-BY-SA-4.0. Data and model excerpt, not the loader code.数据与模型摘录,不是读取代码。 · Original source原始来源 · GitHub: BOP toolkitGitHub:BOP 工具集
    Brachmann et al. Learning 6D Object Pose Estimation Using 3D Object Coordinates. ECCV 2014. · Paper论文 ↩ ↩ ↩
  8. T-LESS data — CC-BY-4.0. Data and object models; attribute the original dataset.数据与物体模型;须标注原始数据集。 · Original source原始来源
    Hodaň et al. T-LESS: An RGB-D Dataset for 6D Pose Estimation of Texture-less Objects. WACV 2017. · Paper论文 ↩
  9. TUD-L data — CC-BY-SA-4.0. Dataset terms are independent of WAPR source.数据许可独立于 WAPR 源码。 · Original source原始来源
    Hodaň et al. BOP: Benchmark for 6D Object Pose Estimation. ECCV 2018. · Paper论文 ↩ ↩ ↩
  10. IC-BIN · Official BOP dataset pageBOP 官方数据页
    Doumanoglou et al. Recovering 6D Object Pose and Predicting Next-Best-View in the Crowd. CVPR 2016. · Paper论文 ↩ ↩ ↩
  11. YCB-Video (YCB-V) · Official BOP dataset pageBOP 官方数据页 · GitHub: PoseCNNGitHub:PoseCNN
    Xiang et al. PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes. RSS 2018. · Paper论文 ↩

Model weights and dataset/task assets may have separate terms. WAPR's source license does not replace them.模型权重、数据集与任务资源可能有独立条款;WAPR 源码许可不替代这些许可。