WAPR, wide-angle pose refinement
WAPR

6D pose localization and detection6D 位姿定位与检测

Localization and detection share estimate_frame_many_categories_many_instances: predict 2D regions, estimate their poses in batches, then select the outputs. Localization supplies object identities and instance counts through obj_ids and inst_count; detection leaves instance counts unconstrained. Both use RGB-D, camera intrinsics and CAD meshes, and share the score thresholds and output format below.

定位与检测共用 estimate_frame_many_categories_many_instances:预测 2D 区域,批量估计位姿,再筛选输出。定位通过 obj_ids 和 inst_count 提供物体类别及实例数量;检测不预设实例数量。两者均使用 RGB-D、相机内参与 CAD 网格,并采用下文相同的分数阈值和输出格式。

Specify categories and instance counts指定类别与实例数量

The table describes API input combinations. Known-count localization requires both obj_ids and inst_count. Passing inst_count=None leaves the number of instances unconstrained; with obj_ids, it restricts the candidate categories rather than specifying the known-count localization task.

下表列出接口支持的输入组合。已知数量的定位需要同时给出 obj_ids 与 inst_count。若设置 inst_count=None,实例数量便不受约束;此时 obj_ids 仅限定候选类别,不构成已知数量的定位任务。

A single category with a single instance is a localization case too, when its region must first be found. Category count and instance count are constraints on the same frame pipeline. If regions are already supplied, use the pose-only entries instead; whether there is a 2D detector is independent of whether you have one or many instances.

需要先寻找区域时,单类单实例也属于定位用法。类别数与实例数是同一整帧流程上的约束。区域已经给定时,直接使用位姿估计入口;是否需要 2D 检测,与处理单个还是多个实例是两个独立维度。

Case情况obj_idsinst_count
One category, one instance单类、单实例[8]{8: 1}
One category, several instances单类、多实例[8]{8: 3}, or None when the count is unknown;数量未知时用 None
Several categories, one instance each多类、每类单实例[8, 12]{8: 1, 12: 1}
Several categories, several instances each多类、每类多实例[8, 12]{8: 3, 12: 2}, or None;或 None
from wapr.frame import estimate_frame_many_categories_many_instances

# Search for driller regions, then retain at most one estimated pose.
# 先搜索电钻区域,再最多保留一个估计位姿。
poses, timing, detections = estimate_frame_many_categories_many_instances(
    estimator, detector, rgb, depth_m, K, meshes,
    obj_ids=[8], inst_count={8: 1},
)

For frame inference, pass a mesh map made from the prepared snapshots. timing["pose_hot_seconds"] excludes setup and warmup; mesh_setup_seconds, pose_warmup_seconds and detector_warmup_seconds record those separately. hot_frame_seconds adds the measured detection/filter segment and pose segment, excluding the intervening warmup gap and final output assembly. It is not a continuous camera latency. A new pose workload receives three discarded warmup calls before its reported pose time.

整帧推理传入由已准备快照构成的网格字典。timing["pose_hot_seconds"] 不包含准备与预热;mesh_setup_seconds、pose_warmup_seconds 和 detector_warmup_seconds 分别记录这些阶段。hot_frame_seconds 相加实测检测/过滤阶段和位姿阶段,排除中间的预热间隔与最终记录整理,不是相机连续延迟。新位姿负载在报告耗时前丢弃三次预热调用。

The count cap is applied after candidate pose estimation, ranked by score_6d. A cap of one neither guarantees one detection nor limits computation to one candidate. The numbers above illustrate configuration, not the contents of every scene.

数量限制在候选位姿估计之后按 score_6d 排序应用。限制为一,不保证一定检测到一个,也不代表只计算一个候选。表中数字是配置示例,不表示每个场景都含这些实例。

examples/04_6d_localization.py is the localization call on this page. The categories and the count of each category are known before the image is read. The frame is LM-O[2] scene 2, image 307. test_targets_bop19.json names eight categories, ids 1, 5, 6, 8, 9, 10, 11, and 12, and each count is 1. The script writes those two values as obj_ids and inst_count, then calls estimate_frame_many_categories_many_instances. A category that is not in obj_ids is not detected and is not estimated. inst_count is the most that call keeps of that category, higher score_6d first. It does not require the detector to find that many. This call is not estimate_one_category_one_instance. The detection configuration below leaves inst_count unset.

examples/04_6d_localization.py 就是这一节的定位。类别,以及每一类的数量,在读图之前就已经知道。这一帧是 LM-O[2] 场景 2 第 307 帧。test_targets_bop19.json 写了八个类别,编号 1、5、6、8、9、10、11、12,每个数量是 1。脚本把这两个值写成 obj_ids 和 inst_count,再调用 estimate_frame_many_categories_many_instances。不在 obj_ids 里的类别不检测,也不估计。inst_count 是这一类这次最多留下几条,score_6d 高的在前。检测器不一定找得到这么多。这次调用不是 estimate_one_category_one_instance。下文的检测配置不传 inst_count。

Build the arguments first, then pass them in. The copy-this form is a custom folder plus one mask file. load_scene in wapr/scene_files.py returns the image, the depth, the camera, and the meshes of every category in that folder. It does not return a mask, and it does not return a pose. examples/06_custom_scene.py calls that same function, then passes rgb, depth_m, K, and meshes to estimate_frame_many_categories_many_instances. That frame call supports one or several categories and instances. For a supplied single region, take one mesh from the dict and call estimate_one_category_one_instance; to find that region first, use the frame entry with one category and a count cap of one.

先把参数造出来,再传进去。可照抄的写法是一个自定义目录加一张 mask。wapr/scene_files.py 的 load_scene 返回图像、深度、相机,以及这个目录里每个类别的网格。它不返回 mask,也不返回位姿。examples/06_custom_scene.py 调用的是同一个函数,然后把 rgb、depth_m、K、meshes 传给 estimate_frame_many_categories_many_instances。整帧调用支持单个或多个类别与实例。已有单个区域时,从字典中取出一份网格,调用 estimate_one_category_one_instance;需要先寻找区域时,使用整帧入口,并指定一个类别及数量上限一。

import cv2
from wapr import WAPREstimator
from wapr.scene_files import load_scene

rgb, depth_m, K, meshes, names = load_scene(scene_dir)
mesh, diameter_m = meshes[obj_id]
mask = cv2.imread(mask_path, cv2.IMREAD_GRAYSCALE)
n_view = 4
n_inplane = 3

est = WAPREstimator(device="cuda:0")
pose_mesh = est.prepare_meshes([mesh])[0]
mesh_setup_seconds = est.last_mesh_setup_seconds
instances = [{"mesh": pose_mesh, "diameter_m": diameter_m, "mask": mask}]
warmup_seconds = est.warmup_pose(rgb, depth_m, K, instances,
                                n_view=n_view, n_inplane=n_inplane)
out = est.estimate_one_category_one_instance(
    rgb, depth_m, K, pose_mesh, diameter_m,
    bbox_xywh=None, mask=mask,
    n_view=n_view, n_inplane=n_inplane,
)

Prepare CAD geometry once with prepare_meshes, then reuse the returned snapshots in every frame. Centering, normals, mesh packing and GL upload belong to setup. Treat snapshot geometry as immutable: edit the source mesh and prepare a new snapshot when the geometry changes. warmup_pose discards three complete calls on the actual input workload; its returned seconds are separate from hot inference. For a pose timer, synchronize CUDA immediately before and after the estimation call. Model setup, mesh setup, warmup, file reading and drawing stay outside that timer. One-instance estimation uses the same GPU batch path as many-instance estimation.

通过 prepare_meshes 一次性准备 CAD,后续各帧复用返回的网格快照。居中、法线计算、网格打包和 GL 上传属于准备阶段。快照几何不可原地修改:形状变化时编辑源网格,再生成新快照。warmup_pose 用实际输入负载完整预热三次并丢弃输出,单独返回预热秒数。测量位姿耗时时,在估计调用前后立即同步 CUDA;模型准备、网格准备、预热、读盘和绘图均在计时外。单实例估计也走与多实例相同的 GPU 批量计算路径。

estimate_frame_many_categories_many_instances starts from 2D detection and accepts any of the four category/instance combinations above. The estimate_one_category_one_instance call above is one category and one instance, and it does not take a class list. obj_ids is applied first. 2D matching scores only those classes, so a class outside the list is never chosen. 6D estimates only the boxes that remain. One id in obj_ids narrows this frame call to one category, and it can still return many instances of that category. That narrowed call is still not estimate_one_category_one_instance. An id that is not in meshes raises KeyError. An id that is not in the template bank raises ValueError. Omit obj_ids, or pass None, and the list is the mesh ids that are also in the bank. The user guideline does this for IC-BIN[4] scene 1, image 26: object 1 only, even though object 2 is in the full IC-BIN set. That example is one category and many instances, not one instance.

estimate_frame_many_categories_many_instances 是多类别、多实例,从 2D 检测来。上面的 estimate_one_category_one_instance 是单类别、单实例,它不接收类别名单。obj_ids 最先用上。2D 匹配只给这些类别打分,名单以外的类别不会被选中。6D 只估计留下来的框。obj_ids 里只放一个 id,会把这次整帧调用限定为单类别,并且仍可以返回这一类的多个实例。收窄之后仍然不是 estimate_one_category_one_instance。不在 meshes 里的 id 会抛出 KeyError。不在模板库里的 id 会抛出 ValueError。不传 obj_ids,或传 None,名单就是既在网格里、也在模板库里的 id。使用指南里 IC-BIN[4] 场景 1 第 26 帧就是这样:只要物体 1,尽管物体 2 也在完整的 IC-BIN 里。那个例子是单类别、多实例,不是单实例。

The two score gates are arguments, and they run after the class list. score_2d_min is the 2D gate. A box stays when its best allowed class is >= that value. Omit it, or pass None, and the gate is score_threshold in wapr/det2d.py, which is 0.1. Pass another float to use a different gate on this call. score_6d_min is the WBPS gate on score_6d. A pose stays when that score is >= the gate. Omit it, or pass None, and no pose is dropped by this score. The LM-O result uses a 0.50 2D input gate and silhouette NMS, without a WBPS score gate. Its extra driller predictions have WBPS scores around 0.25. A WBPS cutoff is not a general confidence probability or correctness threshold: a saved ROBI[3] chrome-screw pose below meets that case's equivalent-pose ADD criterion with a WBPS score of 0.262. Choose a threshold using validation data for the intended setting.

两个分数阈值都是参数,在类别名单之后起作用。score_2d_min 是 2D 检测阈值:允许类别中的最高分不低于该值,候选框才保留。省略或传入 None 时,采用 wapr/det2d.py 的 score_threshold = 0.1;也可为本次调用传入其他数值。score_6d_min 作用于 WBPS 的 score_6d;省略或传入 None 时,不按位姿分数剔除。LM-O 展示结果采用 0.50 的 2D 输入阈值与轮廓 NMS,没有 WBPS 分数阈值;额外的 driller 预测,其 WBPS 分数约为 0.25。WBPS 截断值并非普适的置信概率或正确性阈值:下文ROBI[3] 铬螺丝案例中,一条满足该案例等价姿态 ADD 判据的位姿,其 WBPS 分数仅为 0.262。实际阈值应根据目标场景的验证数据选择。

The per-class count is the last cut, after both score gates. inst_count maps an obj_id to how many poses of that class to keep. Each class is ordered by score_6d, high first, and only that many rows stay. A class missing from the dict is not capped. None does not cap any class. Same-class masks still pass IoU 0.5 inside the detector, before this count. In the configuration below, 14 is an explicit per-class upper limit; it is not a detected count or a guarantee that fourteen poses pass the score gates.

每个类别的数量是最后一道,在两个分数阈值之后。inst_count 把 obj_id 映射到这一类留下几条。每一类按 score_6d 从高到低排,只留这个数量。字典里没有的类别不截断。None 则每一类都不截断。同一类的 mask 仍先在检测器里按 IoU 0.5 抑制,这一步在数量之前。下面配置中的 14 是显式的每类数量上限,不是检出的数量,也不保证十四条位姿都能通过分数阈值。

from wapr.frame import estimate_frame_many_categories_many_instances

obj_ids = [1]
inst_count = {1: 14}
score_2d_min = 0.1
score_6d_min = 0.5

poses, timing, detections = estimate_frame_many_categories_many_instances(
    estimator, detector, rgb, depth_m, K, meshes,
    obj_ids=obj_ids,
    inst_count=inst_count,
    score_2d_min=score_2d_min,
    score_6d_min=score_6d_min,
    n_view=4, n_inplane=3,
)

n_view and n_inplane are one hypothesis setting for the whole call, not a setting per box. Every bounding box in this call is estimated with that same pair. The default pair is 4 tetrahedron views × 3 in-plane spins, 12 hypotheses. within_group picks one of those 12. Pass other integers to change the pair for every box together. Omit them, or pass None, to read wapr/recipe.py.

n_view 和 n_inplane 是本次调用共用的候选采样设置,所有包围盒都使用这两个参数。默认是 4 个正四面体视角 × 3 次面内旋转,一共 12 个候选姿态。within_group 在这 12 个里选一个。要换次数,就改这一次的这两个整数,整批框一起换。不传,或传 None,就读 wapr/recipe.py。

Argument参数 Where it is built在哪里构造 Call调用
rgb, depth_m, K load_scene in wapr/scene_files.py. examples/06_custom_scene.py calls it.wapr/scene_files.py 的 load_scene。examples/06_custom_scene.py 调用它。 rgb, depth_m, K, meshes, names = load_scene(scene_dir)
mesh, diameter_m The same return. meshes[obj_id] is (mesh, diameter_m). Prepare the inference mesh with estimator.prepare_meshes before warmup or timing, and keep the returned snapshot for subsequent calls. Raw meshes remain accepted for compatibility, but their automatic setup is not hot inference.同一次返回。meshes[obj_id] 是 (mesh, diameter_m)。预热和计时前调用 estimator.prepare_meshes,后续估计保留并复用返回的快照。接口也接受未经准备的网格,但自动准备的耗时不能计作预热后推理。 mesh, diameter_m = meshes[obj_id]
rgb, depth_m, K, mesh, diameter_m wapr.bop.load_bop_rgbd reads RGB, metric depth and intrinsics; load_mesh_m reads metric CAD geometry and diameter. Examples 03, 04, 05 and 07 use these BOP readers. Example 10 reads ROBI[3]'s image, depth and camera.yml directly.wapr.bop.load_bop_rgbd 读取 RGB、米制深度和内参;load_mesh_m 读取米制 CAD 与直径。示例 03、04、05、07 使用这些 BOP 读取函数。示例 10 直接读取 ROBI[3] 图像、深度及 camera.yml。 rgb, depth_m, K = load_bop_rgbd(bop_path, dataset, scene_id, im_id)
mesh, diameter_m = load_mesh_m(bop_path, dataset, obj_id)
rgb, depth_m, K, mesh, diameter_m, mask The one-frame sample. examples/02_one_category_one_instance.py main reads samples/pose_lmo/. That folder is meta.json, not a load_scene folder, so this script does not call load_scene.单帧示例。examples/02_one_category_one_instance.py 的 main 读 samples/pose_lmo/。这个目录是 meta.json,不是 load_scene 的目录,所以这个脚本不调用 load_scene。 cv2.cvtColor(cv2.imread("rgb.png"), cv2.COLOR_BGR2RGB). np.load("depth.npy").astype(np.float32). K and diameter_m from meta.json. trimesh.load("model.ply", force="mesh", process=False). cv2.imread("mask.png", cv2.IMREAD_GRAYSCALE).cv2.cvtColor(cv2.imread("rgb.png"), cv2.COLOR_BGR2RGB)。np.load("depth.npy").astype(np.float32)。K 和 diameter_m 来自 meta.json。trimesh.load("model.ply", force="mesh", process=False)。cv2.imread("mask.png", cv2.IMREAD_GRAYSCALE)。
mask No function in wapr reads a mask file. examples/02_one_category_one_instance.py reads mask.png. examples/08_ycbineoat_one_instance.py predicts its frame-0 mask with WAPRDet2D.wapr 里没有函数去读 mask 文件。examples/02_one_category_one_instance.py 读 mask.png。examples/08_ycbineoat_one_instance.py 用 WAPRDet2D 预测首帧 mask。 mask = cv2.imread(mask_path, cv2.IMREAD_GRAYSCALE)
bbox_xywh Supply four pixel coordinates or read rec["bbox_xywh"] from a detection JSON, as in example 07. When a pose input has no mask, the estimator constructs its box region with wapr.estimator.mask_from_bbox. Example 07 also uses this shared function to prepare missing masks for input NMS.提供四个像素坐标,或像示例 07 一样读取检测 JSON 中的 rec["bbox_xywh"]。位姿输入缺少掩码时,估计器通过 wapr.estimator.mask_from_bbox 构造框区域;示例 07 也用该公共函数为输入 NMS 补齐缺失掩码。 bbox_xywh = [x, y, w, h]
mask = None
n_view, n_inplane Arguments of estimate_one_category_one_instance, estimate_many_categories_many_instances, and estimate_frame_many_categories_many_instances. They are not read from the scene folder. None, or omitting them, uses n_view and n_inplane in wapr/recipe.py. Those values are 4 and 3. make_view_rots in wapr/estimator.py turns the two counts into the initial rotations inside the call.estimate_one_category_one_instance、estimate_many_categories_many_instances、estimate_frame_many_categories_many_instances 的参数。不从场景目录读取。None 或不传,就用 wapr/recipe.py 里的 n_view 和 n_inplane。现在是 4 和 3。wapr/estimator.py 的 make_view_rots 在这次调用里把这两个次数变成初始旋转。 n_view = 4
n_inplane = 3
obj_ids, inst_count, score_2d_min, score_6d_min Arguments of estimate_frame_many_categories_many_instances in wapr/frame.py. They are not arguments of estimate_one_category_one_instance. obj_ids is the class list used by both 2D matching and 6D. score_2d_min and score_6d_min are the two score gates. inst_count is the per-class count after those gates.wapr/frame.py 里 estimate_frame_many_categories_many_instances 的参数。不是 estimate_one_category_one_instance 的参数。obj_ids 是 2D 匹配和 6D 共用的类别名单。score_2d_min 和 score_6d_min 是两个分数阈值。inst_count 是这两个阈值之后、每个类别留下的数量。 obj_ids = [1]
inst_count = {1: 14}
score_2d_min = 0.1
score_6d_min = 0.5

rgb, depth_m, K, and mask are read with numpy.asarray. Each one may be a numpy.ndarray, a rectangular nested list, or a torch.Tensor on CPU. A dict is not accepted. A CUDA tensor is not accepted: asarray raises TypeError. bbox_xywh and diameter_m go through Python float, so those two also accept a CUDA tensor. mesh is not an array.

rgb、depth_m、K、mask 用 numpy.asarray 读入。每一个可以是 numpy.ndarray、整齐的嵌套 list,或 CPU 上的 torch.Tensor。不接受 dict。不接受 CUDA 上的 Tensor:asarray 抛出 TypeError。bbox_xywh 和 diameter_m 走 Python float,这两个也可以是 CUDA 上的 Tensor。mesh 不是数组。

Argument参数 Shape形状 Dtype类型 Restriction限制
rgb (H, W, 3). (H, W) is repeated to 3 channels. (H, W, C) with C > 3 keeps channels 0, 1, 2.(H, W, 3)。(H, W) 会复制成 3 个通道。(H, W, C) 且 C > 3 时只留第 0、1、2 通道。 uint8 is already 0–255 and is cast to float32. Any other numeric dtype, including float32 and float64: if the maximum is ≤ 1.5, the values are 0–1 and are multiplied by 255; if the maximum is > 1.5, the values are already 0–255. The stored image is float32.uint8 按 0–255 使用,再转成 float32。其他数值类型,包括 float32 和 float64:最大值 ≤ 1.5 时当作 0–1,乘 255;最大值 > 1.5 时当作已经是 0–255。存下来的是 float32。 Channel order is RGB, not BGR. (3, H, W) and (N, H, W, 3) are not this call. H and W are the depth_m size. This function does not resize.通道顺序是 RGB,不是 BGR。不接受 (3, H, W),也不接受 (N, H, W, 3)。H 和 W 与 depth_m 相同。这个函数不缩放图像。
depth_m (H, W). A third axis keeps index 0, so (H, W, 1) is read as (H, W).(H, W)。有第三维时只取第 0 个,所以 (H, W, 1) 按 (H, W) 读。 Cast to float32 on entry. float64, float32, and integer dtypes such as uint16 are converted. They are not rescaled.入口转成 float32。float64、float32,以及 uint16 这类整数,只做类型转换,不乘系数。 The numbers are already meters. A uint16 millimeter image is not divided by 1000. Some depth inside the mask must be > 0, or estimate_one_category_one_instance raises RuntimeError("depth").数值已经是米。uint16 的毫米深度图不会除以 1000。mask 里要有 > 0 的深度,否则 estimate_one_category_one_instance 抛出 RuntimeError("depth")。
K (3, 3), or 9 numbers in row-major order. Any other count raises ValueError from reshape(3, 3).(3, 3),或按行排开的 9 个数。元素个数不是 9 时,reshape(3, 3) 抛出 ValueError。 Cast to float32. A float64 intrinsic is accepted and stored as float32.转成 float32。传入 float64 可以,存成 float32。 [[fx, 0, cx], [0, fy, cy], [0, 0, 1]], pixels, OpenCV. Not a normalized intrinsic. fx and fy are the divisors in the back-projection.[[fx, 0, cx], [0, fy, cy], [0, 0, 1]],像素,OpenCV。不是归一化内参。反投影时用 fx 和 fy 做除数。
mesh Vertices (V, 3) inside one trimesh.Trimesh.一个 trimesh.Trimesh,顶点 (V, 3)。 A trimesh.Trimesh, or a path string that trimesh.load(..., force="mesh", process=False) reads. Vertices are cast to float32 when the mesh is centered. Not a vertex ndarray, not a dict, not a tensor.trimesh.Trimesh,或 trimesh.load(..., force="mesh", process=False) 能读的路径字符串。居中时顶点转成 float32。不是顶点 ndarray,不是 dict,不是 Tensor。 The vertex unit is meters. This call does not multiply by 0.001. A load that is not one Trimesh raises TypeError.顶点单位是米。这次调用不乘 0.001。读入结果不是单个 Trimesh 时抛出 TypeError。
diameter_m One number. Not an array.一个数。不是数组。 Python int or float, or a NumPy[1] scalar float32 or float64. float() converts it to a Python float.Python int 或 float,或 NumPy[1] 标量 float32、float64。float() 把它转成 Python float。 Meters, and greater than 0. The crop half-edge is refine_crop_ratio * diameter_m / 2.单位米,且大于 0。裁剪半边长是 refine_crop_ratio * diameter_m / 2。
bbox_xywh Four numbers x, y, w, h. A 1-D sequence of length 4: list, tuple, ndarray (4,), or tensor (4,). A longer sequence drops the tail. Shape (1, 4) raises TypeError.四个数 x, y, w, h。一维、长度为 4:list、tuple、ndarray 的 (4,),或 Tensor 的 (4,)。更长就丢掉后面的。形状 (1, 4) 会抛出 TypeError。 Each item is converted with float. float32, float64, and integers are accepted.每一项用 float 转换。float32、float64 和整数都可以。 Pixels, origin at the top-left, y down. This is xywh, not xyxy. It is read only when mask is None. A passed mask ignores it. If both are missing, ValueError.像素,原点在左上,y 向下。这是 xywh,不是 xyxy。只在 mask is None 时读取。传了 mask 就忽略它。两个都没有则 ValueError。
mask (H, W), the same size as depth_m. A third axis keeps index 0. Not (C, H, W).(H, W),与 depth_m 同尺寸。有第三维时只取第 0 个。不是 (C, H, W)。 No cast at entry. Foreground is value > 0, so uint8 1 or 255, bool True, and a positive float32 or float64 all count. 0 is background. The crop later stores the mask as float32.入口不转换类型。前景是 value > 0,所以 uint8 的 1 或 255、bool 的 True、以及正的 float32 或 float64 都算。0 是背景。裁剪时再存成 float32。 A mask selects wapr_w_mask and ignores bbox_xywh. estimate_one_category_one_instance does not check the size. A different H or W fails in the depth reduction. A dimension of length 1 is broadcast, so pass the depth size. See region inputs and checkpoint selection.传了 mask 就用 wapr_w_mask,并忽略 bbox_xywh。estimate_one_category_one_instance 不检查尺寸。H 或 W 不同时,深度归约会失败。某一维长度为 1 会被广播,所以要传和深度相同的尺寸。见 区域输入与模型选择。
n_view One integer. Not an array of directions.一个整数。不是方向数组。 Python int. None reads recipe.n_view, which is 4.Python int。None 读 recipe.n_view,现在是 4。 At least 1. 4, 6, 8, 12, and 20 are the vertices of the tetrahedron, octahedron, cube, icosahedron, and dodecahedron. Any other count is a Fibonacci sphere. The default 4 is the tetrahedron.不小于 1。4、6、8、12、20 分别是正四面体、正八面体、立方体、正二十面体、正十二面体的顶点。其他数量是斐波那契球面。默认的 4 是正四面体。
n_inplane One integer, the number of spins. Not a list of angles.一个整数,旋转次数。不是角度列表。 Python int. None reads recipe.n_inplane, which is 3.Python int。None 读 recipe.n_inplane,现在是 3。 At least 1. The angles are 2π k / n_inplane about camera +Z; the default 3 gives 0°, 120° and 240°. For both backends, n_view * n_inplane must belong to wapr.pose_groups.wbps_group_sizes: 2–25, plus the products of view counts 4/6/8/12/20 with in-plane counts 3/4/5/6 (34 distinct counts, maximum 120). Unsupported products raise ValueError. The default is 12 hypotheses per instance.不小于 1。绕相机 +Z 的角度为 2π k / n_inplane,默认三次对应 0°、120°、240°。两种后端均要求 n_view * n_inplane 属于 wapr.pose_groups.wbps_group_sizes:2–25,以及视角数 4/6/8/12/20 与面内次数 3/4/5/6 的乘积,共 34 种不同组长,最大为 120。不受支持的乘积抛出 ValueError;默认每实例十二个候选姿态。
obj_ids A sequence of class ids, or None. Not a score.一串类别 id,或 None。不是分数。 Each item is converted with int. A Python int is the usual item. This argument belongs to estimate_frame_many_categories_many_instances, and to WAPRDet2D.detect_many_categories_many_instances under the same name.每一项用 int 转换。通常是 Python int。这个参数属于 estimate_frame_many_categories_many_instances,WAPRDet2D.detect_many_categories_many_instances 用同一个名字。 None uses the mesh ids that are also in the template bank. A passed list is the only set that 2D matching and 6D estimation will use. An id missing from meshes raises KeyError. An id missing from the bank raises ValueError.None 用既在网格里、也在模板库里的 id。传入的名单是 2D 匹配和 6D 估计唯一会用的那一套。不在 meshes 里会抛出 KeyError。不在模板库里会抛出 ValueError。
score_2d_min One number, or None. Not per box.一个数,或 None。不是每个框一个。 Python float, or a value float() accepts. None reads score_threshold in wapr/det2d.py.Python float,或 float() 能转的值。None 读 wapr/det2d.py 的 score_threshold。 The recommended default is 0.1, and that is the value of score_threshold. A box stays when its best allowed class is >= this gate. Pass another float to change it. predict takes the same gate as score_min.推荐的默认值是 0.1,也就是 score_threshold 现在的值。允许类别里的最高分不低于这个阈值,包围盒才留下。传入另一个浮点数即可替换该值。predict 用 score_min 接收同一个阈值。
score_6d_min One number, or None. Not per pose until you pass one gate for the whole call.一个数,或 None。传入的是这一次共用的一个阈值。 Python float, or a value float() accepts. None does not filter.Python float,或 float() 能转的值。None 不过滤。 Keeps poses with score_6d at or above the gate. The default None disables this filter. The value 0.5 in the example is illustrative; select an operating threshold on separate validation data for the target setting.保留 score_6d 不低于阈值的位姿。默认 None 不按该分数过滤;示例中的 0.5 仅用于演示,实际阈值需在目标场景的独立验证数据上选择。
inst_count A dict from obj_id to a count, or None.obj_id 到数量的 dict,或 None。 Keys and values are converted with int. A count of 0 keeps none of that class. A negative count raises ValueError.键和值都用 int 转换。数量 0 则这一类一条不留。负数抛出 ValueError。 Applied after both score gates. Within a class, higher score_6d comes first. A class missing from the dict is not capped. None caps nothing.在两个分数阈值之后。同一类里 score_6d 高的在前。字典里没有的类别不截断。None 则不截断。

The supplied-region estimate_one_category_one_instance call returns a dict. pose_4x4 is a NumPy[1] float32 array of shape (4, 4), object to camera. R is float32 (3, 3). t_m is float32 (3,), meters. score_6d is a Python float, (100 - max(between_group)) / 200. within_group picks the hypothesis inside the group. A larger between_group makes the returned score smaller.

给定区域的 estimate_one_category_one_instance 调用返回一个 dict。pose_4x4 是 NumPy[1] float32,形状 (4, 4),物体到相机。R 是 float32 的 (3, 3)。t_m 是 float32 的 (3,),米。score_6d 是 Python float,等于 (100 - max(between_group)) / 200。within_group 在这一组里选候选姿态。between_group 越大,返回的分数越小。

Detection without known instance counts实例数量未知时的检测

examples/05_bop_6d_detection.py applies full 6D detection to one BOP frame without a known instance count. estimate_frame_many_categories_many_instances first runs the 2D detector, then passes masks containing finite, positive depth to one estimate_many_categories_many_instances pose batch. A detection without usable depth remains in the 2D return but cannot yield a pose. The visible mask defines the pose region; the detector box is retained as metadata. This configuration leaves inst_count unset, so it does not cap a category's instance count.

examples/05_bop_6d_detection.py 在不预先知道实例数量的情况下,对一帧 BOP 数据执行完整 6D 检测。estimate_frame_many_categories_many_instances 先运行 2D 检测器,再把 mask 内具有有限正深度的实例送入一次 estimate_many_categories_many_instances 位姿批处理。没有可用深度的检测仍保留在 2D 返回值中,但不能产生位姿。可见 mask 定义位姿区域,检测框只作为元数据保留。此配置不设置 inst_count,因此不限制每类的实例数量。

from wapr.frame import estimate_frame_many_categories_many_instances

poses, timing, detections = estimate_frame_many_categories_many_instances(
    estimator, detector, rgb, depth_m, K, meshes,
    obj_ids=None,
    inst_count=None,
    score_2d_min=0.1,
    score_6d_min=None,
    n_view=4, n_inplane=3,
)

meshes maps obj_id to (mesh, diameter_m). Each item in poses has obj_id, score_2d, bbox_xywh, pose_4x4, R, t_m, score_6d, and time_s. score_2d is the detector ranking score. It is not score_6d. time_s is that pose call divided by the number of instances that entered it, in seconds. Rows dropped by score_6d_min or inst_count are absent, so the kept rows sum to the batch only when nothing was dropped after the batch. Repeated obj_id values stay separate instances. name is included when the mesh metadata has one.

meshes 把 obj_id 映射到 (mesh, diameter_m)。poses 的每一项含 obj_id、score_2d、bbox_xywh、pose_4x4、R、t_m、score_6d、time_s。score_2d 是检测器的排序分,不是 score_6d。time_s 是这次位姿调用总耗时除以进入调用的实例数,单位秒。被 score_6d_min 或 inst_count 删掉的行不在列表里,所以只有 batch 之后没有再删行时,留下的行才加回这次调用。重复的 obj_id 仍是各自的实例。网格 metadata 里有 name 时,位姿里也有它。

The third return, detections, contains every row that passed the 2D gate, including masks without usable depth. Each row has obj_id, score_2d, bbox, and mask. The detector visualization draws blue boxes, tinted masks, and labeled 2D scores; examples 04 and 05 show only their 12 highest-scoring rows to keep the image legible, while the full rows still enter the depth check. The 6D visualization draws red boxes, rendered contours, and labeled WBPS score_6d. Example 05 also limits this illustration to the 12 highest WBPS scores; inference itself is unchanged. recipe.visualize controls whether the 6D image is saved.

第三个返回值 detections 包含通过 2D 阈值的全部检测行,其中也可能有 mask 内无有效深度的行。每行有 obj_id、score_2d、bbox 和 mask。2D 可视化绘制蓝色框、带颜色的 mask,并标注 2D 分数;示例 04、05 的插图只展示分数最高的 12 条,避免标签遮挡,但完整检测结果仍进入深度检查。6D 可视化绘制红色框、渲染轮廓,并标注 WBPS 的 score_6d。示例 05 也只在插图中展示 WBPS 分数最高的 12 条,推理本身不变。recipe.visualize 控制是否保存 6D 图。

Rendered-contour suppression is a separate postprocessing step. The published LM-O figure suppresses overlapping poses of the same CAD when contour IoU exceeds 0.5, retaining the higher score_6d. Example 10 applies the same criterion with rendered_mask_iou_nms. Neither operation is part of estimate_frame_many_categories_many_instances. The shared CUDA implementation in wapr/suppression.py keeps silhouette intersections, ranking and greedy suppression on GPU; only the final indices return to CPU. Silhouettes retain the original image resolution, though pixel-center rasterization can differ at boundary pixels from OpenCV polygon filling.

渲染轮廓抑制是独立的后处理步骤。发布页的 LM-O 图对同一 CAD 的位姿候选比较轮廓 IoU:超过 0.5 时保留 score_6d 较高者。示例 10 通过 rendered_mask_iou_nms 应用相同准则。这两处后处理均不属于 estimate_frame_many_categories_many_instances。wapr/suppression.py 的共享 CUDA 实现将轮廓交集、排序与贪心抑制保留在 GPU,仅将最终下标传回 CPU。轮廓保持原图分辨率,但按像素中心光栅化得到的边界像素可能与OpenCV 多边形填充结果不同。

Example 07 uses greedy_mask_nms_across_categories only for detector = "json", to suppress supplied masks across categories. This input filter is distinct from both same-category 2D mask NMS and pose-contour suppression.

示例 07 仅在 detector = "json" 时通过 greedy_mask_nms_across_categories 对外部掩码执行跨类别抑制。该输入过滤与同类别 2D 掩码 NMS及位姿轮廓抑制分别对应不同阶段。

References and licenses参考文献与许可

  1. numpy — BSD License. Runtime dependency; dependencies and model/data assets retain their own licenses.运行依赖;依赖及模型/数据资源保留各自许可。 · Original source原始来源 · License/notice 1许可/声明 1 ↩ ↩ ↩ ↩
  2. LM-O data — CC-BY-SA-4.0. Data and model excerpt, not the loader code.数据与模型摘录,不是读取代码。 · Original source原始来源 · GitHub: BOP toolkitGitHub:BOP 工具集
    Brachmann et al. Learning 6D Object Pose Estimation Using 3D Object Coordinates. ECCV 2014. · Paper论文 ↩ ↩
  3. ROBI data — Not separately verified / 未单独核实. The saved public poses and dataset are credited to ROBI; an independent grant to redistribute them has not been verified.保存的公开位姿和数据均注明 ROBI 来源;其再分发授权尚未独立核实。 · GitHubGitHub
    Yang et al. ROBI: A Multi-View Dataset for Reflective Objects in Robotic Bin-Picking. IROS 2021. · Paper论文 ↩ ↩ ↩ ↩
  4. IC-BIN · Official BOP dataset pageBOP 官方数据页
    Doumanoglou et al. Recovering 6D Object Pose and Predicting Next-Best-View in the Crowd. CVPR 2016. · Paper论文 ↩ ↩

Model weights and dataset/task assets may have separate terms. WAPR's source license does not replace them.模型权重、数据集与任务资源可能有独立条款;WAPR 源码许可不替代这些许可。