Custom data自定义数据
For pose estimation, provide RGB-D, camera intrinsics, a metric mesh and a mask or box. The integrated detector can obtain regions when none are supplied. Depth and mesh coordinates use meters; camera intrinsics use pixels.
位姿估计需要 RGB-D、相机内参、米制网格及 mask 或框。未提供区域时可使用一体化检测器。深度与网格坐标单位为米,相机内参单位为像素。
To use your own capture, provide an RGB image, depth, camera intrinsics and the CAD meshes to search for. load_scene reads these inputs from a folder; the detector predicts instance regions and the pose estimator processes all candidates from the frame in batches.
接入自己的采集数据时,准备 RGB 图像、深度、相机内参和需要识别的 CAD 网格。load_scene 从目录读取这些输入;检测器预测实例区域,位姿估计器对同帧候选批量推理。
This page uses samples/own_block: HCCEPose RGB-D frame 000003, captured with an iHawk100R RGB-D camera (manufacturer's product page), with several white printed wedges and one supplied demo-bin-picking object-1 mesh. One CAD category corresponds to multiple instances. The mesh is supplied directly; this example performs detection and pose estimation.
本页使用 samples/own_block:HCCEPose 使用 iHawk100R RGB-D 相机采集的第 000003 帧(厂商产品页面),其中有多只白色打印三角块,并提供一份 demo-bin-picking 物体 1 网格。一份 CAD 类别对应多个实例,直接用已提供的网格进行检测与位姿估计。
The author's approximate camera prices are CNY 250 (US$37) used and CNY 1,000 (US$149) new. USD equivalents use the 5 October 2026 reference rate of approximately CNY 6.705 per US dollar.
作者提供的相机参考价格为:二手约人民币 250 元(约 37 美元),全新约人民币 1,000 元(约 149 美元)。美元按 2026-10-05 参考汇率 1 美元约合 6.705 元人民币换算。
Wedge capture, frame 000003: RGB on the left and its depth on the right, on the same 400 × 640 pixel grid. The colorbar reports distance along the camera's optical axis in meters, spanning all valid depths in this frame (0.301–0.662 m): purple is nearer, yellow is farther, and black indicates missing or invalid depth. The source PNG uses millimeters; load_scene applies depth_scale = 1 and divides by 1000. Camera intrinsics and the object mesh model are supplied separately; the predicted detections and poses are shown below.
三角块实拍第 000003 帧:左侧为 RGB,右侧为对应深度,两者使用相同的 400 × 640 像素网格。色条表示沿相机光轴方向的距离,单位为米,覆盖本帧全部有效深度(0.301–0.662 米):紫色较近、黄色较远,黑色表示缺失或无效深度。原始 PNG 的单位为毫米,load_scene 按 depth_scale = 1 读取后除以 1000。相机内参与物体网格模型另行提供;下方展示检测与位姿预测结果。
Input files输入文件
my_scene/
├── rgb.png
├── depth.png
├── camera.json
├── objects.json
└── models/
└── wedge.ply
Provide a mesh in the units declared by mesh_unit. Use the object’s model diameter for diameter_mm when available; the loader otherwise uses the mesh-bounds diagonal. The detector renders CAD templates from this mesh, so its geometry should represent the object to be recognized.
网格单位须与 mesh_unit 一致。已有模型直径时,在 diameter_mm 中填写;省略时,读取器使用网格包围盒对角线。检测器根据这份网格渲染 CAD 模板,因此网格几何应对应需要识别的物体。
Provide an RGB image, a depth image or meter-valued depth.npy, camera.json, objects.json, and the meshes referenced by the object list. The two JSON files from samples/own_block are shown below. Rotational symmetry is defined with the mesh preparation rather than in this frame manifest.
每帧需提供 RGB 图像、深度图或已换算为米的 depth.npy、camera.json、objects.json,以及物体列表引用的网格。下面列出 samples/own_block 的两份 JSON。旋转对称性在网格准备阶段定义,不放在这份帧清单里。
camera.json. fx, fy, cx, cy are pixels. The depth png times depth_scale is millimeters. Here depth_scale is 1, so the png values are already millimeters. The loader then divides by 1000 and passes meters to the pose call.
camera.json。fx、fy、cx、cy 是像素。深度 png 乘 depth_scale 是毫米。这里 depth_scale 是 1,所以 png 里的数已经是毫米。读取器再除以 1000,送给位姿调用的是米。
{
"fx": 390.9188, "fy": 390.9188, "cx": 318.9914, "cy": 202.7166, "depth_scale": 1.0
}
objects.json. mesh_unit is "mm" or "m". This wedge ply is millimeters, so the loader multiplies its vertices by 0.001. One entry defines one class. file is resolved relative to scene_dir; models/wedge.ply points to the mesh inside this scene folder. diameter_mm may be omitted, in which case the loader uses the mesh-bounds diagonal. The wedge entry retains the 118.322 mm diameter provided by HCCEPose.
objects.json 的 mesh_unit 取 "mm" 或 "m"。这份三角块 ply 以毫米计,读取器将顶点乘以 0.001。一条记录定义一个类别。file 相对 scene_dir 解析;models/wedge.ply 指向该场景目录内的网格。diameter_mm 可以省略,此时读取器采用网格包围盒对角线。三角块记录保留 HCCEPose 给出的 118.322 mm 直径。
{
"mesh_unit": "mm",
"objects": [
{"id": 1, "name": "wedge", "file": "models/wedge.ply", "diameter_mm": 118.322}
]
}
Run the example运行示例
Example examples/06_custom_scene.py defaults to samples/own_lmo. On its first run, it downloads the LM-O[2] sample if needed and creates this custom folder from scene 2, frame 307, referencing the eight downloaded CADs. To use your own capture, prepare the input files shown above and set scene_dir to their folder. The wedge capture illustrated on this page is not included in the downloadable sample packs; using samples/own_block requires supplying its RGB-D frame and mesh. Leave template_path empty to build the CAD template bank in GPU memory; set a path to save and reuse the bank.
示例 examples/06_custom_scene.py 默认读取 samples/own_lmo。首次运行时,按需下载 LM-O[2] 小样,并用场景 2 第 307 帧生成这个自定义目录,引用已下载的八份 CAD。接入自己的采集数据时,按上方格式准备输入文件,将 scene_dir 改为对应目录。本页展示的三角块实拍数据不包含在可下载小样包中;使用 samples/own_block 需提供其 RGB-D 图像与网格。template_path 留空时在显存中建立 CAD 模板库;需要保存并复用模板库时再填写路径。
A separate timing run uses this wedge input on RTX 5090, OpenGL + TensorRT[1] FP16, with twelve hypotheses per region, three WAPR updates and two SAPR updates. Three complete warmups precede ten CUDA-synchronized calls with resident CAD meshes: median 2D detection 190.4 ms, pose stage 130.4 ms and complete detection-to-pose call 324.0 ms for ten input regions. The continuous total includes result assembly and per-frame transfers; it excludes loading, template/CAD preparation, warmup, drawing and final pose NMS. The interactive view below illustrates the pose selection and NMS results.
同一三角块输入另行在 RTX 5090、OpenGL + TensorRT[1] FP16 下测速,每区域十二个候选姿态,三次 WAPR 更新及两次 SAPR 更新。CAD 提前准备并驻留,三次完整预热后进行十次 CUDA 同步调用:十个输入区域的 2D 检测中位数为 190.4 毫秒,位姿阶段 130.4 毫秒,检测到位姿的连续调用合计 324.0 毫秒。合计包含结果整理及逐帧传输,不含加载、模板与 CAD 准备、预热、绘图及最终位姿 NMS。下方交互视图展示位姿筛选与 NMS 结果。
For the wedge input, set scene_dir to "samples/own_block" in example 06. The Out excerpt is the recorded input-shape line for that choice; inference logs continue with dictionaries containing obj_id, name, score_2d and score_6d. No hand-rounded pose rows are shown as stdout.
三角块输入需在示例 06 中将 scene_dir 设为 "samples/own_block"。Out 摘录为该设置下已记录的输入形状行;推理日志随后打印含 obj_id、name、score_2d 与 score_6d 的字典,不将手工四舍五入的位姿行展示为标准输出。
rgb, depth_m, K, meshes, names = load_scene(scene_dir)
print("rgb", tuple(rgb.shape), "depth_m", tuple(depth_m.shape), "K", tuple(K.shape), "classes", names, flush=True)Source: examples/06_custom_scene.py, lines 106–107代码来源:examples/06_custom_scene.py,第 106–107 行
mesh_only = {obj_id: pair[0] for obj_id, pair in meshes.items()}
if template_path:
onboard_meshes(mesh_only, template_path, device=device)
template = template_path
else:
template = onboard_meshes(mesh_only, "", device=device)
detector = WAPRDet2D(template, device=device, backend=det_backend)
estimator = WAPREstimator(device=device)
prepared = estimator.prepare_meshes([pair[0] for pair in meshes.values()])
mesh_setup_seconds = estimator.last_mesh_setup_seconds
pose_meshes = {obj_id: (mesh, pair[1]) for (obj_id, pair), mesh in zip(meshes.items(), prepared)}Source: examples/06_custom_scene.py, lines 127–137代码来源:examples/06_custom_scene.py,第 127–137 行
poses, timing, detections = estimate_frame_many_categories_many_instances(estimator, detector, rgb, depth_m, K, pose_meshes)
print("POSE_TIMING", timing, "model_setup_seconds", estimator.model_setup_seconds,
"initial_mesh_setup_seconds", mesh_setup_seconds, flush=True)
print("det_ms", timing["total_ms"], "detections", len(detections), "instances", len(poses), flush=True)
if visualize_det:
preview_detections = sorted(detections, key=lambda row: -float(row["score_2d"]))[:preview_top_k]
print("2d_preview", len(preview_detections), "of", len(detections), flush=True)
visualize_2d_detection(rgb, preview_detections, names=names, image=True, filename="custom_scene_det.jpg")
rows = []
for pose in poses:
row = dict(pose)
row["mesh"] = meshes[int(pose["obj_id"])][0]
print(
{
"obj_id": pose["obj_id"],
"name": pose.get("name", names.get(int(pose["obj_id"]), "")),
"score_2d": pose["score_2d"],
"score_6d": pose["score_6d"],
},
flush=True,
)
rows.append(row)Source: examples/06_custom_scene.py, lines 140–161代码来源:examples/06_custom_scene.py,第 140–161 行
rgb (400, 640, 3) depth_m (400, 640) K (3, 3) classes {1: 'wedge'}Captured stdout excerpt · 已保存标准输出摘录 ·
Above, the frame. Below, the saved poses on the sensor depth.上面是这一帧,下面是保存的位姿叠在传感器深度上。
This view selects nine of ten candidates, omitting the lowest-score candidate (0.432). GPU rendered-mask NMS removes two duplicates, leaving seven displayed poses. Mask IoU above 0.5 suppresses the lower WBPS score; this postprocessing is excluded from the independent timings above. The GPU mask NMS implementation is in wapr.suppression. Each chip and the selected photo label show the original instance index, 2D detection score and 6D pose score; the 6D score uses three decimal places. Click a chip, a box on the photo, or a mesh in the view. The chosen wedge is the red box and the bright mesh. This capture has no pose label. These saved predictions use 4 views, 3 in-plane rotations, 3 WAPR updates and 2 SAPR updates. The measurements above process ten input regions before this view’s filtering. Drag the view to orbit the instance you clicked. Several copies of one object stay separate. Scroll to zoom.
交互视图从十条候选中选取九条,略去分数最低的 0.432 候选;GPU 渲染 mask NMS 去掉两个重复项,保留七个位姿:mask IoU 大于 0.5 时抑制 WBPS 分数较低者;这一后处理不计入上方独立测速。该筛选使用 wapr.suppression 中的 GPU mask NMS。每个控件与图上选中实例的标签同时显示原始实例序号、2D 检测分数和 6D 位姿分数;6D 分数保留三位小数。点一个控件、照片上的一个框,或视图里的一块网格。选中的三角块是红框和亮着的网格。这次采集没有位姿标注,保存的预测使用 4 个视角、3 个平面内旋转、3 次 WAPR 更新与 2 次 SAPR 更新。上方测速处理筛选前的十个输入区域。拖动视图时绕你点中的那一个实例转,同一物体的多份实例分别显示,滚轮缩放。
References and licenses参考文献与许可
- tensorrt-cu12 — Other/Proprietary License. Runtime dependency; dependencies and model/data assets retain their own licenses.运行依赖;依赖及模型/数据资源保留各自许可。 · GitHubGitHub · License/notice 1许可/声明 1 ↩ ↩
- LM-O data — CC-BY-SA-4.0. Data and model excerpt, not the loader code.数据与模型摘录,不是读取代码。 · Original source原始来源 · GitHub: BOP toolkitGitHub:BOP 工具集
Model weights and dataset/task assets may have separate terms. WAPR's source license does not replace them.模型权重、数据集与任务资源可能有独立条款;WAPR 源码许可不替代这些许可。