WAPR, wide-angle pose refinement
WAPR

Datasets and model resources数据集与模型资源

BOP datasetsBOP 数据集

This section collects the seven BOP-Classic core datasets, the included frame excerpts, official full-dataset downloads, ROBI[10] and tracking inputs, and model resources. Single-frame excerpts are for checking input conventions and examples; use the official target lists and complete data for benchmark evaluation.

本节集中介绍七个 BOP-Classic 核心数据集、随包单帧摘录、官方完整集下载,以及 ROBI[10]、跟踪序列与模型资源。单帧摘录用于核对输入格式和运行示例;基准评测须使用官方目标列表与完整数据。

Photographs of LM-O, YCB-V, T-LESS, TUD-L, HB, ITODD, and IC-BIN
A sample image from each of the seven BOP datasets. The dataset name appears below each tile.七个 BOP 数据集的样例图像;各图下方标有数据集名称。

Bundled sample frames随包提供的示例帧

LM-O[4], YCB-V[12], T-LESS[5], and TUD-L[6] each have one demo frame in SEU-WYL/WAPR. Set only in wapr/download_assets.py to that name, then run the command. only = () fetches the weights and every one-frame pack. examples/07_write_bop_pose_csv.py with bop_path left empty fetches the one frame for the dataset you set, into samples/bop/<dataset>/. A full test list still needs the official root in bop_path.

LM-O[4]、YCB-V[12]、T-LESS[5]、TUD-L[6] 在 SEU-WYL/WAPR 里各有一帧演示。把 wapr/download_assets.py 里的 only 设成那个名字,再运行下面的命令。only = () 会获取权重和每一个单帧示例。examples/07_write_bop_pose_csv.py 的 bop_path 为空时,按所填写的 dataset 将该帧下载到 samples/bop/<dataset>/。完整测试列表仍要把官方根目录填进 bop_path。

python -m wapr.download_assets

Unpack a full set so bop_path/<name>/ contains test_targets_bop19.json, models/, and the image split this package reads. T-LESS and HB[13] use test_primesense/. The others use test/. The BOP’19–24 zip is the image subset named in the targets file. Training images stay on the dataset’s BOP section; the examples here do not read them.

完整集解压后,bop_path/<name>/ 里要有 test_targets_bop19.json、models/,以及本包要读的图像划分。T-LESS 和 HB[13] 用 test_primesense/。其他用 test/。BOP’19–24 那个包是 targets 文件点名的图像子集。训练图像留在该数据集的 BOP 小节里,这里的示例不读它们。

LM-O

LM-O (Linemod-Occluded; Brachmann et al., ECCV 2014) is used to evaluate 6D object pose localization and detection in cluttered scenes. It adds pose annotations for eight objects in one LM test set, with models originating from Hinterstoisser et al., ACCV 2012.

Its main challenge is occlusion: only part of a target may be visible, while neighboring objects contribute distracting edges and depth points. For pose correction, the model must align with the visible target surface without being pulled toward an occluder; pose scoring must distinguish plausible projections supported by too little evidence.

LM-O(Linemod-Occluded;Brachmann 等,ECCV 2014)用于评测杂乱场景中的 6D 位姿定位与检测。它在 LM 的一个测试集中为八个物体补充位姿标注,物体模型源自 Hinterstoisser 等的 ACCV 2012 工作。

主要难点是遮挡:目标可能只露出局部,周围物体又会带来干扰边缘和深度点。对位姿修正而言,需要依据目标的可见表面完成对齐,避免被遮挡物带偏;对候选评分而言,需要识别那些投影看似合理、实际观测支持却不足的姿态。

only = ("pose_lmo",) lands in samples/pose_lmo/: scene 2, image 1, object 12 (holepuncher), depth and mesh in meters. examples/02_one_category_one_instance.py reads that folder. only = ("lmo",) is BOP millimeters at samples/bop/lmo/. models/ has the eight LM-O meshes: 1 ape, 5 can, 6 cat, 8 driller, 9 duck, 10 eggbox, 11 glue, 12 holepuncher. The folder has scene 2 images 1 and 307. Image 1 is not a row of test_targets_bop19.json. Image 307 is. cnos-fastsam_scene2_im307.json is the CNOS[16]-FastSAM[14] 2D detections for image 307. The LM-O tile on this page is that same image.

only = ("pose_lmo",) 落到 samples/pose_lmo/:场景 2 第 1 帧,物体 12(打孔器 holepuncher),深度和网格是米。examples/02_one_category_one_instance.py 读这个目录。only = ("lmo",) 是 BOP 毫米版,在 samples/bop/lmo/。models/ 里是 LM-O 的八个网格:1 ape、5 can、6 cat、8 driller、9 duck、10 eggbox、11 glue、12 holepuncher。目录里有场景 2 的第 1 帧和第 307 帧。第 1 帧不在 test_targets_bop19.json 里。第 307 帧在。cnos-fastsam_scene2_im307.json 是第 307 帧的 CNOS[16]-FastSAM[14] 2D 检测。本页拼图里的 LM-O 就是这一帧。

LM-O scene 2 image 307, predicted pose in red and ground truth in green
Scene 2, image 307, 8 instances. The red 3D bbox is the predicted pose. The green 3D bbox is the ground-truth pose.场景 2,第 307 帧,8 个实例。红色 3D 包围盒是预测位姿。绿色 3D 包围盒是真值位姿。

YCB-V

YCB-V (YCB-Video; Xiang et al., RSS 2018) contains RGB-D observations of 21 everyday YCB objects in 92 videos. It supports studying object pose estimation for robotic manipulation; the BOP version evaluates individual frames specified in test_targets_bop19.json.

Clutter, mutual occlusion and object symmetry make both object identification and rotation recovery difficult. For pose correction and scoring, the challenge is to use the remaining appearance and depth evidence while accounting for equivalent poses of symmetric objects. The videos also provide changing viewpoints, although the BOP evaluation here uses its selected frames.

YCB-V(YCB-Video;Xiang 等,RSS 2018)包含 21 个 YCB 日常物体的 92 段 RGB-D 视频,可用于研究机器人操作中的物体位姿求解。BOP 版本按 test_targets_bop19.json 指定的单帧进行评测。

主要难点包括杂乱背景、物体之间的遮挡,以及对称物体的旋转歧义。位姿修正和评分需要结合剩余的外观与深度信息,并考虑对称物体的等价位姿。视频包含连续变化的观察视角,但这里的 BOP 评测使用选定帧。

only = ("ycbv",) lands in samples/bop/ycbv/: scene 50, image 1130. Five official target rows. The mesh in the pack is object 5, 006_mustard_bottle. The YCB-V tile is scene 56, image 1. The robot-arm clips of the cracker box, mustard bottle, and sugar box are YCBInEOAT[9], a separate download from this test split.

only = ("ycbv",) 落到 samples/bop/ycbv/:场景 50,第 1130 帧。官方 target 有五行。示例包里的网格是物体 5,006_mustard_bottle。拼图里的 YCB-V 是场景 56 第 1 帧。机械臂上的饼干盒、芥末瓶和糖盒是 YCBInEOAT[9],和这一份测试划分分开下载。

YCB-V scene 56 image 1, predicted pose in red and ground truth in green
Scene 56, image 1, 4 instances. The red 3D bbox is the predicted pose. The green 3D bbox is the ground-truth pose.场景 56,第 1 帧,4 个实例。红色 3D 包围盒是预测位姿。绿色 3D 包围盒是真值位姿。

T-LESS

T-LESS (Hodan et al., WACV 2017) evaluates detection and 6D pose recovery of 30 industrial objects with little texture or distinctive color. Many objects have symmetries or resemble one another in shape and size; some are components of other objects.

The difficult cases combine similar categories, repeated instances, clutter and substantial occlusion. Appearance alone may not identify a category or determine its rotation, so pose correction and scoring must use geometric detail and visible surfaces while respecting symmetry. The BOP archives used here contain Primesense Carmine images; Kinect and Canon images are available on the dataset site.

T-LESS(Hodan 等,WACV 2017)用于评测 30 个工业物体的检测与 6D 位姿求解。这些物体缺少明显纹理和可区分的颜色,不少具有对称性,或在形状、尺寸上十分接近;部分物体还是其他物体的组成部件。

难例同时包含相似类别、重复实例、杂乱背景和严重遮挡。仅凭外观往往难以判断类别或旋转,因此位姿修正与评分需要利用几何细节和可见表面,并考虑对称性。本页使用的 BOP 压缩包包含 Primesense Carmine 图像;Kinect 和 Canon 图像可从数据集网站获取。

only = ("tless",) lands in samples/bop/tless/: Primesense scene 1, image 1. Four official target rows. The mesh is object 30. The T-LESS tile is image 67 of that scene. Image 1 is the first T-LESS row in samples/validation/det2d/selection.json, a 10-frame smoke list, not a reported score.

only = ("tless",) 落到 samples/bop/tless/:Primesense 场景 1 第 1 帧。官方 target 有四行。网格是物体 30。拼图里的 T-LESS 是该场景第 67 帧。第 1 帧是 samples/validation/det2d/selection.json 里 T-LESS 的第一行,那是 10 帧的冒烟列表,不是公布的分数。

T-LESS Primesense scene 1 image 67, predicted pose in red and ground truth in green
Primesense scene 1, image 67, 4 instances. The red 3D bbox is the predicted pose. The green 3D bbox is the ground-truth pose.Primesense 场景 1,第 67 帧,4 个实例。红色 3D 包围盒是预测位姿。绿色 3D 包围盒是真值位姿。

TUD-L

TUD-L (TUD Light; Hodan, Michel et al., ECCV 2018) records three moving objects under eight lighting conditions. It is used to examine how illumination changes affect 6D object pose recovery.

For an appearance-based pose method, lighting can change brightness, shadows and highlights without a corresponding change in object geometry. This makes matching an observation to a rendered model harder. Pose correction and scoring therefore need to distinguish lighting differences from geometric misalignment as the object moves.

TUD-L(TUD Light;Hodan、Michel 等,ECCV 2018)记录三个运动物体在八种光照条件下的图像序列,用于考察光照变化对 6D 位姿求解的影响。

对依赖外观的位姿算法而言,光照会改变亮度、阴影和高光,而物体几何并未随之变化,这会增加观测图像与模型渲染图之间的匹配难度。位姿修正和评分需要在物体运动时区分光照差异与几何错位。

only = ("tudl",) lands in samples/bop/tudl/: scene 1, image 0. One official target row. The mesh is object 1. The TUD-L tile is image 5481 of that scene. Image 0 is the first TUD-L row of the same smoke list.

only = ("tudl",) 落到 samples/bop/tudl/:场景 1 第 0 帧。官方 target 有一行。网格是物体 1。拼图里的 TUD-L 是该场景第 5481 帧。第 0 帧是同一份冒烟列表里 TUD-L 的第一行。

TUD-L scene 1 image 5481, predicted pose in red and ground truth in green
Scene 1, image 5481, 1 instance. The red 3D bbox is the predicted pose. The green 3D bbox is the ground-truth pose.场景 1,第 5481 帧,1 个实例。红色 3D 包围盒是预测位姿。绿色 3D 包围盒是真值位姿。

Other datasets on the project page项目主页展示的其他数据集

HB, ITODD[7], and IC-BIN[11] appear in the same localization block. download_assets has no one-frame pack for them. A full run uses the BOP archives and bop_path.

HB、ITODD[7] 和 IC-BIN[11] 出现在同一节定位里。download_assets 没有它们的单帧示例。完整运行用 BOP 压缩包和 bop_path。

HB

HB (HomebrewedDB; Kaskman et al., ICCVW 2019) contains 33 toys, household objects and industrial parts in 13 scenes. It studies 6D pose recovery from 3D models, including generalization across occlusion, lighting changes and changes in object appearance.

The challenge is to retain reliable geometric alignment when appearance differs from the model or from training images, and to handle both textured and textureless objects in scenes of varying complexity. This package reads the Primesense split, test_primesense/; Kinect images are also available. Ground-truth poses are public for training and validation, but not for test images.

HB(HomebrewedDB;Kaskman 等,ICCVW 2019)包含 33 个玩具、家用物品及工业零件,共 13 个场景。它用于研究基于 3D 模型的 6D 位姿求解,考察算法面对遮挡、光照变化和物体外观变化时的泛化能力。

难点是在观测外观与模型或训练图像不同的情况下仍保持可靠的几何对齐,同时处理有纹理和无纹理物体,以及不同复杂度的场景。本包读取 Primesense 划分 test_primesense/,数据集另提供 Kinect 图像。训练与验证图像公开真值位姿,测试图像的真值不公开。

HB validation scene 1 image 1, predicted pose in red and ground truth in green
Validation scene 1, image 1, 3 instances. The red 3D bbox is the predicted pose. The green 3D bbox is the ground-truth pose. Test images do not publish ground truth, so this frame is from the validation split.验证集场景 1,第 1 帧,3 个实例。红色 3D 包围盒是预测位姿。绿色 3D 包围盒是真值位姿。测试图像不公开真值,所以这一帧来自验证集。

ITODD

ITODD (MVTec Industrial 3D Object Detection; Drost et al., ICCVW 2017) evaluates industrial object detection and pose recovery, including bin picking and inspection. Its 28 objects span different shapes, sizes, surface reflectance and symmetries; scenes include single objects and multiple instances in clutter.

For pose algorithms, thin or symmetric shapes can leave rotation poorly constrained, while reflective surfaces and occlusion make usable observations harder to obtain. Industrial manipulation also requires accurate 3D alignment rather than just a plausible 2D projection. The BOP version used here provides grayscale images and depth in test/; ground-truth poses are public for training and validation, but not for test images.

ITODD(MVTec Industrial 3D Object Detection;Drost 等,ICCVW 2017)用于评测工业物体检测与位姿求解,面向箱中抓取、物体检测等应用。它包含 28 个物体,覆盖不同形状、尺寸、表面反射特性和对称性,场景从单物体到杂乱环境中的多个实例。

对位姿算法而言,薄片或对称结构可能使旋转约束不足,反光表面与遮挡又增加了获取有效观测的难度。工业操作还要求准确的三维对齐,不能只看二维投影是否相似。本包使用 BOP 版本 test/ 中的灰度图像与深度;训练和验证图像公开真值位姿,测试图像的真值不公开。

ITODD validation scene 1 image 206, predicted pose in red and ground truth in green
ITODD validation image with predicted and reference poses.ITODD 验证图像中的预测位姿与参考位姿。

Validation scene 1, image 206, 2 instances of object 15. The red 3D bbox is the predicted pose. The green 3D bbox is the ground-truth pose. Test images do not publish ground truth, so this frame is from the validation split.

验证集场景 1,第 206 帧,物体 15 的 2 个实例。红色 3D 包围盒是预测位姿。绿色 3D 包围盒是真值位姿。测试图像不公开真值,所以这一帧来自验证集。

IC-BIN

IC-BIN (Doumanoglou et al., CVPR 2016) evaluates 6D object pose recovery in bin-picking scenes. Two object types from IC-MI appear repeatedly in different positions with heavy occlusion.

The key challenge is separating instances of the same category: similar appearance and overlapping surfaces can cause missing detections, duplicate predictions or a pose aligned to the wrong instance. Pose correction and scoring must retain the correspondence between each candidate and the particular visible object it represents. The images and meshes below come from the archives linked by the BOP dataset page; this package reads test/.

IC-BIN(Doumanoglou 等,CVPR 2016)用于评测箱中抓取场景的 6D 位姿求解。来自 IC-MI 的两类物体以多个实例出现在不同位置,并存在严重遮挡。

核心难点是区分同类别的不同实例:相似外观与重叠表面可能造成漏检、重复预测,或将姿态对齐到另一个实例。位姿修正和评分需要保持每个候选与其对应可见物体之间的关系。下面的图像与网格来自 BOP 数据集页链接的压缩包,本包读取 test/。

IC-BIN scene 1 image 26, predicted pose in red and ground truth in green
Scene 1, image 26, 14 instances of object 1. The red 3D bbox is the predicted pose. The green 3D bbox is the ground-truth pose.场景 1,第 26 帧,物体 1 的 14 个实例。红色 3D 包围盒是预测位姿。绿色 3D 包围盒是真值位姿。

ROBIROBI

ROBI(Yang 等,IROS 2021)面向反光零件的机器人箱中抓取,用于研究 6D 位姿求解与多视角深度融合。它使用 Ensenso N35 和 RealSense D415 记录七类物体、63 个场景。

主要难点是弱纹理、强反光、严重遮挡,以及重复实例的密集堆叠。反光会产生干扰边缘并降低深度测量质量,使某些候选姿态看似合理,却与可靠的几何观测不符。对 WAPR 而言,这些案例考察观测不完整或失真时的位姿修正与候选评分。随包 Zigzag 帧提供示例 10 的输入,其他记录案例使用另外获取的帧。

ROBI (Yang et al., IROS 2021) studies 6D object pose recovery and multi-view depth fusion for robotic bin picking of reflective parts. It records seven object types across 63 scenes using Ensenso N35 and RealSense D415 cameras.

Its main challenges are weak texture, strong reflections, severe occlusion and dense piles of repeated instances. Reflections can create misleading image edges and degrade depth measurements, so a candidate pose may appear plausible while disagreeing with reliable geometry. For WAPR, these cases examine pose correction and candidate scoring under incomplete or corrupted observations. The bundled Zigzag frame supplies the input for example 10; other recorded cases use separately obtained frames.

ROBI Zigzag scene 4 view 0, Ensenso left image
Zigzag, scene 4, view 0. The Ensenso left image, 1280×1024.Zigzag,场景 4,第 0 帧。Ensenso 左图,1280×1024。

Bundled ROBI sample frame随包提供的 ROBI 示例帧

样例相机参数为 fx = fy = 1083.097046、cx = 379.326874、cy = 509.437195。CAD_readme.txt 记录 Zigzag.stl 是依据 McMaster-Carr 零件在 SolidWorks 中建立的模型。该样例目录随包单独提供,python -m wapr.download_assets 不下载它。

The sample camera has fx = fy = 1083.097046, cx = 379.326874 and cy = 509.437195. CAD_readme.txt records that Zigzag.stl was modeled in SolidWorks from a McMaster-Carr part. This sample folder is bundled separately and is not fetched by python -m wapr.download_assets.

Full set完整集

The links below are the ones on the dataset page. The Zigzag scene archive is about 3.7 GB. The example reads the one frame above and does not download that archive.

下面的链接就是数据集页上的那些。Zigzag 的场景包大约 3.7 GB。示例读的是上面这一帧,不会下载该数据包。

TACOTACO

TACO[8] (Liu et al., CVPR 2024) records bimanual tool–object manipulation using multi-view sensing and optical motion capture. Its release includes egocentric RGB-D, camera parameters, object models and hand–object pose annotations. We use it to demonstrate 6D tracking of a lint roller and a wooden box that move independently during manipulation and hand occlusion. Dataset background is available on the official project page, with download and file-organization instructions in the official repository.

For 6D tracking, hands can hide the object surface while the tool and manipulated object move independently. The challenge is to maintain a separate pose for each target through contact and occlusion without confusing their regions. The WAPR case uses these observations to study multiple-object pose updates.

TACO[8](Liu 等,CVPR 2024)通过多视角采集与光学动捕记录双手操作工具及物体的过程,提供第一人称 RGB-D、相机参数、物体模型以及手物位姿标注。我们使用其中粘毛滚筒与木盒的序列,展示操作和手部遮挡过程中两个独立运动物体的 6D 位姿跟踪。数据集介绍见官方项目页,下载及文件组织方式见官方仓库。

对 6D 位姿跟踪而言,手部会遮挡物体表面,工具与被操作物体又会各自运动。难点是在接触与遮挡过程中持续维护每个目标的独立位姿,避免混淆不同目标的区域。WAPR 案例使用这些观测研究多物体的位姿更新。

The local excerpt comes from (brush, roller, box)/20231104_007; its acquisition source is the Hugging Face data repository. It retains 114 RGB-D frames, camera parameters, two meshes and motion-capture pose arrays. The excerpt does not include the full release or verified segmentation annotations; the case uses SAM 2[1] to prepare visible regions.

本地摘录选自 (brush, roller, box)/20231104_007,实际获取来源为 Hugging Face 数据仓库,保留 114 帧 RGB-D、相机参数、两份网格及动捕位姿数组。摘录不包含完整发布内容或已核验的分割标注;案例使用 SAM 2[1] 准备可见区域。

Sample field样例信息Value内容
Targets目标物体Lint roller and wooden box粘毛滚筒与木盒
RGB sequenceRGB 序列114 frames · 512×376 · 30 fps114 帧 · 512×376 · 30 fps
Depth conversion深度换算Stored 16-bit values / 4000 → meters; missing depth stays invalid存储的 16 位数值除以 4000 得到米;缺失深度保留为无效值
Reference poses参考位姿tool_076.npy · target_093.npy · one 4×4 transform per object per frame · 每物体每帧一个 4×4 变换
Three sampled frames from the TACO manipulation sequence.TACO 操作序列中的三个采样帧。

TACO excerpt (brush, roller, box)/20231104_007: three sampled frames show the wooden box and lint roller during manipulation. Labels identify the two targets. The RGB-D inputs, object meshes and motion-capture comparisons are described in the TACO case.

TACO 摘录 (brush, roller, box)/20231104_007:三个采样时刻展示操作过程中的木盒与粘毛滚筒,图中文字标出两个目标。RGB-D 输入、物体网格与动捕对照见 TACO 案例。

WAPR and motion-capture poses on the same frame.同一帧的 WAPR 位姿与动捕位姿对照。

Saved pose results on the same frame: WAPR estimates at left and TACO motion-capture references at right. Orange identifies the wooden box; blue shows the estimated roller contour and green its reference contour. This is a visual alignment comparison. The TACO application case provides the looping video, initialization procedure and detailed evaluation conditions.

同一帧的已保存位姿结果:左侧为 WAPR 估计,右侧为 TACO 动捕参考。木盒以橙色区分;滚筒的估计轮廓为蓝色,参考轮廓为绿色。这里展示投影对齐效果;TACO 应用案例提供循环视频、初始化流程及详细评价条件。

YCBInEOAT tracking sequencesYCBInEOAT 跟踪序列

YCBInEOAT evaluates RGB-D based 6D pose tracking during real robot manipulation. It contains five YCB objects manipulated with three types of end effector, with frame-by-frame pose annotations. Each sequence provides RGB-D, camera intrinsics, an object mesh and one annotated target, separately from the BOP YCB-V frames.

The tracking challenges include manipulation-induced occlusion, changing viewpoints and accumulated pose drift. For WAPR, the examples examine pose updates as the object moves and recovery when the current estimate loses alignment. Download each sequence from the official archive; python -m wapr.download_assets does not fetch it. The examples use the three sequences below; protocols, recorded comparisons and initialization conditions are described in the YCBInEOAT case.

YCBInEOAT 用于评测真实机器人操作过程中的 RGB-D 6D 位姿跟踪,包含五个 YCB 物体和三种末端执行器,并逐帧标注物体位姿。每段序列提供 RGB-D、相机内参、物体网格与一个目标的标注,和 BOP YCB-V 帧相互独立。

跟踪难点包括操作过程中的遮挡、观察视角变化,以及连续更新时累积的位姿漂移。WAPR 示例考察物体运动中的位姿更新,以及当前估计失去对齐后的恢复。序列需从官方下载页面分别下载,python -m wapr.download_assets 不获取它们。示例使用下表三段序列,协议、对照结果及初始化条件见 YCBInEOAT 案例。

Sequence序列Target目标Download下载
cracker_box_reorientCracker box饼干盒cracker_box_reorient.tar.gz
mustard_easy_00_02Mustard bottle芥末瓶mustard_easy_00_02.tar.gz
sugar_box1Sugar box糖盒sugar_box1.tar.gz
Saved tracking frame from YCBInEOAT cracker_box_reorient, showing the cracker box
Cracker box · cracker_box_reorient饼干盒 · cracker_box_reorient
Saved tracking frame from YCBInEOAT mustard_easy_00_02, showing the mustard bottle
Mustard bottle · mustard_easy_00_02芥末瓶 · mustard_easy_00_02
Saved recovery frame from YCBInEOAT sugar_box1, showing the sugar-box sequence
Sugar box · sugar_box1糖盒 · sugar_box1

Saved tracking stills from the three sequences above, with supplied frame-0 mask initialization. Red is the estimated pose; green is the annotated pose. The sugar-box panel records a recovery frame. The overlaid timing labels are recorded stage timings; initialization and tracking protocols are explained in the YCBInEOAT case.上表三段序列的已保存跟踪静帧,采用给定首帧掩码初始化。红色是估计位姿,绿色是标注位姿;糖盒面板展示一次补救帧。顶栏显示实测阶段耗时,初始化与跟踪协议见 YCBInEOAT 案例。

SA6D · Coming soonSA6D · 稍后放出

SA6D is the synthetic RGB-D dataset used to train all four pose foundation-model checkpoints in this release: masked WAPR, unmasked WAPR, SAPR, and WBPS. A wide-angle pair on a symmetric object can be far apart in rotation and still look almost the same. If that raw rotation is the training target, two similar pictures receive two different updates. SA6D keeps a rotational symmetry prior with each object and uses it to choose one canonical target before training.

SA6D 是本发布包中四份位姿基础模型权重的合成 RGB-D 训练集:带掩码 WAPR、不带掩码 WAPR、SAPR 和 WBPS 均基于它训练。对称物体上的一对广角样本,旋转可以差很多,看起来却几乎一样。如果把这份原始旋转当成训练目标,两张相近的图会得到两个不同的更新。SA6D 给每个物体留一份旋转对称先验,训练前用它选定一个规范目标。

Two views of a symmetric plate with a large rotation gap and the same canonical target
Why the prior is there. The two plates differ by a large rotation. The visible change is a small piece of texture. The prior treats them as the same pose, so training uses one update.先验为什么要有。两只盘子的旋转差很大,看得见的差别只是一小块纹理。先验把它们当成同一个位姿,训练就只用一个更新。

The set starts from 944 GSO[17] scans. For each scan, annotators set the rotational symmetry type and order, and KASAL places the axis and the center. Geometry scaling and a texture change then make a new instance without a new annotation. One change keeps the texture, so the prior stays texture-aware, as with the plate. The other paints the object a flat color, so only the geometric prior remains, as with the carton. BlenderProc[18] renders these instances into synthetic RGB-D scenes. The paper reports about 50K instances, about 30K of them rotationally symmetric, and about 2M RGB-D images.

数据从 944 个 GSO[17] 扫描出发。每个扫描由标注给出旋转对称的类型和阶数,KASAL 定出轴和中心。接着做几何缩放和纹理改动,得到新实例,不必重新标注。一种改动保留纹理,先验仍看纹理,盘子就是这样。另一种把物体涂成纯色,只留下几何先验,纸盒就是这样。BlenderProc[18] 再把这些实例渲染成合成 RGB-D 场景。论文给出的规模大约是 5 万个实例,其中约 3 万个带旋转对称,以及约 200 万张 RGB-D。

SA6D construction: original scans, symmetry axes, augmented models, and rendered training scenes
How a training image is made. Left: the original scan, then the augmented object with its symmetry axis. Right: those objects rendered into a training scene.一张训练图是怎么来的。左边是原始扫描,然后是带对称轴的扩充物体。右边是这些物体渲染进训练场景。

This dataset is still being organized. It will be released here once that is done.

这个数据集还在整理,整理好了会放出来。

Pose model weights位姿模型权重

The four pose checkpoints are prepared for release in SEU-WYL/WAPR. Their roles, file locations and download commands are listed below.

四份位姿权重的发布目标为 SEU-WYL/WAPR。下方集中说明各权重的用途、文件位置及下载命令。

Four pose checkpoints四份位姿权重

Three checkpoints update a pose. The fourth scores the hypotheses. The repository runs WAPR for 3 updates, then SAPR for 2. One update is one rotation step and one translation step. The counts are wapr_iters and sapr_iters in wapr/recipe.py.

前三个模型用于修正位姿,第四个模型用于给候选姿态评分。仓库先用 WAPR 更新 3 次,再用 SAPR 更新 2 次。一次更新包含一步旋转和一步平移。次数是 wapr/recipe.py 里的 wapr_iters 和 sapr_iters。

WAPR

WAPR (Wide-Angle Pose Refinement) was trained with rotation errors up to about 90°. The wide-angle examples include successful correction from 100° and 120° on individual objects; these demonstrations do not establish a general accuracy guarantee beyond the training range.

WAPR(Wide-Angle Pose Refinement,广角位姿修正)的训练覆盖约 0°–90° 的旋转误差。广角修正展示包含个别物体从 100° 和 120° 成功修正的案例,这些结果不能视为超过训练范围后的普遍精度保证。

Each step applies tanh to the three components of the predicted rotation vector, then multiplies each component by wapr_max_rot_rad = 0.70. This bounds each component to 0.70 rad (about 40°), not the total rotation angle. The vector norm, which sets the rotation angle, is at most √3 × 0.70 rad (about 69.5°) for one step.

每一步对预测旋转向量的三个分量分别施加 tanh,再分别乘以 wapr_max_rot_rad = 0.70。因此每个分量的幅值不超过 0.70 弧度(约 40°),这不是总旋转角度的上限。决定旋转角度的向量模长单步最多为 √3 × 0.70 弧度(约 69.5°)。

With a mask带 mask

wapr_w_mask receives seven channels: RGB, a target mask and xyz. The region can be supplied or predicted by a segmentor such as SAM, FastSAM or Mask R-CNN[15]. Its training uses SAM/FastSAM masks restricted to the input boxes. Mask errors can affect rotation and translation updates; check whether the region contains background or misses the visible target.

wapr_w_mask 接收七通道输入:RGB、目标掩码与 xyz。区域可以外部提供,也可以由 SAM、FastSAM、Mask R-CNN[15] 等分割器预测;训练使用限制在输入框内的 SAM/FastSAM 掩码。掩码误差可能影响旋转与平移更新,应检查区域是否包含背景或遗漏可见物体。

The mask keeps the wide step on the object it covers. That is why this checkpoint is the one to pass when a mask is available.

目标掩码为大角度位姿修正提供目标区域信息。所以有 mask 时用的是这一份权重。

Without a mask不带 mask

wapr_wo_mask drops the mask channel: 6 channels, RGB and xyz. The rest of the training procedure matches the masked model. A call with a mask loads wapr_w_mask. A call without one loads wapr_wo_mask and fills the region from the bbox.

wapr_wo_mask 去掉 mask 通道:6 通道,RGB 和 xyz。除此之外,训练流程和带 mask 的那份一致。调用时带了 mask 就加载 wapr_w_mask。没有 mask 就加载 wapr_wo_mask,并用包围盒填区域。

Use the masked checkpoint for mask-guided pose initialization, correction, localization, and detection. For tracking comparisons with FoundationPose[2], use wapr_wo_mask for subsequent pose updates. A 2D region may still initialize translation at a large frame stride; this does not add a mask channel to the tracking network. Frame-0 registration is a separate initialization step and can use the same supplied mask for both methods.

有可见 mask 引导的姿态初始化、修正、定位和检测使用带 mask 的权重;与 FoundationPose[2] 比较时,后续跟踪更新使用 wapr_wo_mask。大步长更新可以用二维区域确定起始平移,但不会因此给跟踪网络增加 mask 通道。第 0 帧求位姿是单独的初始化步骤,两种方法可以共用同一块给定 mask。

The extra mask channel supplies an instance-specific region cue when several objects share a CAD identity. It does not guarantee correct association; pose accuracy and instance coverage still require evaluation on the target data.

同一 CAD 对应多个实例时,附加掩码通道提供实例区域线索,但不保证关联正确;位姿精度与实例覆盖仍需在目标数据上评估。

SAPR

SAPR is Small-Angle Pose Refinement. It uses the same update as WAPR, with half the per-component rotation bound: sapr_max_rot_rad = 0.35 radians (about 20°) per component. The total angle of one step is at most √3 × 0.35 rad (about 34.7°). Training covered errors from 0° to about 30°. It is the update for a residual of about 20° to 30°, after the three WAPR steps have taken the large turns. The input is 6 channels: RGB and xyz.

SAPR 是小角位姿修正(Small-Angle Pose Refinement)。它和 WAPR 使用同一种更新,每个旋转向量分量的上限减半:sapr_max_rot_rad = 0.35 弧度(约 20°)。单步总旋转角度最多为 √3 × 0.35 弧度(约 34.7°)。训练覆盖的误差大约是 0° 到 30°。它用来修正大约 20° 到 30° 的残余角度偏差,排在三次 WAPR 之后。输入是 6 通道:RGB 和 xyz。

WBPS

WBPS is Within-and-Between Pose Score. Two heads read the same hypotheses. within_group selects a pose inside each group. between_group supplies scores from which the caller computes one comparable score per group. The network processes each group of hypotheses separately; comparison across groups happens on the returned scores. W is within, B is between, and PS is pose score.

WBPS 是组内与组间位姿分数(Within-and-Between Pose Score)。两个分支读同一批候选姿态。within_group 在每个姿态组内选择位姿。调用方根据 between_group 的输出为每组计算一个可比较的分数。网络分别处理各组候选姿态,跨组比较发生在返回分数上。W 是 within,B 是 between,PS 是 pose score,也就是姿态分数。

The default group is n_view × n_inplane. Four view directions and three in-plane rotations make 12 poses. within_group scores those 12. A larger value is better, and the estimate keeps the argmax. That choice is only the selection inside the group.

默认的一组是 n_view × n_inplane。4 个视角、每个视角 3 次面内旋转,就是 12 个姿态。within_group 给这 12 个打分。越大越好,估计结果保留 argmax。这个选择只发生在组内。

Four such groups contain 48 poses. within_group selects the best of 12 in each group. The between_group head returns 12 values per group, without mixing groups in its attention. The caller reduces each row to (100 - max(between_group)) / 200; a larger returned value ranks higher. In this package, each detected object instance is one group. Several instances estimated together produce several group scores that can be compared or filtered. The input is 6 channels: RGB and xyz.

四个这样的组包含 48 个姿态。within_group 在每组 12 个候选姿态中选出最好的一个。between_group 分支对每组返回 12 个值,注意力计算不混合不同组。调用方先取每组输出的最大值,再按 (100 - max(between_group)) / 200 计算该组分数;返回值越大,排序越靠前。本包中,每个被检测到的物体实例是一组。一起估计多个实例时,会得到多个可比较或过滤的组分数。输入是 6 通道:RGB 和 xyz。

The four files四个文件

File文件 Input输入
assets/weights/wapr_w_mask.pth7 channels: RGB, mask, xyz7 通道:RGB、mask、xyz
assets/weights/wapr_wo_mask.pth6 channels: RGB, xyz6 通道:RGB、xyz
assets/weights/sapr.pth6 channels: RGB, xyz6 通道:RGB、xyz
assets/weights/wbps.pth6 channels: RGB, xyz6 通道:RGB、xyz

The source package is GNU LGPL-2.1-only. The four first-party checkpoints use CC BY-ND 4.0; see the scope and attribution notice and full license. Commercial use is free; original redistribution requires attribution, and adapted weights may not be shared under this license. Enterprise versions and customization: Yulin Wang, shopedataset@gmail.com. Third-party resources retain their own terms. Serialized .onnx and .engine files are not supplied with the released checkpoints. Follow the installation instructions to build the TensorRT[3] FP16 engines on the target GPU.

源码包采用 GNU LGPL-2.1-only。四份第一方权重采用 CC BY-ND 4.0,见适用范围与署名声明及完整许可。免费允许商用;原样再发行须署名,本许可不允许共享修改版权重。企业版本与定制需求联系 Yulin Wang:shopedataset@gmail.com。第三方资源仍遵守各自条款。发布的权重不附带序列化 .onnx 与 .engine 文件。请按安装说明在目标 GPU 上构建 TensorRT[3] FP16 引擎。

References and licenses参考文献与许可

  1. SAM 2 — Apache-2.0; cctorch: BSD-3-Clause. Native source and official checkpoints; cctorch carries an additional BSD notice. Ultralytics-converted files require checking their distributor terms.原生源码与官方权重;cctorch 另附 BSD 声明。Ultralytics 转换文件还需核对分发方条款。 · GitHubGitHub · License/notice 1许可/声明 1 · License/notice 2许可/声明 2
    Ravi et al. SAM 2: Segment Anything in Images and Videos. arXiv:2408.00714, 2024. · Paper论文 ↩ ↩
  2. FoundationPose — NVIDIA custom source license. Comparison method source has a custom license; do not describe it as MIT or presume weights share the same grant.对照方法源码使用自定义许可;不能标为 MIT,也不能推定权重有相同授权。 · GitHubGitHub · License/notice 1许可/声明 1
    Wen et al. FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects. CVPR 2024. · Paper论文 ↩ ↩
  3. tensorrt-cu12 — Other/Proprietary License. Runtime dependency; dependencies and model/data assets retain their own licenses.运行依赖;依赖及模型/数据资源保留各自许可。 · GitHubGitHub · License/notice 1许可/声明 1 ↩ ↩
  4. LM-O data — CC-BY-SA-4.0. Data and model excerpt, not the loader code.数据与模型摘录,不是读取代码。 · Original source原始来源 · GitHub: BOP toolkitGitHub:BOP 工具集
    Brachmann et al. Learning 6D Object Pose Estimation Using 3D Object Coordinates. ECCV 2014. · Paper论文 ↩ ↩
  5. T-LESS data — CC-BY-4.0. Data and object models; attribute the original dataset.数据与物体模型;须标注原始数据集。 · Original source原始来源
    Hodaň et al. T-LESS: An RGB-D Dataset for 6D Pose Estimation of Texture-less Objects. WACV 2017. · Paper论文 ↩ ↩
  6. TUD-L data — CC-BY-SA-4.0. Dataset terms are independent of WAPR source.数据许可独立于 WAPR 源码。 · Original source原始来源
    Hodaň et al. BOP: Benchmark for 6D Object Pose Estimation. ECCV 2018. · Paper论文 ↩ ↩
  7. ITODD data — CC-BY-NC-SA-4.0. Non-commercial dataset; WAPR permissions do not replace dataset terms.非商业数据集;WAPR 授权不替代数据条款。 · Original source原始来源
    Drost et al. Introducing MVTec ITODD — A Dataset for 3D Object Recognition in Industry. ICCV Workshops 2017. · Paper论文 ↩ ↩
  8. TACO data — Not separately verified / 未单独核实. Data terms have not been separately verified. The official repository is the original source; the referenced Hugging Face acquisition mirror has no explicit license field. Confirm the data owner's terms for redistribution or commercial use.数据条款尚未单独核实。官方仓库为原始出处;实际获取数据的 Hugging Face 镜像数据卡未声明明确许可字段。再分发或商业使用须核实数据权利方的条款。 · GitHubGitHub
    Liu et al. TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object Understanding. CVPR 2024. · Paper论文 ↩ ↩
  9. YCBInEOAT data — Not separately verified / 未单独核实. Tracking data permission must be checked at the original archive; the comparison code license does not establish the data license.跟踪数据许可须在原始归档核实;对照代码的许可不能作为数据许可。 · GitHubGitHub · GitHubGitHub
    Wen et al. se(3)-TrackNet: Data-driven 6D Pose Tracking by Calibrating Image Residuals in Synthetic Domains. IROS 2020. Introduces YCBInEOAT. · Paper论文 ↩ ↩
  10. ROBI data — Not separately verified / 未单独核实. The saved public poses and dataset are credited to ROBI; an independent grant to redistribute them has not been verified.保存的公开位姿和数据均注明 ROBI 来源;其再分发授权尚未独立核实。 · GitHubGitHub
    Yang et al. ROBI: A Multi-View Dataset for Reflective Objects in Robotic Bin-Picking. IROS 2021. · Paper论文 ↩ ↩
  11. IC-BIN · Official BOP dataset pageBOP 官方数据页
    Doumanoglou et al. Recovering 6D Object Pose and Predicting Next-Best-View in the Crowd. CVPR 2016. · Paper论文 ↩ ↩
  12. YCB-Video (YCB-V) · Official BOP dataset pageBOP 官方数据页 · GitHub: PoseCNNGitHub:PoseCNN
    Xiang et al. PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes. RSS 2018. · Paper论文 ↩ ↩
  13. HomebrewedDB (HB) · Official BOP dataset pageBOP 官方数据页
    Kaskman et al. HomebrewedDB: RGB-D Dataset for 6D Pose Estimation of 3D Objects. ICCV Workshops 2019. · Paper论文 ↩ ↩
  14. FastSAM · GitHubGitHub
    Zhao et al. Fast Segment Anything. arXiv:2306.12156, 2023. · Paper论文 ↩ ↩
  15. Mask R-CNN · GitHubGitHub
    He et al. Mask R-CNN. ICCV 2017. · Paper论文 ↩ ↩
  16. CNOS · GitHubGitHub
    Nguyen et al. CNOS: A Strong Baseline for CAD-based Novel Object Segmentation. ICCV Workshops 2023. · Paper论文 ↩ ↩
  17. Google Scanned Objects (GSO) · Official dataset introduction官方数据集介绍
    Downs et al. Google Scanned Objects: A High-Quality Dataset of 3D Scanned Household Items. arXiv:2204.11918, 2022. · Paper论文 ↩ ↩
  18. BlenderProc · GitHubGitHub
    Denninger et al. BlenderProc2: A Procedural Pipeline for Photorealistic Rendering. Journal of Open Source Software 8(82):4901, 2023. · Paper论文 ↩ ↩

Model weights and dataset/task assets may have separate terms. WAPR's source license does not replace them.模型权重、数据集与任务资源可能有独立条款;WAPR 源码许可不替代这些许可。