Reconstruction and cross-scene pose重建与跨场景位姿
This chapter shows how to build a model of an object without an existing CAD mesh and use it for 6D pose estimation. Starting from an RGB-D image and camera intrinsics, select the object with a pixel click or a sentence. Example 11 reconstructs a mesh, determines its metric dimensions, and estimates the object's position and orientation in that image. It saves the mesh, a matching 4 × 4 pose matrix, and a report.
本章介绍如何为没有现成 CAD 网格的物体建立模型,并将它用于 6D 位姿估计。从一张 RGB-D 图像及相机内参出发,用点击点或一句话指定目标。示例 11 重建网格、确定米制尺寸,并估计物体在这张图上的位置与朝向,保存网格、与之配套的 4 × 4 位姿矩阵和报告。
The saved mesh can then be reused in another scene. Example 12 takes this mesh, a new RGB-D image, and its camera intrinsics; it selects the highest-scoring 2D detection and estimates the object's 6D pose, saving the pose and an overlay. Reconstruction and cross-scene estimation are separate runs, so the same model can be used on new images.
保存的网格可以继续用于其他场景。示例 12 输入这份网格、另一张 RGB-D 图像及其相机内参,选取得分最高的 2D 检测实例,估计物体的 6D 位姿,并保存位姿与叠图。重建与跨场景估计分别运行,同一份模型可以用于新的图像。
| Case案例 | Input输入 | Output输出 |
|---|---|---|
| 11 · Single-frame reconstruction11 · 单帧重建与对齐 | RGB-D, K, a point or sentenceRGB-D、内参、点或句子 | Mesh, source-image pose, report网格、源图位姿、报告 |
| 12 · Cross-scene pose12 · 跨场景位姿估计 | Saved mesh, another RGB-D, K已保存的网格、另一张 RGB-D、内参 | New-image pose and overlay新图位姿与叠图 |
For each RGB-D input, keep the RGB image and depth on the same pixel grid. read_rgb_depth converts source depth to metres; cam_K.txt contains a 3 × 3 pixel-space camera matrix. The SAM 2[2] mask must also have the RGB height and width. Confirm these inputs before running a full UV bake: the bake changes appearance, while mask and depth errors can change the reconstructed shape and its metric scale.
每组 RGB-D 输入须保证彩色图和深度图处于同一像素网格。read_rgb_depth 将原始深度换算为米;cam_K.txt 保存像素坐标系下的 3 × 3 相机内参。SAM 2[2] 掩码也必须与 RGB 图同高同宽。运行耗时的 UV 烘焙前应先核对这些输入:烘焙改变外观,而掩码和深度错误还会影响重建形状及米制尺寸。
Case 11 single-frame reconstruction and alignment案例 11 单帧重建与对齐
Set OBJECT_NAME, FRAME_RGB, FRAME_DEPTH and FRAME_K in examples/11_reconstruct_object.py. The name controls the output directory; changing it does not change the input images. Depth uses the explicit DEPTH_UNIT_M conversion, and K is a 3 × 3 pixel-space matrix. The default input is the first cracker-box frame with a saved integer pixel click.
在 examples/11_reconstruct_object.py 中设置 OBJECT_NAME、FRAME_RGB、FRAME_DEPTH 与 FRAME_K。物体名决定输出目录;只修改名称不会更换输入图像。深度由显式的 DEPTH_UNIT_M 换算为米,内参是 3 × 3 像素矩阵。默认输入为饼干盒首帧及保存的整数像素点击点。
python examples/11_reconstruct_object.py
Select and segment the object选择并分割目标
With SENTENCE = "", SAM 2 uses CLICK_UV, a saved pixel (u, v). With a sentence, Qwen2.5-VL[6] predicts a box and SAM 2 segments within it. No annotated pose is projected to obtain the click, and saved overlay images are not mask inputs. Both prompt routes continue through the same reconstruction stages.
SENTENCE = "" 时,SAM 2 使用保存的像素点 CLICK_UV,坐标为 (u, v)。设置句子时,千问 2.5-VL 预测目标框,再由 SAM 2 分割。程序不从标注位姿投影生成点击点,也不从保存的叠加图提取掩码。两种提示随后执行同一套重建阶段。
Check the mask检查掩码
The predicted mask must be nonempty and match the RGB-D grid. Inspect it when a prompt selects the wrong region or includes the background; the mask affects reconstruction, scale selection, and pose initialization.
预测掩码必须非空,并与 RGB-D 图像尺寸一致。若提示选错目标或包含背景,应先检查掩码;它会影响重建、尺度选择与位姿初始化。
Reconstruct the mesh重建网格
SIZE_PATH_POINT = "unipose" retains full UV baking and UniPose9D[4] dimensions. SIZE_PATH_SENTENCE = "outline" retains the vertex-color mesh and observation-based outline sizing. SIZE_PATH selects the active recipe visibly; either prompt can use either size route when you change that variable. SAM 3D[3] is released before loading the pose estimator.
SIZE_PATH_POINT = "unipose" 保留完整 UV 烘焙与 UniPose9D[4] 尺寸估计;SIZE_PATH_SENTENCE = "outline" 保留顶点色网格与观测轮廓定标。实际路径由可见变量 SIZE_PATH 决定;修改它后,两种提示都可以使用任一尺寸路径。加载位姿估计器前释放 SAM 3D[3]。
Set metric dimensions and select the source-image pose确定米制尺寸并选择源图位姿
First compare isotropic scale trials with batched WAPR, SAPR and WBPS. Then constrain dimensions with the predicted UniPose9D box or the visible mask outline. Refine a batch of pose hypotheses for the sized mesh; WBPS supplies geometric scores and DINOv2[1] compares the actual mesh appearance with the photograph. The selected pose is the object-to-camera transform saved with the output mesh.
先将各向同性尺度候选批量计算送入 WAPR、SAPR 与 WBPS,再根据预测的 UniPose9D 包围盒或可见掩码轮廓校正尺寸。对定标后的网格批量计算修正位姿候选,WBPS 提供几何评分,DINOv2[1] 比较网格实际外观与照片。最终选择的物体到相机位姿与输出网格配套保存。
Update shape with RoMa用 RoMa 更新形状
APPLY_ROMA = True runs observation-only shape fitting. Each round matches a fresh render to the same photograph and fits three axis scales while keeping the selected pose fixed. The solver does not fit an unapplied translation. Scale bounds, match count and stopping choices remain visible. An update is retained only when its applied deformation improves correspondence error and the new render passes fresh-match checks; otherwise the previous mesh is kept and the report records why it stopped.
APPLY_ROMA = True 启用仅用观测的形状拟合。每轮将当前网格的新渲染与同一张照片匹配,在选定位姿固定的情况下拟合三轴缩放。求解器不再拟合一个未实际应用的平移量。缩放范围、匹配数量与停止条件均显式保留。只有实际变形降低对应误差、且新渲染通过重新匹配检查时才保留更新;否则保留上一轮网格,并在报告中记录停止原因。
Save a matched mesh and pose保存配套的网格与位姿
outputs/reconstruct_object/<object>/
report.json
prediction/
mesh.bin, mesh.json, mesh.jpg # JPEG only for UV / UV 网格才有 JPEG
pose.npy # 4 × 4 object-to-camera / 物体到相机
mask.png, overlay.png
provenance.json
Mesh vertices and pose translation are meters. The original reconstructed-object coordinate frame is retained; there is no reference-CAD alignment, printed-face turn or later-frame shape fit. The report records prompt, input hashes, size path, selected pose and RoMa[5] decisions. The overlay uses the serialized mesh and saved pose. Neither entry writes to pages/demo/. The selected-pose scores describe the mesh before RoMa; the report labels that stage, and RoMa records its own correspondence errors.
网格顶点与位姿平移的单位均为米,保留原重建物体坐标系;不执行参考 CAD 坐标对齐、印刷面翻转或后续帧形状拟合。报告记录提示、输入摘要、尺寸路径、所选位姿与 RoMa[5] 决策。叠图使用实际序列化的网格和保存的位姿。两个入口都不写入 pages/demo/。 所选位姿的评分对应 RoMa 更新前的网格,报告明确记录该阶段;RoMa 另行记录对应关系误差。
Case 12 cross-scene pose estimation案例 12 跨场景位姿估计
examples/12_cross_scene_pose.py reads example 11's independent prediction bundle. Set its mesh, new RGB-D and camera file paths. Its supplied target scene is YCB-V[8] scene 50, frame 1130, for the mustard bottle. Prepare the mustard bundle with example 11 before running that recipe: set the object name to mustard, the RGB-D to the first frame of mustard0, K to that sequence's cam_K.txt, and a saved point or SENTENCE = "黄色的瓶子". Choose SIZE_PATH explicitly.
examples/12_cross_scene_pose.py 读取示例 11 的独立预测包,输入路径指定网格、新场景 RGB-D 与相机文件。所附配置使用 YCB-V[8] 场景 50、第 1130 帧中的芥末瓶。运行前,先用示例 11 生成芥末瓶预测包:物体名设为 mustard,RGB-D 指向 mustard0 首帧,内参指向该序列的 cam_K.txt,并指定保存的点击点或 SENTENCE = "黄色的瓶子";尺寸路径由 SIZE_PATH 显式选择。
python examples/12_cross_scene_pose.py
Templates are built from the saved mesh and rebuilt when its geometry or texture changes. The detector selects the instance with the highest score_2d. That predicted mask initializes one batch of 6D hypotheses; WAPR, SAPR, WBPS and DINOv2 choose the new-image pose without repeating the pose pipeline for the same instance. This case does not segment from a prompt, reconstruct a mesh, or update its shape.
模板由保存的网格生成;网格几何或贴图变化时重建模板库。检测器选择 score_2d 最高的实例,使用其预测掩码初始化一批 6D 位姿候选,再通过 WAPR、SAPR、WBPS 与 DINOv2 选择新图位姿;同一实例不会重复执行位姿流程。此案例不从提示分割,不重新建模,也不更新形状。
Outputs are outputs/cross_scene_pose/<object>/pose.npy, mask.png, overlay.png and report.json. An empty detection set records no_detection. REFERENCE_MASK_PATH = None keeps evaluation optional; when supplied, that mask is read only after prediction files are saved, and its IoUs go to a separate evaluation.json. IoU never chooses the instance or pose.
输出为 outputs/cross_scene_pose/<物体>/pose.npy、mask.png、overlay.png 与 report.json。没有检测时记录 no_detection。REFERENCE_MASK_PATH = None 表示不要求评测掩码;指定后也只在预测文件保存后读取,将 IoU 写入独立的 evaluation.json。IoU 不参与选择实例或位姿。
- Load the independent mesh and provenance; read the new RGB-D and convert depth to meters using the camera scale.读取独立网格及来源记录;加载新 RGB-D,按相机深度系数换算为米。
- Hash the geometry and texture for template reuse; warm the detector and decode the highest-scoring predicted mask.以几何和纹理摘要判断模板是否可复用;预热检测器,解码最高分预测掩码。
- Initialize the 4 × 3 pose group with sensor depth; visibly call masked WAPR, SAPR and WBPS on the whole group.用传感器深度初始化 4 × 3 位姿组;依次显式调用带 mask 的 WAPR、SAPR 与 WBPS,对整组计算。
- Keep positive within-group scores, or the full group if none pass; DINOv2 ranks rendered appearance. Save pose, mask, overlay and relative input paths before optional evaluation.保留组内评分为正的候选;若均未通过则保留全组,由 DINOv2 比较渲染外观。先保存位姿、掩码、叠图及相对输入路径,再进行可选评测。
Reconstructed mesh and new-scene pose重建网格与新场景位姿结果
These saved figures show the same cross-scene experiment as the project page. The source is the first YCBInEOAT[7] mustard0 frame; the target is YCB-V scene 50, frame 1130, with a different camera, lighting and background. The saved mesh was constrained by a predicted box. The following result details specify the candidate filtering and pose selection used in this comparison.
这些保存图片对应宣传页的同一组跨场景实验。源图为 YCBInEOAT[7] 的 mustard0 首帧,目标图为 YCB-V 场景 50、第 1130 帧,相机、光照与背景不同。保存的网格经预测包围盒约束定标;下文给出本对照的候选筛选与位姿选择设置。
The saved result record contains 27 detector candidates, with four scores at or above 0.50. The highest score_2d, 0.767, selects the mustard region. WAPR refines twelve pose hypotheses for it; positive within_group candidates remain and DINOv2 selects the render closest to the photograph. Its returned score_6d is 0.863, within_group 63.7 and DINOv2 cosine 0.775. Offline mask IoUs are 0.941 for detection and 0.885 for the projected pose; neither IoU selects the region or pose. This one frame does not measure performance over the complete test set.
保存的结果记录包含 27 个检测候选,其中四个分数不低于 0.50。最高的 score_2d 为 0.767,据此选择芥末瓶区域。WAPR 对其修正十二个候选姿态,保留 within_group 为正的候选,再由 DINOv2 选择与照片最接近的渲染。返回的 score_6d 为 0.863、within_group 为 63.7、DINOv2 余弦为 0.775。事后计算的掩码 IoU 分别为检测 0.941、位姿投影 0.885,均不参与选择区域或位姿。该单帧结果不代表完整测试集上的性能。
Reconstruction results and optional diagnostics重建结果与可选诊断
Cracker-box mesh and pose饼干盒网格与位姿
The same saved cracker-box mesh shown on the project page is included here. Before, After and Ground truth display the mesh before metric sizing, the saved sized reconstruction and the reference CAD. Reference geometry is used for this offline comparison, not to choose reconstructed dimensions. Its sorted sides are 80.3 × 164.0 × 218.1 mm, compared with 71.7 × 164.0 × 213.5 mm for the reference; the checked surface distance is 4.7 mm.
这里直接展示宣传页的同一份饼干盒网格。对齐前、对齐后与真值分别展示米制定标前网格、保存的定标后重建与参考 CAD。参考几何用于事后对照,不参与选择重建尺寸。重建排序三边为 80.3 × 164.0 × 218.1 mm,参考为 71.7 × 164.0 × 213.5 mm;表面距离为 4.7 mm。


Cracker-box reconstruction: stages and measurements饼干盒重建:步骤与测量
The cracker-box reconstruction uses the following procedure. In the point-prompt route, SAM 2 segments the selected object. The saved comparison extracts the mask from its stored prompt overlay. SAM 3D Objects builds a mesh from the mask, then bakes a UV texture from 100 Gaussian views at 1024 pixels for 2500 steps. The raw mesh is not in meters. Depth inside the mask tries four isotropic sizes, and the higher score_6d keeps one. WAPR estimates the pose at that size. The pose stays. UniPose9D reads the same mask and the same depth and returns a 3D bbox: a rotation, a center, and three edge lengths. The mesh is scaled along those bbox axes until the posed mesh fills the bbox. The real mesh is not used to choose the size. WAPR also estimates the real mesh on this photo. Those two poses place the scaled mesh in the real mesh's coordinates. The cracker's sorted sides are 80.3 × 164.0 × 218.1 mm. The real box is 71.7 × 164.0 × 213.5 mm. The sum of absolute differences is 13.2 mm, and the mean surface distance is 4.7 mm. WAPR refines 12 poses. Poses with within_group above 0 stay, and DINOv2 keeps the render closest to the photo. That contour covers 0.952 of the mask. The pictures below show that render, with a magnified view. The bar is 10 cm.
饼干盒重建采用以下流程。点提示路径由 SAM 2 分割所选物体;保存的对比结果从提示图提取掩码。SAM 3D Objects 用所得掩码生成网格,再用 100 个 1024 像素的高斯视角烘焙 UV 贴图,优化 2500 步。原始网格的单位尚非米。mask 里的深度在四个各向同性尺度上尝试,score_6d 保留分数更高者。WAPR 在这个尺寸上估计位姿,然后位姿保持不动。UniPose9D 读同一个 mask 和同一份深度,返回一个包围盒:旋转、中心,以及三条边长。网格沿这三条包围盒轴缩放,直至已变换到该位姿的网格填满该包围盒。真实网格不参与选尺寸。WAPR 再在这张照片上估计真实网格。这两个位姿把缩放后的网格放进真实网格的坐标系。饼干盒排序后的三边是 80.3 × 164.0 × 218.1 mm。真实盒子是 71.7 × 164.0 × 213.5 mm。绝对差之和是 13.2 mm,表面平均距离是 4.7 mm。WAPR 修正 12 个位姿。within_group 大于 0 的候选姿态予以保留,再由 DINOv2 保留与照片最接近的渲染。这个轮廓覆盖 mask 的 0.952。下面的图是这次渲染,旁边是放大。标尺是 10 cm。
Single-frame reconstruction of the sugar box and mustard bottle糖盒与芥末瓶的单帧重建
The geometric comparisons evaluate the displayed meshes. Sorted side lengths are in the reference CAD frame. Surface distance is the symmetric mean nearest-point distance between 4000 sampled points per mesh; the reported value is the median of ten repetitions. The offline check gives 4.7 mm for cracker, 3.3 mm for sugar and 4.9 mm for mustard.
几何对比使用展示的网格:排序边长在参考 CAD 坐标系中计算;表面距离为每网格采样 4000 点后的对称平均最近点距离,报告十次重复的中位数。离线复核记录给出饼干盒 4.7 mm、糖盒 3.3 mm、芥末瓶 4.9 mm。
The sugar box and the mustard bottle use the same reconstruction as the cracker: the original SAM 3D Objects mesh, then a UV texture from 100 Gaussian views at 1024 pixels for 2500 steps. Sugar's sorted sides are 52.3 × 93.8 × 170.7 mm, against the real 45.4 × 92.6 × 176.1 mm. The sum of absolute differences is 13.5 mm, and the mean surface distance is 3.3 mm. Mustard is 56.4 × 90.7 × 185.7 mm, against the real 57.7 × 95.7 × 191.5 mm. The sum is 12.1 mm, and the mean surface distance is 4.9 mm. The real meshes were not used to choose the sizes. Each pose picture keeps the hypotheses with within_group above 0, then DINOv2 keeps the render closest to the photo. That contour covers 0.971 of the sugar mask and 0.902 of the mustard mask. The bar is 10 cm.
糖盒和芥末瓶的重建与饼干盒相同:原来的 SAM 3D Objects 网格,再用 100 个 1024 像素的高斯视角烘焙 UV 贴图,优化 2500 步。糖盒排序后的三边是 52.3 × 93.8 × 170.7 mm,真实盒子是 45.4 × 92.6 × 176.1 mm。绝对差之和是 13.5 mm,表面平均距离是 3.3 mm。芥末瓶是 56.4 × 90.7 × 185.7 mm,真实瓶子是 57.7 × 95.7 × 191.5 mm。绝对差之和是 12.1 mm,表面平均距离是 4.9 mm。真实网格都没有用来选尺寸。每张位姿图先留下 within_group 大于 0 的候选姿态,再由 DINOv2 留下和照片最接近的渲染。这个轮廓覆盖糖盒 mask 的 0.971,覆盖芥末瓶 mask 的 0.902。标尺是 10 cm。
Choose Before, After or Ground truth to inspect the corresponding box dimensions. Drag to rotate and scroll to zoom. 选择对齐前、对齐后或真值,查看对应包围盒尺寸;拖动旋转,滚轮缩放。
View the reconstruction results查看重建结果图
The first three panels show the input mask, reduced vertex-color mesh, and UV-baked mesh. Baking retains more of the package print; the blue reference outline is for comparison. The two texture panels show filling unbaked texels. Texture repair does not move vertices.
前三个面板为输入掩码、简化的顶点色网格,以及经过 UV 烘焙的网格。烘焙保留了更多包装印刷细节;蓝色参考轮廓仅用于对照。两个贴图面板展示补齐空纹素前后的结果。修补贴图不会移动顶点。
In this run, ×0.75 projects inside the box and ×2.00 extends beyond it. The retained ×1.40 candidate has the highest score, 0.589, and mask IoU, 0.845. These values describe this cracker-box run.
这次运行中,×0.75 的投影小于盒子轮廓,×2.00 超出轮廓。保留的 ×1.40 候选具有最高分数 0.589 和最高 mask IoU 0.845;这些数值仅对应这次饼干盒运行。
The reference side lengths are 71.7 × 164.0 × 213.5 mm for the cracker box, 45.4 × 92.6 × 176.1 mm for the sugar box, and 57.7 × 95.7 × 191.5 mm for the mustard bottle. These measurements do not select the reconstructed dimensions.
参考三边尺寸分别为:饼干盒 71.7 × 164.0 × 213.5 mm,糖盒 45.4 × 92.6 × 176.1 mm,芥末瓶 57.7 × 95.7 × 191.5 mm。这些测量值不参与选择重建网格的尺寸。
In this diagnostic, the three axis scale factors are 0.850, 1.000, and 0.893. The second axis is left unchanged: its projected extent is only 15.6 px, about 0.15 of the longest axis, below AXIS_OBSERVABLE_FRACTION = 0.40. This view does not sufficiently constrain that axis's scale.
这次尺度诊断中,三轴缩放系数为 0.850、1.000、0.893。第二条轴保持原尺度:它在图像上的投影仅为 15.6 像素,约为最长轴的 0.15,低于 AXIS_OBSERVABLE_FRACTION = 0.40。这张图对该轴的尺度约束不足。
This scale diagnostic used reference-CAD assistance. It compares mustard vertices after box sizing with the RoMa diagnostic result at the same pose; it is not the output of the observation-only default stage.
这次尺度诊断使用参考 CAD 辅助,对照包围盒定标后的芥末瓶与同一位姿下的 RoMa 诊断结果;它不代表默认仅用观测的形状更新结果。
References and licenses参考文献与许可
- DINOv2 — Apache-2.0. Code and official DINOv2 weights; retain copyright and license.源码与官方 DINOv2 权重;保留版权和许可。 · GitHubGitHub · License/notice 1许可/声明 1
- SAM 2 — Apache-2.0; cctorch: BSD-3-Clause. Native source and official checkpoints; cctorch carries an additional BSD notice. Ultralytics-converted files require checking their distributor terms.原生源码与官方权重;cctorch 另附 BSD 声明。Ultralytics 转换文件还需核对分发方条款。 · GitHubGitHub · License/notice 1许可/声明 1 · License/notice 2许可/声明 2
- SAM 3D Objects — SAM License (custom). Source and model materials follow the upstream custom agreement, including redistribution, attribution and use restrictions.源码及模型材料遵循上游自定义协议,包括再分发、署名与使用限制。 · GitHubGitHub · License/notice 1许可/声明 1
- UniPose9D — Apache-2.0. Inference repository and its release; helper models retain their own terms.推理仓库及其发布材料;辅助模型保留各自条款。 · GitHubGitHub · License/notice 1许可/声明 1
- RoMa — MIT. Matching source; separately obtained checkpoints and dependencies retain upstream terms.匹配源码;另行获取的权重与依赖保留上游条款。 · GitHubGitHub · License/notice 1许可/声明 1
- Qwen2.5-VL-3B-Instruct — Qwen Research License (custom). This exact 3B model permits non-commercial research/evaluation; commercial use requires permission from Alibaba Cloud. Do not infer its terms from other Qwen sizes.本项目选用的 3B 模型限定非商业研究/评估;商业使用须向 Alibaba Cloud 申请。不能套用其他 Qwen 规格的许可。 · Original source原始来源 · License/notice 1许可/声明 1
- YCBInEOAT data — Not separately verified / 未单独核实. Tracking data permission must be checked at the original archive; the comparison code license does not establish the data license.跟踪数据许可须在原始归档核实;对照代码的许可不能作为数据许可。 · GitHubGitHub · GitHubGitHub
- YCB-Video (YCB-V) · Official BOP dataset pageBOP 官方数据页 · GitHub: PoseCNNGitHub:PoseCNN
Model weights and dataset/task assets may have separate terms. WAPR's source license does not replace them.模型权重、数据集与任务资源可能有独立条款;WAPR 源码许可不替代这些许可。