6D pose tracking6D 位姿跟踪
For large inter-frame motion, the recorded stride-32 mustard comparison shows WAPR preserving accurate tracking where FoundationPose loses it. The sequence results and separate sugar-box recovery example document their conditions and failures.
面对较大的帧间运动,录制的步长 32 芥末瓶对照展示了 WAPR 保持准确跟踪、FoundationPose 部分帧跟踪失效的表现。序列结果与独立糖盒补救案例说明各自条件和失败情况。
6D pose tracking updates object-to-camera poses over an RGB-D sequence while preserving instance identities. Initialize poses on the first frame, carry each instance's pose and optional region state forward, and update all active tracks together on each subsequent frame. The same entry supports one or several categories and instances.
6D 位姿跟踪在连续 RGB-D 帧中更新物体到相机的位姿,并保持实例编号一致。首帧估计初始位姿,随后维护各实例的位姿与可选区域状态,在每个后续帧中批量更新全部活跃轨迹。同一入口支持单个或多个类别与实例。
Initialize poses and track identities初始化位姿与轨迹编号
Use estimate_frame_many_categories_many_instances to initialize from predicted 2D regions. Assign a unique, persistent track_id to each selected instance. Each track stores its CAD mesh, diameter_m and previous pose_4x4; obj_id is optional category metadata. Same-category instances share a mesh and retain distinct track IDs. Predictions determine the inputs; annotated poses and masks are reserved for evaluation.
通过 estimate_frame_many_categories_many_instances 从预测的 2D 区域初始化位姿,为每个选定实例分配唯一且持久的 track_id。每条轨迹保存 CAD 网格、diameter_m 与上一帧 pose_4x4,obj_id 为可选类别信息。同类别实例可共用网格,但使用不同轨迹 ID。输入由预测结果确定,标注位姿和掩码仅用于评估。
from wapr.frame import estimate_frame_many_categories_many_instances
from wapr.tracking import track_many_categories_many_instances
# Run detection and batched pose estimation once on frame 0.
# 在首帧执行一次检测与批量计算位姿估计。
initial_poses, timing, detections = estimate_frame_many_categories_many_instances(
estimator, detector, rgb0, depth0_m, K, meshes,
obj_ids=None, inst_count=None,
)
tracks = [
{"track_id": index, "obj_id": pose["obj_id"],
"mesh": meshes[pose["obj_id"]][0],
"diameter_m": meshes[pose["obj_id"]][1],
"pose_4x4": pose["pose_4x4"]}
for index, pose in enumerate(initial_poses)
]
For a complete two-object sequence with initialization, association and recovery, run example 09: TACO multi-object tracking[4].
完整的双物体序列初始化、关联与恢复流程见 示例 09:TACO 多物体跟踪[4]。
Update all tracks on one frame同帧轨迹的批量更新
On each subsequent frame, call track_many_categories_many_instances(estimator, rgb, depth_m, K, tracks) once for the active track list. RGB has shape (H, W, 3), depth has shape (H, W) in meters, and K is a 3×3 pixel-valued intrinsic matrix. Poses are 4×4 object-to-camera transforms with translation in meters. The returned list preserves input order and track_id.
后续每帧将活跃轨迹列表一次传入 track_many_categories_many_instances(estimator, rgb, depth_m, K, tracks)。RGB 形状为 (H, W, 3),深度形状为 (H, W)、单位米,K 为像素单位的 3×3 内参矩阵。位姿为物体到相机的 4×4 变换,平移单位米。返回列表保持输入顺序与 track_id。
# One call updates all active tracks on the current RGB-D frame.
# 一次调用更新当前 RGB-D 帧中的全部活跃轨迹。
updates = track_many_categories_many_instances(estimator, rgb, depth_m, K, tracks)
# Update stored poses only; rendering and neural inference are batched above.
# 这里只更新保存的位姿;渲染与神经网络推理在上方批量计算完成。
for track, update in zip(tracks, updates):
track["pose_4x4"] = update["pose_4x4"]
Each instance retains six candidates: the previous pose, the depth-centered pose, centered WAPR, centered SAPR, SAPR after centered WAPR, and SAPR after previous-pose WAPR. Shared prefixes form two WAPR rows and three SAPR rows per instance. Rendering and refinement batch across instances; WBPS scores all six-candidate groups together, with attention confined to each instance. Iteration counts and rotation caps follow wapr/recipe.py.
每个实例保留六个候选:上一帧位姿、深度居中位姿、居中位姿经 WAPR、居中位姿经 SAPR、居中位姿经 WAPR 后再经 SAPR、上一帧位姿经 WAPR 后再经 SAPR。共用前缀形成每实例两行 WAPR 和三行 SAPR 输入。渲染与修正跨实例批量计算,WBPS 同时处理全部六候选组,注意力仅在各实例内部计算。迭代次数与旋转上限采用 wapr/recipe.py 中的设置。
TensorRT[3] uses wbps_batch.engine or the optional fixed-six wbps_tracking.engine, exported on the inference GPU from the existing checkpoint. The released profiles support up to twenty instances per chunk; larger lists are split at the engine capacity. PyTorch[2] also scores instance groups together. The original four initialization engines remain in use.
TensorRT[3] 使用从现有权重在推理 GPU 上导出的 wbps_batch.engine,或可选的固定六候选 wbps_tracking.engine。发布的 profile 每块最多支持二十个实例,较大列表按引擎容量分块;PyTorch[2] 同样批量处理各实例评分组。首帧初始化继续使用原有四份引擎。
python -m wapr.export_batch_engine
Each update includes pose_4x4, R, t_m, hypothesis, hypothesis_index, and the selected candidate's between_group and score_6d. A larger between_group indicates a weaker match; score_6d = (100 - between_group) / 200. These scores support a quality policy and are not calibrated correctness probabilities. Empty input returns an empty list.
每条更新包含 pose_4x4、R、t_m、hypothesis、hypothesis_index,以及被选候选的 between_group 和 score_6d。between_group 越大表示匹配越弱,score_6d = (100 - between_group) / 200。这些分数用于制定质量判断,不是经过校准的正确概率。输入空列表时返回空列表。
Target-mask tracking and translation initialization目标掩码跟踪与平移初始化
Target-mask tracking is an optional 2D step called separately from pose updates. DINOv2[1] matches the target's patch features from the previous and first frames to the current frame, returning an approximate mask and its image center. If used, encode the current RGB once and share its features across tracks. center_pose_4x4 may supply translation from the associated region and sensor depth while preserving the previous rotation. Without this field, both starting candidates use the previous pose. wapr.region_tracking supplies the DINOv2 implementation used by example 08; its main block retains initialization, frame iteration, track state and recovery decisions.
目标掩码跟踪是可选的 2D 步骤,与位姿更新入口分别调用。DINOv2[1] 将上一帧和首帧的目标图块特征与当前帧匹配,得到近似目标 mask 及其图像中心。使用时,当前 RGB 只编码一次,各轨迹共享图像特征。center_pose_4x4 可提供由关联区域与传感器深度估计的平移,同时保留上一帧旋转;未提供该字段时,两种起始候选均采用上一帧位姿。wapr.region_tracking 提供示例 08 使用的 DINOv2 实现;初始化、帧循环、轨迹状态与恢复决策仍在示例主入口中展开。
wapr.region_tracking.dino_tokens encodes one RGB frame. one_instance_mask_to_patches and propagate_one_instance_region associate each track's stored region with those shared tokens. Example 08 explicitly selects dinov2_vits14, a 448 px short side, 14 px patches, patch_cover = 0.20, cosine_min = 0.50 and min_matches = 8. RGB is divided by 255 and normalized with mean (0.485, 0.456, 0.406) and standard deviation (0.229, 0.224, 0.225). Encoding uses FP16; matching uses FP32 tokens with 384 channels.
wapr.region_tracking.dino_tokens 编码一帧 RGB;one_instance_mask_to_patches 与 propagate_one_instance_region 将各轨迹保存的区域关联到这些共享特征。示例 08 显式选择 dinov2_vits14、448 像素短边、14 像素图块,以及 patch_cover = 0.20、cosine_min = 0.50、min_matches = 8。RGB 除以 255 后,用均值 (0.485, 0.456, 0.406) 和标准差 (0.229, 0.224, 0.225) 归一化。编码采用 FP16,匹配采用 384 通道的 FP32 特征。
dino_tokens(dino, rgb, device) returns tokens shaped (grid_h, grid_w, 384) and grid with gh, gw, nh, nw, height, width. On frame 0 the selected mask initializes both reference states. template_tokens and template_patch stay on that first mask. prev_tokens and prev_patch start as the same pair and then move forward.
dino_tokens(dino, rgb, device) 返回形状 (grid_h, grid_w, 384) 的 tokens,以及带 gh、gw、nh、nw、height、width 的 grid。第 0 帧的选定掩码同时初始化两份参考状态。template_tokens 和 template_patch 停在这块第一帧 mask 上。prev_tokens 和 prev_patch 一开始是同一对,之后逐帧往前走。
On a later frame the shared inputs are the new rgb (uint8, H×W×3, RGB), depth_m (meters, H×W), and K (3×3). dino_tokens runs once on that RGB. diameter_px is diameter_m * K[0, 0] / max(pose[2, 3], 0.05). propagate_one_instance_region matches the previous patches and the frame-0 patches into this grid, then keeps the densest cluster. The cluster cell is max(0.5 * diameter_px, 24) pixels. The return is (mask, center_uv, n_match). center_uv is the median pixel of that cluster. Fewer than 8 matches returns None.
后面每一帧,大家共用的输入是新的 rgb(uint8,H×W×3,RGB)、depth_m(米,H×W)和 K(3×3)。dino_tokens 对这张 RGB 只跑一次。diameter_px 是 diameter_m * K[0, 0] / max(pose[2, 3], 0.05)。propagate_one_instance_region 把上一帧图块和第 0 帧图块都匹配到这一帧的网格上,再留下最密的那一团。分团的格子是 max(0.5 * diameter_px, 24) 像素。返回 (mask, center_uv, n_match)。center_uv 是这一团的像素中位数。对上的图块少于 8 个时返回 None。
tokens, grid = dino_tokens(dino, rgb, device)
diameter_px = diameter_m * float(K[0, 0]) / max(float(pose[2, 3]), 0.05)
propagated = propagate_one_instance_region(
prev_tokens, prev_patch,
template_tokens, template_patch,
tokens, grid, diameter_px,
)
The depth helper returns None for small regions or unusable depth. Example 08 keeps the prior translation for its normal update and searches DINOv2 candidates.
区域过小或深度无效时,深度辅助函数返回 None。示例 08 在常规更新中保留上一平移,并搜索 DINOv2 候选。
Normal updates compare six WAPR/SAPR hypotheses using WBPS within_group. The complete-group between_group maximum describes quality. Region propagation updates prev_tokens and prev_patch independently; patch recovery supplies a pose and does not replace the region template.
常规更新由 WBPS 的 within_group 在六个 WAPR/SAPR 候选中选择,以完整组 between_group 最大值表示质量。区域传播独立更新 prev_tokens 和 prev_patch;图块补救提供位姿,不替换区域模板。
Association and lost-track recovery轨迹关联与丢失恢复
DINOv2 propagates the target region through every raw frame. At pose-update frames, missing regions, unusable depth, front-surface depth disagreement over 0.18 m, depth-normalized area below 0.70 of the preceding valid region, or positive WAPR group error trigger search. Complete the normal update first. Search up to six local and six global depth-consistent patches, preserve the previous rotation, batch-refine the translation seeds, and compare them with the normal result using WBPS within_group. Search does not produce a replacement mask. No ground truth or 2D detection is used for recovery.
DINOv2 在每个原始帧传播目标区域。位姿更新时,区域缺失、深度无效、前表面深度偏差超过 0.18 m、按深度归一化的面积低于上一有效区域的 0.70 倍,或 WAPR 组间误差为正时触发搜索。先完成常规更新,再从局部和全图各取最多六个深度一致的图块候选,保留上一旋转,批量修正平移起点,并与常规结果一起由 WBPS 的 within_group 选择。搜索不产生替代掩码;补救不用真值,也不调用 2D 检测。
Run a reference sequence运行参考序列
Example 08 provides a single-instance recipe for initialization, target-mask tracking and pose updates. Set seq_dir, mesh_path, positive diameter_m in meters, and stride in the script. The sequence supplies rgb/, millimeter depth in depth/ and cam_K.txt. In 08_ycbineoat_init_clicks.json, save the first RGB filename stem and a manually selected integer click_uv: [u, v] under the sequence name. The first-frame detector selects the highest-scoring predicted mask covering that point; later frames track the target mask. Neither dataset initialization masks nor annotated poses enter inference. The sequence must be downloaded separately.
示例 08 演示单实例的首帧初始化、目标掩码跟踪与位姿更新。在脚本中填写 seq_dir、mesh_path、以米为单位的物体直径 diameter_m(须为正数)及 stride。序列需提供 rgb/、depth/ 中的毫米深度与 cam_K.txt。在 08_ycbineoat_init_clicks.json 中按序列名保存首帧 RGB 文件名主干和人工选定的整数 click_uv: [u, v]。首帧检测器在覆盖该点的预测掩码中选择最高分项,后续帧跟踪目标掩码;数据集初始化掩码与标注位姿不参与推理。序列数据需另行下载。
python examples/08_ycbineoat_one_instance.py
DINOv2 propagates the target region through every raw frame. At pose-update frames, missing regions, unusable depth, front-surface depth disagreement over 0.18 m, depth-normalized area below 0.70 of the preceding valid region, or positive WAPR group error trigger search. Complete the normal update first. Search up to six local and six global depth-consistent patches, preserve the previous rotation, batch-refine the translation seeds, and compare them with the normal result using WBPS within_group. Search does not produce a replacement mask. No ground truth or 2D detection is used for recovery.
DINOv2 在每个原始帧传播目标区域。位姿更新时,区域缺失、深度无效、前表面深度偏差超过 0.18 m、按深度归一化的面积低于上一有效区域的 0.70 倍,或 WAPR 组间误差为正时触发搜索。先完成常规更新,再从局部和全图各取最多六个深度一致的图块候选,保留上一旋转,批量修正平移起点,并与常规结果一起由 WBPS 的 within_group 选择。搜索不产生替代掩码;补救不用真值,也不调用 2D 检测。
Tracking questions跟踪常见问题
How does WAPR compare with FoundationPose under large motion?大运动下,WAPR 与 FoundationPose 的对照表现如何?
In the recorded large-motion mustard sequence mustard_easy_00_02, updating poses every 32 frames gives WAPR 100% ADD recall and 4.6 mm mean ADD, versus FoundationPose’s 68.2% and 36.3 mm. FoundationPose loses accurate tracking in parts of this comparison. Both methods start from the same predicted mask; neither reinitializes poses in this sequence. Recall uses ADD below 0.1 object diameter. These are sequence-specific results.
在录制的大运动芥末瓶序列 mustard_easy_00_02 中,每 32 帧更新一次位姿,WAPR 的 ADD 达标率为 100%、平均 ADD 为 4.6 mm;FoundationPose 分别为 68.2% 和 36.3 mm,对照中部分帧出现跟踪失效。双方从同一预测掩码开始,该序列均不重新初始化位姿。达标阈值为物体直径的 0.1 倍;这些是此序列的结果。
Does tracking recovery always succeed?跟踪补救是否一定成功?
Lost-track recovery is a separate optional workflow using DINOv2 candidate search and pose scoring. The sugar-box example retains failed predictions and demonstrates recovery behavior; it does not guarantee recovery for every sequence.
跟踪丢失补救是单独的可选流程,使用 DINOv2 候选搜索及位姿评分。糖盒案例保留失败预测并展示补救表现,不保证所有序列都能恢复。
References and licenses参考文献与许可
- DINOv2 — Apache-2.0. Code and official DINOv2 weights; retain copyright and license.源码与官方 DINOv2 权重;保留版权和许可。 · GitHubGitHub · License/notice 1许可/声明 1
- torch — BSD License. Runtime dependency; dependencies and model/data assets retain their own licenses.运行依赖;依赖及模型/数据资源保留各自许可。 · Original source原始来源 · License/notice 1许可/声明 1 · License/notice 2许可/声明 2 ↩ ↩
- tensorrt-cu12 — Other/Proprietary License. Runtime dependency; dependencies and model/data assets retain their own licenses.运行依赖;依赖及模型/数据资源保留各自许可。 · GitHubGitHub · License/notice 1许可/声明 1 ↩ ↩
- TACO data — Not separately verified / 未单独核实. Data terms have not been separately verified. The official repository is the original source; the referenced Hugging Face acquisition mirror has no explicit license field. Confirm the data owner's terms for redistribution or commercial use.数据条款尚未单独核实。官方仓库为原始出处;实际获取数据的 Hugging Face 镜像数据卡未声明明确许可字段。再分发或商业使用须核实数据权利方的条款。 · GitHubGitHub
Model weights and dataset/task assets may have separate terms. WAPR's source license does not replace them.模型权重、数据集与任务资源可能有独立条款;WAPR 源码许可不替代这些许可。