YCBInEOAT single-object pose trackingYCBInEOAT 单物体位姿跟踪
In the recorded large-motion mustard sequence mustard_easy_00_02, updating poses every 32 frames gives WAPR 100% ADD recall and 4.6 mm mean ADD, versus FoundationPose’s 68.2% and 36.3 mm. FoundationPose loses accurate tracking in parts of this comparison. Both methods start from the same predicted mask; neither reinitializes poses in this sequence. Recall uses ADD below 0.1 object diameter. These are sequence-specific results.
在录制的大运动芥末瓶序列 mustard_easy_00_02 中,每 32 帧更新一次位姿,WAPR 的 ADD 达标率为 100%、平均 ADD 为 4.6 mm;FoundationPose 分别为 68.2% 和 36.3 mm,对照中部分帧出现跟踪失效。双方从同一预测掩码开始,该序列均不重新初始化位姿。达标阈值为物体直径的 0.1 倍;这些是此序列的结果。
Lost-track recovery is a separate optional workflow using DINOv2 candidate search and pose scoring. The sugar-box example retains failed predictions and demonstrates recovery behavior; it does not guarantee recovery for every sequence.
跟踪丢失补救是单独的可选流程,使用 DINOv2 候选搜索及位姿评分。糖盒案例保留失败预测并展示补救表现,不保证所有序列都能恢复。
This case explains single-object tracking on YCBInEOAT[6], with visual walkthroughs, accuracy and timing statistics. For initialization and batch updates, see 6D pose tracking. Each comparison specifies whether its initial mask is predicted from RGB or supplied as an input.
本案例介绍 YCBInEOAT[6] 单物体跟踪,展示处理过程、精度与耗时。初始化与批量更新的使用方法见 6D 位姿跟踪。各项对照注明首帧掩码由 RGB 预测还是作为输入提供。
Data and evaluation scope数据与评测范围
The batch update entry accepts several known tracks in one GPU call; the recorded YCBInEOAT example remains a single-instance workflow. Each reported YCBInEOAT run uses one sequence and one obj_id. The nine YCBInEOAT sequences on disk each store one 4×4 pose and one mask per frame: cracker_box_reorient, cracker_box_yalehand0, sugar_box1, sugar_box_yalehand0, tomato_soup_can_yalehand0, mustard0, mustard_easy_00_02, bleach0, bleach_hard_00_03_chaitanya. cracker_box_reorient also shows other objects on the table. Those objects are absent from annotated_poses. The tracker predicts its frame-0 mask with WAPRDet2D. The clips on the project page are three of these one-instance sequences, with warmed predicted-mask initialization for cracker and mustard. The selective sugar-box clip retains the shared supplied frame-0 mask and its separate full-path warm protocol.
批量更新入口可在一次 GPU 调用中处理多个身份已知的轨迹;现有 YCBInEOAT 示例仍是单实例流程。这里报告的 YCBInEOAT 结果每次对应一个序列和一个 obj_id。磁盘上的九段 YCBInEOAT 每一帧都保存一个 4×4 位姿矩阵和一块 mask:cracker_box_reorient、cracker_box_yalehand0、sugar_box1、sugar_box_yalehand0、tomato_soup_can_yalehand0、mustard0、mustard_easy_00_02、bleach0、bleach_hard_00_03_chaitanya。cracker_box_reorient 的桌子上还有别的物体。它们不在 annotated_poses 里。跟踪示例用 WAPRDet2D 预测第 0 帧 mask。项目主页上的三段就是这些单实例序列里的三段,其中饼干盒和芥末瓶使用预热后的预测掩码初始化记录;选择性糖盒片段保留共享的第 0 帧提供掩码,采用另行完整预热的协议。
Batch-update compute checks批量更新的计算检查
RTX 5090 compute check on wapr-test: fourteen repeated predicted-pose inputs took a median 0.060 s in one batch versus 0.520 s serially (8.59×, three warmups and five timed calls). The largest pose difference was 0.096 mm and 0.179°; 6 of fourteen hypothesis indices differed. The 21- and 64-instance checks exercised chunk boundaries. A separate two-frame robot check verified one DINO encode per camera frame and batched updates of two tracks. These checks do not measure full-sequence accuracy or episode latency.
wapr-test 的 RTX 5090 计算检查:十四个重复预测位姿输入,一次批量计算中位数 0.060 秒,逐实例中位数 0.520 秒,约 8.59 倍加速;预热三次、测量五次。最大位姿差 0.096 mm、0.179°,十四个候选索引中 6 个不同。21 和 64 实例检查覆盖分块边界。另用机器人两帧检查确认,每相机帧只编码一次 DINO,两条轨迹批量更新。这些检查不代表完整序列精度或回合耗时。
Validation with predicted initial masks预测掩码初始化对照
This demonstration's WAPR initializer begins with twelve hypotheses, runs two masked updates, retains one in-plane hypothesis per view, then runs one masked update and two SAPR updates on four hypotheses. Normal updates use one mask-free WAPR step and two SAPR steps. FoundationPose[4] uses two tracking iterations and five registration iterations, with its native scorer. These schedules differ from the three-refinement selective sugar reassociation experiment below. Warmup discards three to six calls per path; it removes cold calls but does not guarantee constant latency. The ordinary public initialization keeps all twelve hypotheses through its three WAPR and two SAPR updates.
本演示的 WAPR 初始化从十二条候选姿态开始,带掩码更新两次后,每个视角保留一个面内候选,再对四条候选姿态执行一次带掩码更新与两次 SAPR;常规更新为一次不带掩码 WAPR 加两次 SAPR。FoundationPose[4] 跟踪迭代两次、注册迭代五次,仅用原有评分器。这与下方各修正三次的糖盒选择性重关联实验不同。每条路径丢弃三至六次预热调用,排除预热前的调用,但不保证延迟恒定。常规公共初始化则保留全部十二条候选姿态,执行三次 WAPR 与两次 SAPR。
Cracker and mustard initialize from masks predicted from RGB by WAPRDet2D, with a saved manual click specifying the instance. Sugar supplies a shared initial mask and independently searches DINOv2[2] candidates. Both methods use their own pose state and scorer; annotated poses do not enter inference.
饼干盒与芥末瓶由 WAPRDet2D 根据 RGB 预测掩码,保存的人工点击点指定实例。糖盒共享给定初始掩码,并独立搜索 DINOv2[2] 候选。双方使用各自的位姿状态和评分器,标注位姿不参与推理。
| Sequence / stride序列 / 步长 | WAPR recallWAPR 召回率 | FoundationPose[4] recallFoundationPose[4] 召回率 |
|---|---|---|
| cracker_box_reorient / 1 | 100% | 100% |
| mustard_easy_00_02 / 32 | 100% | 68.2% |
| sugar_box1 / 32 | 93.1% (27/29) | 93.1% (27/29) |
ADD recall uses sampled poses below 0.1 object diameter. In sugar_box1, both methods pass on 27/29 poses. WAPR selects search candidates on seven updates, FoundationPose on three. The video retains failures and shows measured RTX 5090 update times; DINOv2 region tracking is timed separately.
ADD 达标率按采样位姿误差低于物体直径 0.1 倍统计。糖盒中双方均为 27/29 个位姿达标,WAPR 在七次更新中选中搜索候选,FoundationPose 为三次。视频保留失败,并显示 RTX 5090 的实测更新时间;DINOv2 区域跟踪另计。
The script predicts the initial mask from RGB without reading init_mask.png or gt_mask. Each sequence has a manually selected point in examples/08_ycbineoat_init_clicks.json. Among predicted masks covering that point, the script selects the highest score_2d and saves the mask and score under outputs/tracking_init/. See the predicted-mask comparison above.
脚本从 RGB 预测初始化掩码,不读取 init_mask.png 或 gt_mask。每段序列的人工选点保存在 examples/08_ycbineoat_init_clicks.json;脚本在覆盖该点的预测掩码中选择 score_2d 最高者,将掩码与分数保存到 outputs/tracking_init/。结果见上方预测掩码初始化对照。
The example uses the selected pose’s score from the six-candidate WBPS group. Recovery compares the tracked and re-estimated poses in a two-candidate group; WBPS requires L>1. Sequence accuracy for this example’s grouped recovery configuration has not been measured.
示例使用六候选 WBPS 组中被选位姿的评分。恢复时,将跟踪位姿与重估位姿组成两候选组比较;WBPS 要求 L>1。该示例的成组恢复配置尚未完成序列精度评测。
Cracker box stride 1饼干盒 步长 1
Mustard bottle stride 32芥末瓶 步长 32
Sugar box recovery stride 32糖盒补救 步长 32
Selective reassociation and recovery experiment选择性重关联与恢复实验
预测掩码初始化对照:糖盒补救视频,步长 32
Tracking with predicted initial masks: sugar-box recovery video, stride 32
This diagram describes the sugar-box experiment, which starts from a supplied frame-0 mask and updates poses every 32 frames. DINOv2 tracks the target region through the intervening frames. Each pose update uses valid region depth to seed translation, or retains the previous pose as its seed, then runs mask-free WAPR/SAPR and WBPS. Reassociation runs only when the normal WBPS raw between-group error is at least −50 and there is a missing region, depth mismatch, region collapse or abrupt error increase. Candidates are compared against the normal pose in the same scoring group; an accepted candidate must reduce that error. These raw errors can be negative and are different from the public higher-is-better score_6d.
下图对应糖盒的选择性跟踪实验:使用提供的第 0 帧掩码,位姿更新步长为 32;DINOv2 在中间帧持续跟踪目标区域。每次更新用有效区域深度设置起始平移,否则以上一位姿为起点,再执行无掩码 WAPR/SAPR 修正与 WBPS 评分。仅当常规位姿的 WBPS 原始组间误差 ≥ −50,且存在区域缺失、深度不符、区域骤缩或误差突增时,才尝试重新关联。候选与常规位姿在同一组内评分,误差降低才接受。原始误差可以为负,与公开的、越高越好的 score_6d 不同。
DINOv2 propagates the target region through every raw frame. At pose-update frames, missing regions, unusable depth, front-surface depth disagreement over 0.18 m, depth-normalized area below 0.70 of the preceding valid region, or positive WAPR group error trigger search. Complete the normal update first. Search up to six local and six global depth-consistent patches, preserve the previous rotation, batch-refine the translation seeds, and compare them with the normal result using WBPS within_group. Search does not produce a replacement mask. No ground truth or 2D detection is used for recovery.
DINOv2 在每个原始帧传播目标区域。位姿更新时,区域缺失、深度无效、前表面深度偏差超过 0.18 m、按深度归一化的面积低于上一有效区域的 0.70 倍,或 WAPR 组间误差为正时触发搜索。先完成常规更新,再从局部和全图各取最多六个深度一致的图块候选,保留上一旋转,批量修正平移起点,并与常规结果一起由 WBPS 的 within_group 选择。搜索不产生替代掩码;补救不用真值,也不调用 2D 检测。
DINOv2 propagates the target region through every raw frame. At pose-update frames, missing regions, unusable depth, front-surface depth disagreement over 0.18 m, depth-normalized area below 0.70 of the preceding valid region, or positive WAPR group error trigger search. Complete the normal update first. Search up to six local and six global depth-consistent patches, preserve the previous rotation, batch-refine the translation seeds, and compare them with the normal result using WBPS within_group. Search does not produce a replacement mask. No ground truth or 2D detection is used for recovery.
DINOv2 在每个原始帧传播目标区域。位姿更新时,区域缺失、深度无效、前表面深度偏差超过 0.18 m、按深度归一化的面积低于上一有效区域的 0.70 倍,或 WAPR 组间误差为正时触发搜索。先完成常规更新,再从局部和全图各取最多六个深度一致的图块候选,保留上一旋转,批量修正平移起点,并与常规结果一起由 WBPS 的 within_group 选择。搜索不产生替代掩码;补救不用真值,也不调用 2D 检测。
For the same saved 480×640 YCB image and DINOv2-L CAD matcher, the hot 2D call took a median 220.7 ms with GroundingDINO[1] + SAM2[3] and 85.0 ms with FastSAM[7]-s on RTX 5090 (3 warmups, 12 timed calls). These calls stop after 2D matching: no 6D pose or video tracking is measured. The faster proposal path has not by itself established better pose accuracy or recovery.
相同的 480×640 YCB 图像、相同 DINOv2-L CAD 匹配器下,RTX 5090 的热 2D 调用中位数为 GroundingDINO[1] + SAM2[3] 方案 220.7 ms、FastSAM[7]-s 方案 85.0 ms(预热三次、计时十二次)。这些调用在 2D 匹配后结束,未测 6D 位姿或整段跟踪;更快的候选生成本身不能证明位姿精度或找回率更高。
One repeated frame, 12 timed calls after 3 warmups. Dark segments are proposal and segmentation; light segments are later DINOv2 encoding, CAD matching, and postprocessing. Full per-call results are in the research benchmark files.
同一帧预热 3 次后计时 12 次。深色为候选生成与分割,浅色为后续 DINOv2 编码、CAD 匹配及后处理;逐次结果见研究基准文件。
The number of regions sent to full pose initialization also affects recovery cost. The inference and renderer scaling experiment measures the current grouped-WBPS path, with twelve hypotheses per instance and up to 1000 instances. It separates neural execution from rendering and varies CAD categories at a fixed total workload. These measurements cover full masked initialization, rather than routine mask-free tracking updates.
送入完整位姿初始化的区域数量也会影响补救成本。推理与渲染后端扩展实验测量当前成组 WBPS 路径,每实例十二个候选,规模覆盖至 1000 个实例;分别比较网络执行与渲染,并在固定总负载下改变 CAD 类别。该实验测量完整带掩码初始化,与常规无掩码跟踪更新的范围不同。
Single-instance timing and accuracy comparison单实例耗时与精度对照
The sugar-box comparison uses stride 32 and the same supplied initial mask. Both methods maintain their own previous pose and DINOv2 region state. Missing regions, front-surface depth disagreement over 0.18 m or depth-normalized area below 0.70 of the preceding valid region trigger search independently. WAPR additionally checks whether the complete-group between_group maximum is positive. Each completes its normal update before comparing search candidates. Ground-truth poses are used only to evaluate the outputs.
糖盒对照采用步长 32,共享给定的初始掩码。两侧各自维护上一位姿和 DINOv2 区域状态,区域缺失、前表面深度偏差超过 0.18 m,或按深度归一化的面积低于上一有效区域的 0.70 倍时独立触发搜索。WAPR 还检查完整组 between_group 的最大值是否为正。双方先完成常规更新,再与搜索候选比较;真值位姿仅用于评价输出。
DINOv2 returns up to six local and six global patch candidates, requiring similarity at least 0.35, depth difference at most 0.18 m and separation at least 25 pixels. The local radius is 0.55 projected object diameters. Each seed retains that method’s previous rotation and adjusts translation from sensor depth. WAPR batch-refines the seeds with one mask-free WAPR and two SAPR iterations, then ranks them with the normal result using WBPS within_group. FoundationPose batch-refines and ranks its own candidates with its own scorer. A selected search candidate is not necessarily an accurate recovered pose. No 2D detector is called.
DINOv2 从局部和全图各取最多六个图块候选,要求相似度至少 0.35、深度差不超过 0.18 m、候选间隔至少 25 像素;局部半径为投影物体直径的 0.55 倍。各起点保留对应方法的上一旋转,用传感器深度调整平移。WAPR 对搜索起点批量执行一次无掩码 WAPR 和两次 SAPR,再与常规结果一起由 WBPS 的 within_group 选择。FoundationPose 批量修正自身候选,并用自身评分器排序。选中搜索候选不等于准确恢复;补救不调用 2D 检测。
Both methods pass on 27/29 sampled poses. Mean ADD, including failed poses, is 7.1 mm for WAPR and 70.1 mm for FoundationPose. WAPR searches on 12 updates and FoundationPose on 4. At frame 864 both fail; at 896 WAPR passes with ADD 18.4 mm, while FoundationPose remains inaccurate at 1158.2 mm. The ADD threshold is 19.8 mm. Failed poses remain in the table and video.
双方均为 27/29 个采样位姿达标。包含失败位姿的平均 ADD 为 WAPR 7.1 mm、FoundationPose 70.1 mm;分别在 12 次和 4 次更新中搜索。第 864 帧双方未达标;第 896 帧 WAPR 达标,ADD 为 18.4 mm,FoundationPose 仍为错误位姿,ADD 为 1158.2 mm。ADD 阈值为 19.8 mm;失败位姿保留在表格和视频中。
After warmup on RTX 5090, normal-update medians are 27.8 ms for WAPR and 369.1 ms for FoundationPose. Full updates containing search take 64.6–172.5 ms and 420.0–445.5 ms; initialization takes 47.7 and 431.9 ms. CUDA-synchronized update times include loss checks, normal computation, patch search, seed construction, refinement, own scoring and state writes. DINOv2 encoding, matching and region propagation are measured separately at 6.74 and 6.69 ms per raw frame. Loading, I/O, offline evaluation and drawing are excluded. The video rounds each measured update upward to 0.1 s; these are not camera FPS.
RTX 5090 预热后,常规更新中位数为 WAPR 27.8 ms、FoundationPose 369.1 ms;含搜索的完整更新分别为 64.6–172.5 ms 和 420.0–445.5 ms,首帧初始化为 47.7、431.9 ms。CUDA 同步更新计时包含失效检查、常规计算、图块搜索、起点构造、修正、各自评分及状态写回。DINOv2 编码、匹配和区域传播另计,分别为 6.74、6.69 ms/原始帧;不含加载、文件读取、离线评价和绘图。视频将每次实测更新向上取整至 0.1 秒,这些时间不表示相机帧率。
| Sequence序列 | Method方法 | ADD | Mean平均mm | First首帧ms | Later之后ms |
|---|---|---|---|---|---|
cracker_box_reorientstride 1步长 1 |
WAPR | 100% | 6.0 | 90.7 | 69.6 |
| FoundationPose[4] | 100% | 5.1 | 552.0 | 68.5 | |
mustard_easy_00_02stride 32步长 32 |
WAPR | 100% | 4.6 | 87.4 | 71.8 |
| FoundationPose | 68.2% | 36.3 | 559.6 | 73.2 | |
sugar_box1stride 32, DINOv2 recovery步长 32,DINOv2 补救 |
WAPR | 93.1% | 7.1 | 47.7 | 27.8 / 64.6–172.5 |
| FoundationPose | 93.1% | 70.1 | 431.9 | 369.1 / 420.0–445.5 |
Sugar-box first-frame initialization time糖盒首帧初始化耗时
WAPR first-frame initialization on RTX 5090 takes 89.5 ms, the median of ten CUDA-synchronized calls after three discarded warmups, with a measured range of 43.8–180.4 ms. The measurement uses OpenGL and TensorRT[5] FP16, the original RGB-D frame, camera intrinsics, CAD mesh and supplied initial mask. It follows the initialization schedule used in the comparison: twelve initial candidates, two WAPR corrections, selection of one candidate per view (four retained), one further WAPR correction, two SAPR corrections and WBPS selection. It excludes model loading, mesh upload, disk I/O, detection, DINOv2 tracking and drawing; prepared observation tensors are reused after warmup.
WAPR 在 RTX 5090 上的首帧初始化耗时为89.5 ms,为丢弃三次预热调用后,十次 CUDA 同步计时的中位数,实测范围为 43.8–180.4 ms。使用 OpenGL 与 TensorRT[5] FP16,以及原实验的 RGB-D 帧、相机内参、CAD 网格和给定初始掩码。初始化流程为:十二个初始候选,先做两次 WAPR 修正,每个视角保留一个候选(剩余四个),再做一次 WAPR 修正、两次 SAPR 修正并由 WBPS 选择。不含模型加载、网格上传、磁盘读取、检测、DINOv2 跟踪与绘图;预热后复用已准备的观测张量。
Both frame-0 values in the table are synchronized, warmed initialization calls from the linked per-frame records. The final column shows the normal-update median followed by the range of complete recovery updates; the latter includes 2D detection and pose initialization. The raw records retain discarded warmup samples and prediction poses.
表中两种方法的首帧数值均来自所链接逐帧记录中的同步、预热后初始化调用。最后一列先列常规更新中位数,再列完整恢复更新范围;后者包含 2D 检测与位姿初始化。原始记录同时保留丢弃的预热样本及预测位姿。
Illustrated sequence comparisons序列对照的图解过程
The In cells quote Example 08. Images and paired videos are separate experiment visualizations, not its console output. Example 08 uses six normal hypotheses and predicted initialization; the sugar comparison uses twelve normal WAPR hypotheses and a supplied initial mask. Both use DINOv2 search followed by their own pose scorer.
In 单元摘录示例 08 原文。图片与双方法视频是独立实验可视化,不是该脚本的控制台输出。示例 08 使用六个常规候选和预测初始化;糖盒对照使用十二个常规 WAPR 候选及给定初始掩码。二者均以 DINOv2 搜索候选,再由自身位姿评分器选择。
These pictures use the warmed RTX 5090 predicted-initialization validation for cracker_box_reorient, stride 1. Both methods share the frame-0 mask predicted by WAPRDet2D and selected with a saved manual click. Later each refines its own previous pose. DINOv2 tracks a target mask for illustration but does not reset either translation. Red is the estimate and green is the offline reference. The times shown belong to these recorded calls; they exclude detection, DINOv2 mask tracking and drawing. The public batch tracking interface is a separate workload.
下图采用 RTX 5090 预热后的 cracker_box_reorient 步长 1 预测初始化验证。两种方法共用 WAPRDet2D 预测、由保存的人工点击选中的第 0 帧掩码;后续各自修正上一帧位姿。DINOv2 跟踪得到的目标掩码仅供观察,不重设两侧平移。红色是估计,绿色是离线参考。显示时间对应这些已记录调用,不含检测、DINOv2 掩码跟踪与绘图;公共批量跟踪接口是另一种负载。
Initialize from frame 0用首帧初始化位姿
Left is WAPR; right is FoundationPose. Both initialize from the same predicted frame-0 mask selected by a saved manual click. The warmed frame-0 calls take 90.7 ms for WAPR and 552.0 ms for FoundationPose's register.
左侧为 WAPR,右侧为 FoundationPose。两者用同一块由保存的人工点击选中的预测首帧 mask 初始化。预热后的第 0 帧调用分别为 WAPR 90.7 ms、FoundationPose register 552.0 ms。
torch.cuda.synchronize(device)
init_started = time.perf_counter()
init = estimator.estimate_one_category_one_instance(
rgb, depth_m, K, mesh, float(diameter_m), mask=mask0,
)
torch.cuda.synchronize(device)
init_hot_seconds = time.perf_counter() - init_started
pose = np.asarray(init["pose_4x4"], dtype=np.float32)
print(
{"frame": stem0, "kind": "init", "pose_hot_seconds": init_hot_seconds,
"t_m": pose[:3, 3].reshape(3).tolist()},
flush=True,
)Source: examples/08_ycbineoat_one_instance.py, lines 360–372代码来源:examples/08_ycbineoat_one_instance.py,第 360–372 行
The two pose methods share the frame-0 mask. The bar values and pose outlines come from the same warmed RTX 5090 validation.两种方法共用第 0 帧 mask。顶栏数值与位姿轮廓来自同一次 RTX 5090 预热后调用验证。
Track the target mask in the next frame在下一帧跟踪目标掩码
One frame later there is no new given mask. DINOv2 ViT-S/14 resizes the new RGB, takes one layer of patch tokens in FP16, and matches two sets into that grid: the patches of the previous frame, and the patches of frame 0. Attention is the GPU flash kernel. Each object patch keeps its nearest current patch when the cosine is high enough. The densest cluster of those matches is painted back as rectangles. That is why the yellow region on the right is blocky. The yellow region is a 2D mask, not a pose. The displayed next-frame region update is 7 ms on this RTX 5090; its sequence median is 6.1 ms. This stride-1 region update is a different workload from the selective sugar experiment's dense region stream.
再过一帧,没有新的给定 mask。DINOv2 ViT-S/14 把新的 RGB 缩放后,用 FP16 取出一层图块特征,再把两组图块匹配到这一帧:上一帧的图块,和第 0 帧的图块。注意力用 GPU 的 flash 核。余弦够高时,每个物体图块留下当前帧上最近的图块。匹配最密的那一团涂回成矩形。所以右边的黄色区域是一块一块的。黄色区域是 2D mask,不是位姿。这台 RTX 5090 显示的下一帧区域更新为 7 毫秒,整段中位数为 6.1 毫秒;步长 1 区域更新与选择性糖盒实验的密集区域流负载不同。
for frame_index, stem in enumerate(frame_stems(seq_dir)[1:], start=1):
rgb, depth_m = read_rgbd(seq_dir, stem)
last_rgb = rgb
tokens, grid = dino_tokens(dino, rgb, device, dino_short_side=dino_short_side, dino_patch=dino_patch)
z_prev = max(float(pose[2, 3]), 0.05)
diameter_px = float(diameter_m) * float(K[0, 0]) / z_prev
propagated = propagate_one_instance_region(
prev_tokens, prev_patch, template_tokens, template_patch, tokens, grid, diameter_px,
cosine_min=cosine_min, min_matches=min_matches, dino_patch=dino_patch,
)Source: examples/08_ycbineoat_one_instance.py, lines 403–412代码来源:examples/08_ycbineoat_one_instance.py,第 403–412 行
Left is frame 0. The yellow fill is the predicted initial mask, and the bar rounds 5.7 ms to 6 ms. Right is the next frame. The yellow rectangles are the matched ViT-S patches, and the bar reads 7 ms. Both times are this 5090, after warmup.左边是第 0 帧。黄色填充是预测的初始 mask,顶栏将 5.7 毫秒取整显示为 6 毫秒。右边是下一帧,黄色矩形是匹配到的 ViT-S 图块,顶栏 7 毫秒。两个时间都是这台 5090,预热之后。
Compare pose updates on the same frame比较同一帧的位姿更新
Both methods now update on the same next frame from their own previous pose. WAPR applies one wapr_wo_mask step followed by two SAPR steps; FoundationPose applies track_one. Neither method takes a new mask or resets translation from the yellow DINOv2 region at stride 1. This recorded next-frame call takes 18.4 ms for WAPR and 24.0 ms for FoundationPose; these individual calls differ from the sequence medians.
下一帧,两种方法都从各自上一帧位姿更新。WAPR 执行一次 wapr_wo_mask 和两次 SAPR 修正;FoundationPose 调用 track_one。步长 1 时两者都不接收新 mask,也不用黄色 DINOv2 区域重设平移。记录的这一下一帧调用分别为 WAPR 18.4 ms、FoundationPose 24.0 ms;单次调用不同于整段中位数。
center_pose = pose.copy()
if translation is not None:
center_pose[:3, 3] = translation
# Finish the normal update first; it remains a search comparison candidate.
# 先完成常规更新,保留其结果与搜索候选比较;不使用真值触发或筛选。
tracked = track_many_categories_many_instances(estimator, rgb, depth_m, K,
[{"track_id": sequence_name, "mesh": mesh, "diameter_m": float(diameter_m),
"pose_4x4": previous_pose, "center_pose_4x4": center_pose}])[0]
pose = np.asarray(tracked["pose_4x4"], dtype=np.float32)
kind = HYPOTHESES[int(tracked["hypothesis_index"])]
score = float(tracked["between_group"])
if score > recover_between_group:
reasons.append("wbps_quality")Source: examples/08_ycbineoat_one_instance.py, lines 445–457代码来源:examples/08_ycbineoat_one_instance.py,第 445–457 行
The next tracking frame. Both estimates begin from the preceding pose; the yellow region is not used to initialize either update. The displayed times are measured warm calls from this trace.跟踪的下一帧。两种估计都从前一帧位姿出发,黄色区域不参与这次位姿更新的初始化。显示时间为该轨迹中实测的预热后调用。
Initialize the same sequence with 2D detection用 2D 检测初始化同一序列
This example initializes tracking with a mask predicted from frame-0 RGB. WAPRDet2D uses the known cracker-box CAD template bank. A saved manual click specifies the target instance; the highest score_2d among predicted masks covering that click is selected. Annotated masks and poses are reserved for evaluation. On RTX 5090 with Python 3.10, detection takes a median 278.3 ms over ten synchronized calls after three warmups. The displayed initial pose takes 90.7 ms in its warmed frame-0 call. Later frames in this stride-1 clip refine the preceding pose without another detection.
本例使用首帧 RGB 预测的掩码初始化跟踪。WAPRDet2D 使用已知饼干盒的 CAD 模板库;保存的人工点击点指定目标实例,在覆盖该点的预测掩码中选择 score_2d 最高者。标注掩码和位姿仅用于评估。Python 3.10、RTX 5090 上,检测预热三次后同步计时十次,中位数为 278.3 ms;图中首帧位姿估计来自预热后的一次调用,耗时 90.7 ms。这个步长为 1 的片段在后续帧直接修正上一帧位姿,不重复检测。
detections, timing = detector.detect_many_categories_many_instances(rgb, profile=True)
if not detections:
raise RuntimeError("No frame-0 detection for %s / 第 0 帧没有检测候选" % stem0)
clicked_detections = []
for detection in detections:
visible = np.asarray(mask_utils.decode(detection["mask"])) > 0
if visible.shape != rgb.shape[:2]:
raise RuntimeError("Detection mask shape / 检测掩码尺寸错误")
if visible[click_v, click_u] and int(np.count_nonzero(visible)) >= min_region_px:
clicked_detections.append((detection, visible))
if not clicked_detections:
raise RuntimeError("No detector mask covers the saved click / 没有检测掩码覆盖保存的点击点")
top, visible = max(clicked_detections, key=lambda item: float(item[0]["score_2d"]))
mask0 = visible.astype(np.uint8) * 255Source: examples/08_ycbineoat_one_instance.py, lines 316–329代码来源:examples/08_ycbineoat_one_instance.py,第 316–329 行
Left: the selected predicted mask and its white detection box, with score_2d 0.805. Right: the initial pose estimated from that mask. Red denotes the estimate; green denotes ground truth. The detector label is a repeated-call median; the pose label is the displayed frame's measured call.左侧为选中的预测掩码与白色检测框,score_2d 为 0.805;右侧为该掩码初始化得到的位姿。红色表示估计,绿色表示真值。检测时间是多次调用的中位数,位姿时间是图中这一帧的实测调用时间。
print(
{
"frame": stem,
"kind": kind,
"matches": int(nmatch),
"between_group": score,
"recovery_trigger": reasons,
"candidate_count": candidate_count,
"recovery_accepted": recovery_accepted,
"search_ms": search_ms,
"pose_hot_seconds": pose_hot_seconds,
"t_m": pose[:3, 3].reshape(3).tolist(),
},
flush=True,
)Source: examples/08_ycbineoat_one_instance.py, lines 483–497代码来源:examples/08_ycbineoat_one_instance.py,第 483–497 行
Left: DINOv2 tracks the predicted initial target mask, shown in yellow. Right: WAPR refines its own previous pose without resetting translation from the region. Both labels are measured on RTX 5090 for this frame; the 18.4 ms update is distinct from the sequence median of 69.6 ms. There is no second detection.左侧为 DINOv2 跟踪得到的目标掩码,以黄色标出;右侧为 WAPR 直接修正自身上一帧位姿,不根据该区域重设平移。两侧均为这一帧的 RTX 5090 实测时间;18.4 ms 是本帧更新耗时,区别于整个序列的 69.6 ms 中位数。此帧不再次检测。
Recorded tracking stills已记录跟踪静帧
The stills below use the same warm records as the project videos: predicted initialization for cracker and mustard, and the supplied-initial-mask reassociation experiment for sugar. They do not measure the public batch interface. Sequence inputs and download links are listed in Data and examples.
下方静帧采用与主页视频相同的预热后调用记录:饼干盒与芥末瓶为预测初始化,糖盒为提供初始掩码的重关联实验;不能作为公共批量计算接口的测速。序列输入与下载链接统一见数据与示例。
Three recorded sequences三段记录序列
Both methods pass on 27/29 sampled poses. Mean ADD, including failed poses, is 7.1 mm for WAPR and 70.1 mm for FoundationPose. WAPR searches on 12 updates and FoundationPose on 4. At frame 864 both fail; at 896 WAPR passes with ADD 18.4 mm, while FoundationPose remains inaccurate at 1158.2 mm. The ADD threshold is 19.8 mm. Failed poses remain in the table and video.
双方均为 27/29 个采样位姿达标。包含失败位姿的平均 ADD 为 WAPR 7.1 mm、FoundationPose 70.1 mm;分别在 12 次和 4 次更新中搜索。第 864 帧双方未达标;第 896 帧 WAPR 达标,ADD 为 18.4 mm,FoundationPose 仍为错误位姿,ADD 为 1158.2 mm。ADD 阈值为 19.8 mm;失败位姿保留在表格和视频中。
References and licenses参考文献与许可
- GroundingDINO — Apache-2.0. Vendored detector source; original notices retained.随包检测源码;保留原始声明。 · GitHubGitHub · License/notice 1许可/声明 1 · License/notice 2许可/声明 2
- DINOv2 — Apache-2.0. Code and official DINOv2 weights; retain copyright and license.源码与官方 DINOv2 权重;保留版权和许可。 · GitHubGitHub · License/notice 1许可/声明 1
- SAM 2 — Apache-2.0; cctorch: BSD-3-Clause. Native source and official checkpoints; cctorch carries an additional BSD notice. Ultralytics-converted files require checking their distributor terms.原生源码与官方权重;cctorch 另附 BSD 声明。Ultralytics 转换文件还需核对分发方条款。 · GitHubGitHub · License/notice 1许可/声明 1 · License/notice 2许可/声明 2
- FoundationPose — NVIDIA custom source license. Comparison method source has a custom license; do not describe it as MIT or presume weights share the same grant.对照方法源码使用自定义许可;不能标为 MIT,也不能推定权重有相同授权。 · GitHubGitHub · License/notice 1许可/声明 1
- tensorrt-cu12 — Other/Proprietary License. Runtime dependency; dependencies and model/data assets retain their own licenses.运行依赖;依赖及模型/数据资源保留各自许可。 · GitHubGitHub · License/notice 1许可/声明 1 ↩ ↩
- YCBInEOAT data — Not separately verified / 未单独核实. Tracking data permission must be checked at the original archive; the comparison code license does not establish the data license.跟踪数据许可须在原始归档核实;对照代码的许可不能作为数据许可。 · GitHubGitHub · GitHubGitHub
- FastSAM · GitHubGitHub
Model weights and dataset/task assets may have separate terms. WAPR's source license does not replace them.模型权重、数据集与任务资源可能有独立条款;WAPR 源码许可不替代这些许可。