Inference and renderer scaling推理与渲染后端性能
More instances increase the number of pose hypotheses; more CAD categories also change the geometry submitted to the renderer. This experiment examines those two workload axes, first by varying instance count, then by holding 1000 instances fixed and distributing them across categories. It uses the release's SA6D-trained masked WAPR, SAPR and WBPS checkpoints. Every instance has twelve initial rotations, followed by three masked WAPR updates, two SAPR updates and independent-group WBPS selection.
实例数量增加,会扩大候选位姿的计算规模;CAD 类别增加,还会改变送入渲染器的几何数据。本实验分别考察这两个负载维度:先改变实例数量,再固定总计 1000 个实例、改变类别数。实验使用发布版中基于 SA6D 训练的带掩码 WAPR、SAPR 与 WBPS 权重。每个实例使用十二个初始旋转,随后执行三次带掩码 WAPR、两次 SAPR 与按实例独立分组的 WBPS 选择。
All curves report warmed RTX 5090 compute-stage medians. They exclude 2D detection, NMS and setup; the timing scope below defines the included work. Backends are labeled with their measured precision. On narrow screens, scroll figures horizontally; select a figure to open its scalable original.
所有曲线均为 RTX 5090 预热后的计算阶段中位耗时,不计 2D 检测、NMS 与准备工作;具体范围见下文计时说明。各后端同时标明实际测量精度。窄屏可横向滚动图表,点击图像可查看可缩放原图。
Instance count and computation time实例数量与计算耗时
The sampled instance counts are 1, 10, 100 and 1000. The left panel repeats one CAD. The right uses min(N, 100) distinct CADs, distributed evenly and interleaved within GPU chunks. The 100-CAD pool contains 30 T-LESS[11], 28 ITODD[12], 21 YCB-V[15] and 21 HB[16] models. They are centered, normalized to a 0.20 m bounding-box diagonal and assigned the same gray vertex color, retaining their original triangles. The observed image and region are repeated to isolate computation. This is not an accuracy evaluation or a claim that 1000 objects are visible in the frame.
实例数采样为 1、10、100、1000。左图重复同一 CAD;右图采用 min(N, 100) 个不同 CAD,将实例均分并交错放入 GPU 分块。100 个 CAD 包含 30 个 T-LESS[11]、28 个 ITODD[12]、21 个 YCB-V[15] 与 21 个 HB[16] 模型;统一居中、将包围盒对角线缩放为 0.20 m,并赋予相同灰色顶点颜色,保留原有三角面。实验重复使用观测图像与区域,以考察计算扩展,不评价精度,也不表示图中可见 1000 个物体。
Top: OpenGL. Bottom: nvdiffrast[5] CUDA. Each panel fixes the renderer and compares four inference backends. Points are measured medians; whiskers show the minimum and maximum timed calls. Time uses a linear axis starting at zero, shared across all four panels; the instance-count axis remains logarithmic and is labeled accordingly. The callouts give measured times and speedups at 100 and 1000 inputs. Changing precision and changing runtime are separate comparisons.
上排为 OpenGL,下排为 nvdiffrast[5] CUDA;每个子图固定渲染器,比较四种推理后端。数据点为实测中位数,误差线表示逐次计时的最小值至最大值;耗时纵轴采用从零开始的线性尺度,四幅子图共用相同范围;实例数横轴保留对数尺度并明确标注。图内列出 100 与 1000 个输入时的实测耗时及加速比。精度变化与执行后端变化需分别理解。
For TensorRT[8] with OpenGL and one repeated CAD, increasing from 100 to 1000 instances multiplies pose-stage time by 10.07. Above ten instances, a full chunk already contains 120 hypothesis rows; larger workloads mainly increase the number of chunks. Thus the nearly proportional growth reflects capacity-limited batching, rather than one unrestricted GPU batch containing all 12000 rows.
在 TensorRT[8]、OpenGL 与同一 CAD 条件下,实例数由 100 增至 1000,位姿阶段耗时增至 10.07 倍。十个实例已填满包含 120 条候选的分块;更大负载主要增加分块数量。因此,近似成比例的增长反映了容量受限的批量计算方式,而不是将全部 12000 条候选放入一个不受限的 GPU batch。
| Executor执行后端 | One CAD同一 CAD | 100 CAD categories100 类 CAD | |
|---|---|---|---|
| OpenGL | OpenGL | nvdiffrast[5] CUDA | |
| PyTorch[6] FP32 | 42.099 | 41.695 | 58.285 |
| ONNX Runtime[9] CUDA FP32 | 47.036 | 46.702 | 63.099 |
| PyTorch AMP FP16 | 22.099 | 21.790 | 38.425 |
| TensorRT[8] FP16 | 11.497 | 11.170 | 27.936 |
For 1000 instances across 100 CADs with OpenGL fixed, TensorRT's pose-stage speedup is 3.73× relative to PyTorch[6] FP32, 1.95× relative to PyTorch AMP FP16 and 4.18× relative to ONNX Runtime[9] FP32. The FP32-to-FP16 comparison includes a precision change; the AMP comparison gives a closer indication of the runtime difference at reduced precision. These are compute-stage ratios for this workload, not end-to-end detection speedups.
固定 OpenGL、使用 100 类 CAD 的 1000 个实例时,TensorRT 位姿阶段相对 PyTorch[6] FP32、PyTorch AMP FP16、ONNX Runtime[9] FP32 的加速分别为 3.73 倍、1.95 倍和 4.18 倍。FP32 与 FP16 的对照包含精度变化;AMP 对照更接近低精度条件下的执行后端差异。这些比值对应本组负载的计算阶段,不是完整检测流程的加速比。
With 1000 instances across 100 CADs and OpenGL fixed, ONNX Runtime FP32 takes 1.12× the time of native PyTorch FP32. Exporting to ONNX does not itself guarantee lower latency: execution kernels, graph optimization and precision still matter. For this workload, the measured TensorRT FP16 path is the fastest of the four configurations.
固定 OpenGL、使用 100 类 CAD 的 1000 个实例时,ONNX Runtime FP32 的耗时为PyTorch FP32 的 1.12 倍。导出 ONNX 本身并不保证延迟降低,执行算子、图优化及数值精度仍会影响结果。本组负载中,实测 TensorRT FP16 路径在四种配置中耗时最少。
One RGB/depth render pass at 160×160 pixels per hypothesis, without neural inference. For 1000 instances and 100 CADs: OpenGL 274.3 ms, nvdiffrast 3578.3 ms. All required capacity chunks are included. Time uses a shared linear axis starting at zero; instance count uses a labeled logarithmic axis. Callouts show measured times and renderer speedups at 1000 inputs. Whiskers span the minimum and maximum timed calls.
每个候选输出 160×160 像素 RGB/深度,测量一次渲染,不含网络推理。100 类 CAD 的 1000 个实例:OpenGL 274.3 ms,nvdiffrast 3578.3 ms。包含全部容量分块;耗时纵轴共用从零开始的线性尺度,数量横轴为明确标注的对数尺度。图内列出 1000 个输入时的实测耗时与渲染加速比;误差线表示逐次计时的最小值至最大值。
For 1000 instances across 100 CADs, switching from PyTorch AMP to TensorRT saves 10.62 s with OpenGL and 10.49 s with nvdiffrast. The full-stage speedups nevertheless differ: 1.95× and 1.38×. Similar absolute savings produce a smaller ratio when the remaining stage costs are higher. This is why the inference-backend curves and renderer curves should be read together.
在 100 类 CAD 的 1000 个实例上,由 PyTorch AMP 改用 TensorRT,OpenGL 路径节省 10.62 秒,nvdiffrast 路径节省 10.49 秒;整体加速比分别为 1.95 倍与 1.38 倍。绝对节省的时间相近,但其余环节耗时较高时,加速比会较小。因此,推理后端与渲染器曲线需要结合解读。
Changing categories with 1000 instances fixed固定 1000 个实例,改变类别数量
The category counts are 1, 10, 50 and 100, corresponding to 1000, 100, 20 and 10 instances per category. All points therefore process the same 12000 initial hypothesis rows. Neural input size stays fixed, while mixed-mesh packing and submitted triangle counts may change. The category pools are prefixes of the same model list; different CADs retain different triangle counts. A change in these curves should therefore be interpreted together with geometry complexity, rather than attributed to category count alone.
类别数为 1、10、50、100,每类分别包含 1000、100、20、10 个实例;各点均处理 12000 条初始候选位姿。网络输入规模不变,但混合网格打包与提交的三角面数量可能变化。各组使用同一模型列表的前缀,不同 CAD 保留不同的面数。因此,曲线变化需要结合几何复杂度解释,不能全部归因于类别数量。
The first CAD has 84260 triangles, while the 100-category pool averages about 14279 per CAD. Thus this category sweep also changes mean geometry complexity. Backend comparisons at the same point share exactly the same CADs and hypothesis rows. The renderer curves characterize the provided implementations, including their mesh packing and memory costs, rather than an intrinsic limit of OpenGL or nvdiffrast.
第一类 CAD 含 84260 个三角面,100 类模型平均每个约 14279 个;因此类别扫描也改变了平均几何复杂度。同一数据点上的后端对照使用完全相同的 CAD 与候选行。渲染曲线反映本项目现有实现的网格打包与内存开销,不代表 OpenGL 或 nvdiffrast 的固有性能上限。
The provided nvdiffrast path also changes representation: instanced mode shares indexed vertices for one CAD, whereas the mixed-CAD range path expands triangle corners for each hypothesis and transforms their coordinates and attributes on the GPU. Mixed geometry is cached only for the current chunk layout; alternating layouts also incur repacking. These different memory workloads help explain why the single-CAD and mixed-CAD curves need separate interpretation. This experiment measures both paths as provided; it does not claim that category count alone explains their timing.
现有 nvdiffrast 路径还会改变几何表示:同一 CAD 的实例模式共享带索引的顶点;多 CAD 的 range 路径则按每条候选展开三角形角点,并在 GPU 上变换坐标与属性。混合几何仅缓存当前分块布局,布局交替时还会计入重新打包的成本。两种路径的内存负载不同,因此同一 CAD 与混合 CAD 曲线需要分别解读。本实验测量项目提供的实现,不将它们的耗时差异全部归因于类别数量。
Top: pose-stage medians with each renderer fixed. Bottom left: one-pass render medians. Bottom right: submitted CAD triangles, showing the geometry change accompanying this category sweep. Both axes use linear scales, with all vertical axes starting at zero. All sampled points retain 1000 instances and twelve hypotheses per instance; timing whiskers span the minimum and maximum calls.
上排分别固定两种渲染器,比较位姿阶段中位耗时;左下为单次渲染中位耗时,右下为提交的 CAD 面数,用于观察类别扫描伴随的几何规模变化。横纵轴均采用线性尺度,各纵轴从零开始。所有采样点均保持 1000 个实例、每实例十二个候选;计时误差线表示最小值至最大值。
| CAD categoriesCAD 类别 | Instances per category每类实例数 | TRT + OGL · s | TRT + nvdiffrast[5] · s | OGL render · msOGL 渲染 · ms | nvdiffrast render · msnvdiffrast[5] 渲染 · ms |
|---|---|---|---|---|---|
| 1 | 1000 | 11.497 | 25.559 | 337.9 | 2661.3 |
| 10 | 100 | 11.179 | 52.540 | 274.0 | 7173.0 |
| 50 | 20 | 11.168 | 30.148 | 275.2 | 3953.9 |
| 100 | 10 | 11.170 | 27.936 | 274.3 | 3578.3 |
Backend settings and timing scope后端设置与计时范围
At 100 instances, OpenGL + TensorRT FP16 takes 1.142 s for one repeated CAD and 1.108 s for 100 distinct CADs: 87.6 and 90.2 instances/s, respectively. These rates count input instances, each with twelve hypotheses, three masked WAPR updates, two SAPR updates and WBPS. They describe the controlled pose stage, rather than successful detections per second or a 100-object frame's complete latency. The paper's reported 25 instances/s uses its own evaluation setting; the rates alone do not establish its inference backend or a matched speedup.
100 个实例时,OpenGL + TensorRT FP16 对同一 CAD 和 100 个不同 CAD 分别耗时 1.142 秒与 1.108 秒,折合每秒 87.6 与 90.2 个实例。吞吐量按输入实例计数,每实例包含十二个候选、三次带掩码 WAPR、两次 SAPR 与 WBPS。这对应控制实验的位姿阶段,不能等同于每秒正确检测数量或含 100 个物体图像的完整处理延迟。论文报告的每秒 25 个实例属于其评测设定,仅凭数值不能确认论文执行后端,也不能形成同条件加速比。
All configurations run sequentially on the same RTX 5090 under Python 3.10. PyTorch FP32 and ONNX Runtime CUDA FP32 disable TF32; PyTorch AMP FP16 provides a reduced-precision reference for the TensorRT FP16 engines. ONNX is a model format, so the measured executor is ONNX Runtime's CUDA Execution Provider, using GPU I/O binding rather than copying crops through NumPy[7]. Geometry and pose updates remain FP32 for every configuration. The checkpoint weights and twelve-hypothesis groups are unchanged.
所有配置在同一张 RTX 5090、Python 3.10 环境中依次运行。PyTorch FP32 与 ONNX Runtime CUDA FP32 关闭 TF32;另测 PyTorch AMP FP16,作为 TensorRT FP16 引擎的低精度对照。ONNX 是模型格式,因此这里测量的是 ONNX Runtime 的 CUDA 执行后端,并通过 GPU I/O 绑定传递裁剪张量,避免经 NumPy[7] 往返复制。所有配置的几何计算与位姿更新均保持 FP32,模型权重与每实例十二个候选姿态的分组不变。
PyTorch enables cuDNN benchmarking. ONNX Runtime uses full graph optimization and HEURISTIC convolution search. A separate 120-row check compares it with EXHAUSTIVE search, with three warmups and ten timed calls for each network. The three networks' warmed median times differ by less than one percent, and their outputs match in that check. Algorithm search itself is outside timing.
PyTorch 启用 cuDNN 算法基准搜索;ONNX Runtime 使用完整图优化与 HEURISTIC 卷积搜索。另以 120 条候选行,对每个网络预热三次、计时十次,检查它与 EXHAUSTIVE 搜索的差异。三个网络的预热中位耗时差异均小于百分之一,输出在该检查中一致。算法搜索过程不计时。
The timed pose stage includes crop construction, rendering, refinement, scoring, selection and transfer of final outputs. The RGB-D frame, its xyz map, CADs and initial region/depth seed are prepared before timing. It excludes 2D detection, NMS, loading, compilation and downloads. Each workload discards two complete warmup calls and records three CUDA-synchronized calls. The separate renderer measurement draws the same initial hypotheses once, including crop-window construction, with three warmups and ten timed calls; it is not the sum of all rendering inside pose estimation.
位姿阶段计时包含裁剪构造、渲染、修正、评分、候选选择及最终结果传回。RGB-D 图像、xyz 图、CAD 和初始区域对应的深度平移在计时前准备;不计 2D 检测、NMS、加载、编译与下载。每种负载舍弃两次完整预热,记录三次 CUDA 同步调用。纯渲染实验则将同一批初始候选姿态渲染一次,包含裁剪窗口构造,预热三次、计时十次;它并非位姿流程内全部渲染耗时之和。
Both renderers use GPU batching. OpenGL submits mixed CADs through the release's tile renderer; nvdiffrast uses RasterizeCudaContext with instanced mode for one CAD and range mode for mixed geometry. This is supported by nvdiffrast's batching API. Chunks contain at most 120 hypothesis rows, matching the current engine capacity: ten instances per full chunk, or 100 chunks for 1000 instances. A necessary capacity loop does not turn each object into a separate network call.
两种渲染器均采用 GPU 批量计算。OpenGL 使用发布版的分块渲染器提交混合 CAD;nvdiffrast 使用 CUDA 光栅上下文,对同一 CAD 使用实例模式,对混合几何使用 range 模式,见其批量渲染接口。每块最多 120 条候选位姿,与现有引擎容量一致:满块包含十个实例,1000 个实例分为 100 块。容量分块所需的循环不等同于逐物体调用网络。
Relating batch throughput to application latency批量吞吐量与应用耗时的关系
The release's WAPR estimator defaults to TensorRT FP16 and renders through OpenGL. The LM-O[10] and ROBI[13] measurements prepare resident CAD meshes before warmup. Their hot pose stages include frame transfer, mask/translation preparation and one observed xyz map per frame. The controlled backend measurements start with GPU-resident RGB-D, xyz and initial translations. A shared crop/output check establishes numerical agreement on its tested inputs. Read each timing with its stated workload and transfer scope.
发布版 WAPR 估计器默认使用 TensorRT FP16,并通过 OpenGL 渲染。LM-O[10] 与 ROBI[13] 测速均在预热前准备并驻留 CAD 网格;预热后的位姿阶段包含图像传输、实例掩码与初始平移准备,以及每帧一次观测 xyz 计算。后端控制实验从 GPU 上已准备的 RGB-D、xyz 与初始平移开始。裁剪与输出一致性检查验证所测输入的数值一致;阅读各项耗时需同时查看其负载与传输范围。
| Workload负载 | Pose inputs位姿输入数 | Pose stage位姿阶段 | Measured total已测阶段合计 |
|---|---|---|---|
| Controlled, one CAD控制实验,同一 CAD | 10 | 0.115 | — |
| Controlled, 100 CADs控制实验,100 类 CAD | 100 | 1.108 | — |
| LM-O[10], frame 307LM-O[10],第 307 帧 | 10 | 0.134 | 0.419* |
| ROBI[13] Zigzag, mixed proposalsROBI[13] Zigzag,混合候选 | 92 | 1.856 | 2.475 |
| ROBI screws, mixed proposalsROBI 螺丝,混合候选 | 101 | 2.003 | 2.700 |
| IC-BIN[14], frame 26IC-BIN[14],第 26 帧 | 42 | 0.596 | 0.899* |
| Custom white wedges自定义白色三角块 | 10 | 0.130 | 0.324* |
* LM-O, IC-BIN[14] and custom wedges measure detection + pose, without final pose NMS in these totals. IC-BIN and custom wedges each discard three complete warmup calls and record ten CUDA-synchronized calls with prepared, resident meshes; their hot mesh preparation/packing/upload counters remain unchanged. Dashes indicate an unmeasured stage. All rows use twelve hypotheses per input, but CAD geometry, masks, resolution and input preparation differ. Pose-input counts describe processed regions, not the number of visible objects or final retained predictions; a WBPS display threshold does not reduce work already performed. Stage medians need not add to the median of complete calls.
* LM-O、IC-BIN[14] 与三角块记录检测 + 位姿,合计均不含最终位姿 NMS。IC-BIN 和三角块各舍弃三次完整预热调用,再记录十次 CUDA 同步调用;网格提前准备并驻留,计时中准备、打包、上传计数均不增加。横线表示该阶段未测。各行每输入均使用十二个候选,但 CAD 几何、掩码、分辨率及输入准备不同。位姿输入数是参与计算的区域数,不是可见物体数或最终保留的预测数;提高 WBPS 展示阈值不会减少已经完成的计算。各阶段中位数也不必等于完整调用中位数的分解。
The case measurements reuse prepared, resident CAD meshes and one observed XYZ map per frame. Centering, normals, mesh packing and GL upload are outside the hot timer. Mesh preparation/packing/upload counters do not increase in the timed calls. For 67 screw regions with 4×4 hypotheses, the 6D-stage median is 1.824 s. The 101-input 4×3 row above uses live mixed proposals and is a different workload.
案例测速复用准备好并驻留的 CAD,每帧共用一份观测 XYZ 图。居中、法线、网格打包和 GL 上传在预热后的计时外,计时调用中这三项网格计数均不增加。67 个螺丝区域采用 4×4 候选配置时,6D 阶段中位数为 1.824 秒。上方 101 输入、4×3 一行使用现场混合候选,是另一种负载。
The OpenGL + TensorRT label applies to WAPR pose computation. FoundationPose[4], DINOv2[2] mask tracking, GroundingDINO[1]/SAM2[3], reconstruction and CPU baselines have their own execution paths. Backend and timing scope are specified for each result; the WAPR backend label does not describe these other components.
OpenGL + TensorRT 标识对应 WAPR 位姿计算。FoundationPose[4]、DINOv2[2] 掩码跟踪、GroundingDINO[1]/SAM2[3]、重建与 CPU 基线分别具有自己的执行路径。各项结果分别注明后端与计时范围,WAPR 的后端标识不代表这些组件的执行方式。
Experiment conditions实验条件
The isolated experiment preserves the release's crop, normalization and update rules. Its OpenGL/TensorRT crops, final poses and scores match the released estimator exactly on the one- and ten-instance checks. ONNX FP32 outputs are checked against the checkpoint modules at both group sizes. The renderer check uses ten actual CADs: silhouette IoU 0.9994, shared-valid-pixel depth MAE 0.011 mm. Boundary pixels can differ between rasterizers; this check does not establish dataset-level accuracy equivalence.
独立实验保持发布版的裁剪、归一化与更新规则。在一实例及十实例检查中,其 OpenGL/TensorRT 裁剪、最终位姿和分数与正式估计器一致;ONNX FP32 输出也在两种组数下与权重模块对照。渲染检查使用十个实际 CAD:轮廓 IoU 为 0.9994,共同有效像素的深度平均绝对差为 0.011 mm。两种光栅器的边界像素可能不同;该检查不代表整个数据集上的精度等价。
The configurations use a Python 3.10 CUDA environment on one RTX 5090. Each backend runs sequentially with the same input frame, CADs, twelve-hypothesis groups and update counts. The measurements use FP32 ONNX graphs and GPU-specific TensorRT FP16 engines, including grouped WBPS. The published estimator uses OpenGL + TensorRT; the ONNX Runtime and nvdiffrast curves describe the backend comparison experiment.
各配置在单张 RTX 5090、Python 3.10 CUDA 环境中依次运行,使用相同输入帧、CAD、每实例十二个候选以及相同的更新次数。实验使用 FP32 ONNX 图与适配 GPU 的 TensorRT FP16 引擎,包含按实例分组的 WBPS。公开估计器使用 OpenGL + TensorRT;ONNX Runtime 与 nvdiffrast 曲线对应后端对照实验。
Measured software: PyTorch 2.8.0+cu128, CUDA 12.8, ONNX Runtime 1.23.2, TensorRT 10.9.0.34 and nvdiffrast 0.3.3.1. The CPU is an Intel Xeon Platinum 8470Q with four CPU threads used for the experiment. The RTX 5090 has 32607 MiB VRAM and driver 595.71.05. The plotted values are medians of warmed, CUDA-synchronized calls with the workload and timing scope specified above.
实测软件为 PyTorch 2.8.0+cu128、CUDA 12.8、ONNX Runtime 1.23.2、TensorRT 10.9.0.34 和 nvdiffrast 0.3.3.1。CPU 为 Intel Xeon Platinum 8470Q,实验使用四个 CPU 线程;RTX 5090 显存为 32607 MiB,驱动为 595.71.05。图中数值为预热后 CUDA 同步调用的中位数,负载与计时范围见上方说明。
References and licenses参考文献与许可
- GroundingDINO — Apache-2.0. Vendored detector source; original notices retained.随包检测源码;保留原始声明。 · GitHubGitHub · License/notice 1许可/声明 1 · License/notice 2许可/声明 2
- DINOv2 — Apache-2.0. Code and official DINOv2 weights; retain copyright and license.源码与官方 DINOv2 权重;保留版权和许可。 · GitHubGitHub · License/notice 1许可/声明 1
- SAM 2 — Apache-2.0; cctorch: BSD-3-Clause. Native source and official checkpoints; cctorch carries an additional BSD notice. Ultralytics-converted files require checking their distributor terms.原生源码与官方权重;cctorch 另附 BSD 声明。Ultralytics 转换文件还需核对分发方条款。 · GitHubGitHub · License/notice 1许可/声明 1 · License/notice 2许可/声明 2
- FoundationPose — NVIDIA custom source license. Comparison method source has a custom license; do not describe it as MIT or presume weights share the same grant.对照方法源码使用自定义许可;不能标为 MIT,也不能推定权重有相同授权。 · GitHubGitHub · License/notice 1许可/声明 1
- nvdiffrast — NVIDIA Source Code License (custom). Optional renderer benchmark dependency; consult the upstream source license.可选渲染后端实验依赖;以其上游源码许可为准。 · GitHubGitHub · License/notice 1许可/声明 1
- torch — BSD License. Runtime dependency; dependencies and model/data assets retain their own licenses.运行依赖;依赖及模型/数据资源保留各自许可。 · Original source原始来源 · License/notice 1许可/声明 1 · License/notice 2许可/声明 2 ↩ ↩ ↩
- numpy — BSD License. Runtime dependency; dependencies and model/data assets retain their own licenses.运行依赖;依赖及模型/数据资源保留各自许可。 · Original source原始来源 · License/notice 1许可/声明 1 ↩ ↩
- tensorrt-cu12 — Other/Proprietary License. Runtime dependency; dependencies and model/data assets retain their own licenses.运行依赖;依赖及模型/数据资源保留各自许可。 · GitHubGitHub · License/notice 1许可/声明 1 ↩ ↩ ↩
- onnxruntime-gpu — MIT. Optional runtime, optional inference runtime.可选推理后端,可选推理后端。 · Original source原始来源 · License/notice 1许可/声明 1 ↩ ↩ ↩
- LM-O data — CC-BY-SA-4.0. Data and model excerpt, not the loader code.数据与模型摘录,不是读取代码。 · Original source原始来源 · GitHub: BOP toolkitGitHub:BOP 工具集
- T-LESS data — CC-BY-4.0. Data and object models; attribute the original dataset.数据与物体模型;须标注原始数据集。 · Original source原始来源
- ITODD data — CC-BY-NC-SA-4.0. Non-commercial dataset; WAPR permissions do not replace dataset terms.非商业数据集;WAPR 授权不替代数据条款。 · Original source原始来源
- ROBI data — Not separately verified / 未单独核实. The saved public poses and dataset are credited to ROBI; an independent grant to redistribute them has not been verified.保存的公开位姿和数据均注明 ROBI 来源;其再分发授权尚未独立核实。 · GitHubGitHub
- IC-BIN · Official BOP dataset pageBOP 官方数据页
- YCB-Video (YCB-V) · Official BOP dataset pageBOP 官方数据页 · GitHub: PoseCNNGitHub:PoseCNN
- HomebrewedDB (HB) · Official BOP dataset pageBOP 官方数据页
Model weights and dataset/task assets may have separate terms. WAPR's source license does not replace them.模型权重、数据集与任务资源可能有独立条款;WAPR 源码许可不替代这些许可。