CAD templates and feature encodingCAD 模板与特征编码
A CAD template bank associates object identities and rendering viewpoints with template images and their visual features. An encoder maps template images and observed object regions into a comparable feature space for CAD identity matching. Templates and queries must use the same encoder and consistent preprocessing. The current implementation uses DINOv2[2] ViT-L/14 for this encoding stage.
CAD 模板库关联物体类别、渲染视角、模板图像及其视觉特征。编码器将模板图像与观测中的物体区域映射到可比较的特征空间,用于后续 CAD 类别匹配。模板与查询区域需使用相同编码器和一致的预处理;当前实现采用 DINOv2[2] ViT-L/14 完成这一编码阶段。
Template generation and preprocessing模板生成与预处理
onboard_meshes in wapr/det2d.py builds the template bank. By default it stays on the GPU: the rendered views and the DINOv2 features are not written to disk. examples/05_bop_6d_detection.py leaves template_path empty, so each run renders and encodes into that process. Set template_path to outputs/cache/det2d/lmo.pt only when a later run should load the file instead of rendering again. The input is the eight LM-O[6] meshes in samples/bop/lmo/models/. Those ply files are millimeters. The reader multiplies by 0.001, so the vertices that enter onboard_meshes are meters. The object ids are 1, 5, 6, 8, 9, 10, 11, 12. The counts below are the variables at the top of wapr/det2d.py, used in this order.
wapr/det2d.py 中的 onboard_meshes 构建模板库。默认将渲染视图与编码特征保留在显存,不写入磁盘。examples/05_bop_6d_detection.py 将 template_path 留空,每次运行时生成模板;如需跨进程复用,可设置为 outputs/cache/det2d/lmo.pt。示例输入为 samples/bop/lmo/models/ 中的八份 LM-O[6] 网格,物体编号为 1、5、6、8、9、10、11、12。PLY 顶点以毫米计,读取后乘以 0.001 转换为米。以下说明采用 wapr/det2d.py 中的模板生成参数。
Each CAD mesh is centered and rendered from template_views = 42 Fibonacci-sphere viewpoints with four in-plane rotations, producing 168 templates per object. The camera is placed relative to the mesh diameter so objects of different physical sizes occupy comparable crops. Object pixels come from the rendered depth mask. Camera and crop choices are in wapr/det2d.py; lighting choices are in wapr/recipe.py.
各 CAD 网格先居中,再从 template_views = 42 个斐波那契球面视角及四种面内旋转渲染,每个物体得到 168 张模板。相机位置相对网格直径设定,使物理尺寸不同的物体在裁剪图中具有可比尺度;目标像素由渲染深度掩码确定。相机与裁剪设置见 wapr/det2d.py,光照设置见 wapr/recipe.py。
Each rendered view is then cropped. The crop is a square around the mask box. The square’s edge is the longer side of that box times crop_margin, and crop_margin is 1.1. Bilinear sampling resizes the square to 224×224. RGB outside the mask is black. inplane_turns is [0, 1, 2, 3]: that square is rotated by 0°, 90°, 180°, and 270°. One object therefore has 42 × 4 = 168 images. The sampling is this Fibonacci sphere. It is not a regular polyhedron.
每一张渲染图再裁一次。裁剪是包住 mask 包围盒的一个正方形。正方形的边长是这个包围盒较长的一边乘 crop_margin,crop_margin 是 1.1。双线性采样把这个正方形缩到 224×224。mask 外面的 RGB 是黑的。inplane_turns 是 [0, 1, 2, 3]:这个正方形再转 0°、90°、180°、270°。一个物体因此有 42 × 4 = 168 张图。采样用的是这个斐波那契球面,不是正多面体。
Left, the 42 camera positions on the unit sphere. The color is the view index. The black dot is the mesh center. Right, object 8, the driller, at the view whose mask has the most pixels. The four squares are rot90 of that one crop: 0°, 90°, 180°, 270°. They are not four new camera orbits.
左边是单位球面上的 42 个相机位置。颜色是视角序号。黑点是网格中心。右边是物体 8,电钻 driller,取 mask 像素最多的那个视角。四个正方形是这一张裁剪的 rot90:0°、90°、180°、270°。不是四次新的相机环绕。
Black-surface template rendering黑色表面的模板渲染
The background of a template image is black because RGB outside the mask is set to 0. The mask itself comes from depth, so the outline is still there when the surface is dark. What happens to the surface color is decided in wapr/ogl.py before the picture is encoded.
模板图的背景是黑的,因为 mask 外面的 RGB 被置成 0。mask 来自深度,所以表面很暗时,轮廓还在。表面颜色怎么处理,在编码之前就由 wapr/ogl.py 定下来。
A mesh with no usable color is painted gray. Missing vertex color, a uniform black vertex color, and the constant 102/255 that trimesh[4] writes when no color exists are placeholders. DEFAULT_UNTEXTURED_VERTEX_COLOR is 0.5, so every vertex becomes that gray. A texture that has no UV is first replaced by its mean color; a black mean is the same uniform black, and it becomes gray too. The light is base × 0.8 + base × max(n · light, 0) × 0.2, so the gray still shows the surface normals. On the ape view below, the brightest pixel inside the mask is 127.
没有可用颜色的网格会被涂成灰。缺失的顶点色、整片纯黑的顶点色,以及 trimesh[4] 在没有颜色时写下的常数 102/255,都是占位。DEFAULT_UNTEXTURED_VERTEX_COLOR 是 0.5,所以每个顶点都变成这个灰。没有 UV 的纹理先换成它的平均色;平均色是黑的,就和整片纯黑一样,也会变成灰。光照是 base × 0.8 + base × max(n · light, 0) × 0.2,所以这个灰仍然能看出表面法向。下面这张猩猩的视角里,mask 内最亮的像素是 127。
A black texture that has UV is kept by the renderer. base is 0, so that render is black: the brightest pixel inside the mask is 0, the same black as the background. That black image is not encoded. When the foreground maximum is below black_foreground_max (0.01), onboard_meshes replaces it with gray shading from the depth of that same view. The albedo is DEFAULT_UNTEXTURED_VERTEX_COLOR, 0.5. The light is the same formula, base × 0.8 + base × max(n · light, 0) × 0.2, and the normal comes from the depth. A front face is 0.5. The right panel is that replacement. DINOv2 encodes it, and the bank is written from it. A view already brighter than 0.01 is encoded as rendered. The pose renderer in wapr/ogl.py is not changed.
带 UV 的黑色纹理,渲染器会保留。base 是 0,所以那张渲染是黑的:mask 内最亮像素是 0,和背景一样黑。这张黑图不拿去编码。前景最大值低于 black_foreground_max(0.01)时,onboard_meshes 把它换成同一视角深度算出的灰色明暗。反照率是 DEFAULT_UNTEXTURED_VERTEX_COLOR,0.5。光照是同一条公式,base × 0.8 + base × max(n · light, 0) × 0.2,法向从深度来。正对相机的面是 0.5。右图就是这张替换后的图。DINOv2 编码它,库也按它来写。已经亮过 0.01 的视角按渲染结果编码。位姿渲染 wapr/ogl.py 不改。
A query crop is handled on the same rule, after the mask has zeroed the background. If the RGB inside the mask is also below 0.01, those pixels are painted 0.5, so a pure-black crop is a gray silhouette on black, and it is not a constant-black embedding. A photograph whose masked pixels rise above 0.01 keeps its RGB. That photo is matched to the shaded gray template, which is the same pairing already used for a mesh that had no color. The background stays black on both sides. Normals are not estimated from the photograph, and the template background is not turned white.
查询裁剪用同一条规则,在 mask 把背景置黑之后。mask 内的 RGB 也低于 0.01 时,这些像素涂成 0.5,所以一块纯黑裁剪是黑底上的灰色剪影,不是一张恒定黑图的特征。照片里 mask 内的像素亮过 0.01,就保留这张 RGB。这张照片去和灰色明暗的模板匹配,和本来就没有颜色的网格是同一种配对。两边的背景都保持黑色。不从照片估计法向,也不把模板背景改成白色。
Object 1, the ape, one of the 42 views. Each panel is the image that is encoded. Left, the color stored on the mesh. The brightest pixel inside the mask is 171. Middle, the vertex color was set to uniform black, and the loader painted it gray 0.5. The brightest pixel is 127. Right, a black texture with UV. The renderer returns black, and that black image is not encoded. The panel is the image DINOv2 receives: gray shading from this view's depth, albedo 0.5, the same light. The brightest pixel inside the mask is 127.
物体 1,猩猩 ape,42 个视角里的一个。每一张都是送去编码的图。左,网格上原来的颜色。mask 内最亮像素是 171。中,顶点色被设成整片纯黑,加载时涂成灰 0.5。最亮像素是 127。右,带 UV 的黑色纹理。渲染器返回全黑,那张黑图不编码。这一张是 DINOv2 收到的图:由这个视角的深度算出的灰色明暗,反照率 0.5,光照相同。mask 内最亮像素是 127。
Feature representation and encoding特征表示与编码
The combination of DINOv2 CLS features and GeM[8]-pooled patch features follows MUSE, Section 3.2[7] (Cho, Park and Oh, 2025), including the pooling exponent of 1.5. GeM pooling itself originates from Radenović, Tolias and Chum, Fine-tuning CNN Image Retrieval with No Human Annotation; MUSE applies it to the integration of features for model-based object detection and segmentation.
这里将 DINOv2 的 CLS 特征与 GeM[8] 汇聚后的图块特征结合,参考 MUSE 第 3.2 节[7](Cho、Park 和 Oh,2025),池化指数同样采用 1.5。GeM 池化本身源自 Radenović、Tolias 和 Chum 的 Fine-tuning CNN Image Retrieval with No Human Annotation;MUSE 将它用于基于模型的物体检测与分割中的特征组合。
In the current implementation, the 168 rendered templates are encoded by NativeDino with DINOv2 ViT-L/14. The template bank is built in PyTorch[3]; a TensorRT[5] engine can later encode query crops. A 224-pixel crop becomes a 16×16 patch grid, and patches mostly outside the object mask are excluded from matching. Normalization and coverage thresholds are defined in wapr/det2d.py.
当前实现中,每个物体的 168 张渲染模板由 NativeDino 使用 DINOv2 ViT-L/14 编码。模板库在 PyTorch[3] 中构建;后续查询裁剪可使用 TensorRT[5] 引擎。224 像素裁剪形成 16×16 图块网格,主要落在目标掩码之外的图块不参与匹配。归一化和覆盖率阈值见 wapr/det2d.py。
One image becomes four tensors. cls is the CLS token, float32, shape (1024,). gem pools the 256 patch tokens with GeM, exponent 1.5, and is also float32, shape (1024,). patches stores each patch token after L2 normalization, in float16, shape (256, 1024). An invalid patch is stored as zeros. valid is the bool mask of those 256 patches, shape (256,).
每张模板的编码结果包含四个张量。cls 是 CLS token,float32,形状 (1024,)。gem 把 256 个图块 token 做 GeM,指数 1.5,也是 float32,形状 (1024,)。patches 存 L2 归一化之后的图块 token,float16,形状 (256, 1024)。无效图块存成 0。valid 是这 256 个图块的布尔掩码,形状 (256,)。
Object 8, the driller, ViT-L/14, the view whose mask is largest. The crop is the 224 image that is encoded. Valid patches are the 16×16 cells where the mask covers more than half. Patch PCA paints the 1024-d token at each cell with its first three components, so the handle and the body come out as different colors. The last panel is the cosine of every patch token with the patch nearest the center.
物体 8,电钻,ViT-L/14,mask 最大的那个视角。裁剪就是被编码的那张 224 的图。有效图块是 16×16 里 mask 盖过一半的格子。图块 PCA 把每个格子上的 1024 维 token 画成它的前三个主成分,所以手柄和机身是不同的颜色。最后一块是每个图块 token 和最靠近中心的那个图块的余弦。
The two stored global vectors, on the eight meshes at this same camera. Color encodes cosine similarity from 0 to 1; darker blue is higher. The diagonal is each object with itself. Mean cosine between different objects is 0.283 for CLS and 0.901 for GeM.
存下来的两个全局向量,八个物体网格模型,同一台相机。颜色表示 0 到 1 的余弦相似度,蓝色越深越高。对角线是物体和自己。不同物体之间的平均余弦,CLS 是 0.283,GeM 是 0.901。
Current encoder configuration当前编码器配置
The published configuration uses DINOv2 ViT-L/14 with CLS and GeM features. ViT-S/14 and ViT-B/14 are available choices, but changing the encoder also requires a matching weight, template bank, and, for TensorRT, engine. The figure below compares features on one camera view; it does not establish which model has the highest detection accuracy. The exact model dimensions remain in the encoder implementation.
已公布的配置使用 DINOv2 ViT-L/14 的 CLS 与 GeM 特征。ViT-S/14 和 ViT-B/14 也是可选项,但更换编码器时必须配套更换权重、模板库,以及 TensorRT 引擎。下图只比较同一相机视角的特征,不能据此判断哪个模型的检测精度最高;具体模型维度见编码器实现。
The same driller crop and the same eight renders. Each PCA has its own three axes, so the hues are not a shared legend. The cosine color is shared, 0 to 1. Mean CLS cosine between different objects on this camera is 0.355 for S, 0.287 for B, and 0.283 for L. This is that one camera, not a detection-AP table.
同一张电钻裁剪,同一批八个渲染。每个 PCA 用自己的三个轴,所以颜色不能互相对着看。余弦的颜色是共用的,0 到 1。这一台上,不同物体之间的 CLS 平均余弦,S 是 0.355,B 是 0.287,L 是 0.283。这是这一台相机,不是检测 AP 表。
torch.stack stacks the eight objects. That dictionary is the bank, and it stays on the GPU. torch.save runs only when cache_path is set, and the file is replaced through a temporary file. LM-O has N = 8 objects and 168 views, so the bank is:
torch.stack 把八个物体叠起来。这个字典就是库,留在显存里。只有设置了 cache_path 才调用 torch.save,文件经过一个临时文件再替换。LM-O 是 N = 8 个物体、168 个视角,所以库是:
| Key键 | Shape形状 | Dtype类型 |
|---|---|---|
cls | (8, 168, 1024) | float32 |
gem | (8, 168, 1024) | float32 |
patches | (8, 168, 256, 1024) | float16 |
valid | (8, 168, 256) | bool |
obj_ids | (8,) | int64 |
obj_ids is those eight ids in sorted order. patches is most of the bank. When a .pt file is written, outputs/cache/det2d/lmo.pt.json beside it records the bank hash, the DINOv2 weight hash, the object ids, and views = 168. WAPRDet2D checks those hashes when it loads that file.
obj_ids 是这八个编号,按从小到大排列。patches 占了库的绝大部分。如果写出了 .pt,旁边的 outputs/cache/det2d/lmo.pt.json 记下库的哈希、DINOv2 权重的哈希、物体编号,以及 views = 168。WAPRDet2D 加载这个文件时核对这些哈希。
Matching starts from the RGB crop inside a SAM mask, with the background set to black. DINOv2 compares its CLS, GeM, and valid patch features with the CAD views. Fused similarity is combined with GroundingDINO[1] objectness, then confidence filters low-scoring rows; obj_ids can restrict eligible categories before this choice. Visible-mask overlap suppresses duplicates within each category, while allowing several instances of one category. The score weights and suppression threshold are explicit in wapr/det2d.py and its recipe.
匹配输入是 SAM 掩码内的 RGB 裁剪,背景置黑。DINOv2 将其 CLS、GeM 和有效图块特征与 CAD 视角比较;融合相似度再结合 GroundingDINO[1] 的目标分数,由 confidence 过滤低分结果。obj_ids 可在选择前限定类别。可见掩码的重叠度用于同类别去重,但同一类别仍可保留多个实例。分数权重与抑制阈值保留在 wapr/det2d.py 及其配置中。
References and licenses参考文献与许可
- GroundingDINO — Apache-2.0. Vendored detector source; original notices retained.随包检测源码;保留原始声明。 · GitHubGitHub · License/notice 1许可/声明 1 · License/notice 2许可/声明 2
- DINOv2 — Apache-2.0. Code and official DINOv2 weights; retain copyright and license.源码与官方 DINOv2 权重;保留版权和许可。 · GitHubGitHub · License/notice 1许可/声明 1
- torch — BSD License. Runtime dependency; dependencies and model/data assets retain their own licenses.运行依赖;依赖及模型/数据资源保留各自许可。 · Original source原始来源 · License/notice 1许可/声明 1 · License/notice 2许可/声明 2 ↩ ↩
- trimesh — MIT License. Runtime dependency; dependencies and model/data assets retain their own licenses.运行依赖;依赖及模型/数据资源保留各自许可。 · GitHubGitHub · License/notice 1许可/声明 1 ↩ ↩
- tensorrt-cu12 — Other/Proprietary License. Runtime dependency; dependencies and model/data assets retain their own licenses.运行依赖;依赖及模型/数据资源保留各自许可。 · GitHubGitHub · License/notice 1许可/声明 1 ↩ ↩
- LM-O data — CC-BY-SA-4.0. Data and model excerpt, not the loader code.数据与模型摘录,不是读取代码。 · Original source原始来源 · GitHub: BOP toolkitGitHub:BOP 工具集
- MUSE
- GeM pooling
Model weights and dataset/task assets may have separate terms. WAPR's source license does not replace them.模型权重、数据集与任务资源可能有独立条款;WAPR 源码许可不替代这些许可。