WAPR, wide-angle pose refinement
WAPR

Detection and pose setup检测与位姿环境

For the published wheel, start with the 0.0.5 installation guide. The repository paths and manual source instructions below apply to source checkouts.

已发布 wheel 请先按0.0.5 安装引导操作。下方仓库路径和手工源码步骤适用于源码副本。

Wheel resourceswheel 资源目录

Installed-wheel resources default to the user cache (~/.cache/wapr, or XDG_CACHE_HOME/wapr). Set WAPR_CACHE_DIR to an absolute directory before running to use another location. Required public resources are obtained by the relevant package entry points; gated models require the user’s own access or local files.

安装 wheel 后资源默认存入用户缓存(~/.cache/wapr 或 XDG_CACHE_HOME/wapr)。运行前设置绝对路径 WAPR_CACHE_DIR 可指定其他目录。所需公开资源由对应包入口获取;受控模型需用户自己的访问权限或本地文件。

python -m wapr.bootstrap --feature det2d --check
python -m wapr.bootstrap --feature det2d

2D detection and 6D pose estimation, localization and detection use the same setup. Pose refinement, scoring, GPU rendering and TensorRT[6] deployment are installed together; no separate environment is needed for each task.

2D 检测及 6D 位姿估计、定位、检测使用同一套环境。位姿修正、评分、GPU 渲染和 TensorRT[6] 推理部署一起准备,无需按任务重复安装。

PyTorch, TensorRT and OpenCV are prepared by the user. Ordinary dependencies have no exact version pins. The installer uses requirements.txt for pose/export/build dependencies and adds requirements-detector.txt only when 2D detection is enabled. ONNX is needed to export engines; it is not a basic wheel dependency. TensorRT engines must match the target GPU and TensorRT environment.

PyTorch、TensorRT 和 OpenCV 由用户准备,普通依赖不固定精确版本。安装脚本使用 requirements.txt 补齐位姿、导出与编译依赖,只有启用 2D 检测时才增加 requirements-detector.txt。ONNX 用于导出引擎,不属于基础 wheel 依赖。TensorRT 引擎须与目标 GPU 和 TensorRT 环境匹配。

Pose weights and inference engines位姿权重与推理引擎

The installer downloads the four pose checkpoints and LM-O[8] sample packs and compiles the renderer. python -m wapr.export_engines builds the four standard engines and wbps_batch.engine on the inference GPU. Weights and engines go to assets/weights/; sample data goes to samples/.

安装脚本下载四份位姿权重与 LM-O[8] 样例包,并编译渲染器。python -m wapr.export_engines 在推理 GPU 上构建四份常规引擎及 wbps_batch.engine。权重与引擎保存在 assets/weights/,样例数据保存在 samples/。

For pose-only use, set prepare_2d_detector = False in wapr/tools/install_wapr.py before installation. To update an existing four-engine setup, rebuild the renderer with the installer and run python -m wapr.export_batch_engine to add the grouped scorer.

只使用位姿估计时,在安装前将 wapr/tools/install_wapr.py 中的 prepare_2d_detector = False。更新已有四引擎环境时,用安装脚本重建渲染器,再运行 python -m wapr.export_batch_engine 补充批量计算评分引擎。

python -m wapr.download_assets fetches all supported small sample packs and pose weights. Full evaluation datasets and tracking sequences are separate downloads; see Datasets and model resources.

python -m wapr.download_assets 下载全部支持的小样例包和位姿权重。完整评测数据集及跟踪序列另行下载,见数据集与模型资源。

Fetch获取

The command below checks the four-checkpoint pack. If any checkpoint is missing, it fetches the whole pack again and replaces the four destination files. Files available under the local assets/hf/ directory are copied; the others come from the Hugging Face model repo SEU-WYL/WAPR. The download tries https://huggingface.co first and continues from https://hf-mirror.com if that host does not connect. Set HF_ENDPOINT to try one address first. Destinations are listed in assets/weights/PATHS.txt. examples/02_one_category_one_instance.py performs the same pack check at startup and checks its LM-O sample pack separately.

下面的命令按四份权重组成的数据包检查。只要有一份缺失,就重新获取整个数据包,并替换四份目标文件。本地 assets/hf/ 中已有的文件直接复制,其余文件从 Hugging Face 模型仓库 SEU-WYL/WAPR 获取。下载首先访问 https://huggingface.co;连接失败时,同一次运行接着尝试 https://hf-mirror.com。设置 HF_ENDPOINT 可以指定优先尝试的地址。目标路径列在 assets/weights/PATHS.txt 中。examples/02_one_category_one_instance.py 启动时也检查这份数据包,并单独检查 LM-O 样例数据包。

These paths apply to source installation. Run the installer and python -m wapr.export_engines before creating WAPREstimator; source use requires prepared weights and engines. A ready resource pack causes no Hub request. Public downloads do not send a locally stored Hugging Face token to either endpoint.

上述路径适用于源码安装。构造 WAPREstimator 前,先运行安装脚本和 python -m wapr.export_engines,准备权重与引擎。资源包齐全时不会请求 Hub。公开资源下载不会向任一地址发送本机保存的 Hugging Face 令牌。

only = ("wapr_sapr_wbps",)
python -m wapr.download_assets

only is the tuple in wapr/download_assets.py. only = ("wapr_sapr_wbps",) downloads these four checkpoints. only = () also downloads the one-frame packs named on the BOP Challenge page.

only 是 wapr/download_assets.py 里的元组。only = ("wapr_sapr_wbps",) 下载这四份权重。only = () 还会下载BOP 挑战页列出的单帧示例。

2D detector sources2D 检测源码

The default installer needs compatible GroundingDINO[1], Ultralytics[2] v8.3.70 and DINOv2[3] source trees under third_party/. It checks these trees, downloads detector weights and builds the DINOv2 engine; it does not clone or adapt upstream source code. Use the supplied compatible trees when available. For a fresh or offline setup, use the source instructions and offline file list below.

默认安装需要 third_party/ 中的兼容 GroundingDINO[1]、Ultralytics[2] v8.3.70 和 DINOv2[3] 源码。脚本检查这些目录,下载检测权重并构建 DINOv2 引擎,但不会克隆或适配上游源码。已有兼容源码时直接使用;从头准备或离线安装时,直接按下方源码配置及离线文件清单准备。

Prepare GroundingDINO and DINOv2 source trees, then copy the supplied attention adaptation and JSON configurations before running the installer.

Source preparation and offline resources源码准备与离线资源

GroundingDINO

GroundingDINO uses the supplied PyTorch attention adaptation. Copy the two JSON configurations and attention module as shown below. Do not use pip install -e for GroundingDINO: WAPR imports this source tree directly.

GroundingDINO 使用随附的 PyTorch 注意力适配。按下方命令复制两份 JSON 配置及注意力模块;无需对 GroundingDINO 执行 pip install -e,WAPR 直接导入该源码目录。

git clone https://github.com/IDEA-Research/GroundingDINO.git third_party/GroundingDINO
cp wapr/tools/compat/groundingdino_swin*.json third_party/GroundingDINO/groundingdino/config/
cp wapr/tools/compat/ms_deform_attn.py third_party/GroundingDINO/groundingdino/models/GroundingDINO/ms_deform_attn.py

The supplied attention module routes both CPU and CUDA execution through multi_scale_deformable_attn_pytorch, avoiding GroundingDINO’s separately compiled _C extension. Copying it with the command above applies this adaptation; no further source edits are needed.

随附的注意力模块在 CPU 和 CUDA 上均调用 multi_scale_deformable_attn_pytorch,无需另行编译 GroundingDINO 的 _C 扩展。执行上方复制命令即可完成适配,不需要再次修改源码。

output = multi_scale_deformable_attn_pytorch(
    value, spatial_shapes, sampling_locations, attention_weights
)

SAM

SAM 2.1[4]-L is loaded by Ultralytics v8.3.70. Use the local tree or clone that tag into third_party/ultralytics if absent. Do not install Ultralytics with pip or clone facebookresearch/sam2 for this detector. The installer fetches sam2.1_l.pt; ultralytics-thop in requirements.txt is a separate helper package.

SAM 2.1[4]-L 由 Ultralytics v8.3.70 加载。使用本地源码;若缺失,再将该 tag 克隆到 third_party/ultralytics。此检测路径不用 pip 安装 Ultralytics,也不用克隆 facebookresearch/sam2。安装脚本会获取 sam2.1_l.pt;requirements.txt 中的 ultralytics-thop 是单独的辅助包。

git clone --branch v8.3.70 --depth 1 https://github.com/ultralytics/ultralytics.git third_party/ultralytics

DINOv2

Use the local DINOv2 tree or clone facebookresearch/dinov2 into third_party/dinov2 if absent. It must contain hubconf.py. xFormers is not needed for this inference path; its absence may produce a startup warning. The installer fetches dinov2_vitl14_pretrain.pth and builds the TensorRT engine by default. Set build_dino_engine_now = False in the installer to defer that build until the first detector call.

使用本地 DINOv2 源码;若缺失,可将 facebookresearch/dinov2 克隆到 third_party/dinov2,其中必须有 hubconf.py。这条推理路径不需要 xFormers;缺失时可能出现启动警告。安装脚本会获取 dinov2_vitl14_pretrain.pth,默认预构建 TensorRT 引擎;在脚本中将 build_dino_engine_now 设为 False,即可推迟到首次检测调用时构建。

git clone https://github.com/facebookresearch/dinov2.git third_party/dinov2

BERT

GroundingDINO’s text encoder is google-bert/bert-base-uncased, a model rather than a git checkout. The installer places its config, model, and tokenizer files in assets/weights/det2d/bert-base-uncased/. wapr/det2d.py points text_encoder_type to that local directory. For manual or offline setup, the five file links are in the file list below.

GroundingDINO 的文本编码器是 google-bert/bert-base-uncased 模型,不是 git 源码。安装脚本将配置、模型和 tokenizer 文件放在 assets/weights/det2d/bert-base-uncased/;wapr/det2d.py 将 text_encoder_type 指向该目录。手动或离线安装所需的五份文件链接见下方文件清单。

2D detector assets2D 检测器资源

GroundingDINO, SAM, DINOv2 and BERT[7] provide the detector components. After preparing the source trees above, use the following file locations, direct download links and checksums for an offline installation.

检测器使用 GroundingDINO、SAM、DINOv2 与 BERT[7]。完成上方源码配置后,离线安装可直接按以下文件位置、下载链接和校验值准备。

Weight files权重文件

Create assets/weights/det2d/. These weights are not in SEU-WYL/WAPR. You can place the published files yourself, or let WAPRDet2D download a missing one. Pass dino and grounding on that constructor. A missing file is written into assets/weights/det2d/, and backend="trt" builds the DINOv2 engine for that weight when none matches. The required weight files and destinations are listed below. The lists below are the published vitl14 and swinb pair.

先建 assets/weights/det2d/。这些权重不在 SEU-WYL/WAPR。可以自己放上已发表的文件,也可以让 WAPRDet2D 去下载缺的那一份。在构造时传入 dino 和 grounding。缺的文件写进 assets/weights/det2d/,backend="trt" 在没有匹配引擎时为这份 DINOv2 权重构建引擎。所需权重文件及目标位置列在下方。下面列出的是已发表的 vitl14 和 swinb。

GroundingDINO

The Swin-B config is copied automatically from third_party/GroundingDINO/groundingdino/config/groundingdino_swinb.json to assets/weights/det2d/ when the detector weights are prepared. A fresh upstream clone still calls a CUDA extension on GPU; the attention adaptation selects multi_scale_deformable_attn_pytorch instead.

准备检测器权重时,Swin-B 配置会自动从 third_party/GroundingDINO/groundingdino/config/groundingdino_swinb.json 复制到 assets/weights/det2d/。新克隆的上游源码在 GPU 上仍调用 CUDA 扩展;注意力适配使其改用 multi_scale_deformable_attn_pytorch。

SAM

SAM 2.1-L, loaded by Ultralytics v8.3.70. The checkout is the tag above. The weight is the Ultralytics asset, not a clone of facebookresearch/sam2.

SAM 2.1-L,由 Ultralytics v8.3.70 加载。检出是上面那个 tag。权重是 Ultralytics 的资源文件,不是 facebookresearch/sam2 的克隆。

DINOv2

ViT-L/14. The checkout needs hubconf.py at third_party/dinov2/hubconf.py. The xFormers warning at startup can be ignored. Do not install xFormers for this package.

ViT-L/14。检出后要有 third_party/dinov2/hubconf.py。启动时的 xFormers 警告可以忽略。不必为本包安装 xFormers。

BERT

GroundingDINO uses Google's bert-base-uncased text encoder. Its upstream name can trigger a Hub request during model construction. This release points text_encoder_type to a local directory instead. The installer fetches the five published files below; interrupted downloads and incomplete local directories are common causes of a first-run failure.

GroundingDINO 使用 Google 发布的 bert-base-uncased 文本编码器。上游名称可能在构建模型时触发 Hub 请求;本项目将 text_encoder_type 指向本地目录。安装脚本获取下面五份公开文件;下载中断或本地目录不完整,常会使首次运行失败。

The installer places these five files in assets/weights/det2d/bert-base-uncased/; for an offline installation, copy them there before starting the detector. wapr/det2d.py sets text_encoder_type to that directory. If an otherwise complete offline setup still attempts a Hub request, set HF_HUB_OFFLINE=1 and TRANSFORMERS_OFFLINE=1. The loader permits the known missing bert.embeddings.position_ids key.

安装脚本会将这五份文件放在 assets/weights/det2d/bert-base-uncased/;离线安装时应在启动检测器前手动放齐。wapr/det2d.py 将 text_encoder_type 设为该目录。文件齐全但仍尝试访问 Hub 时,可设置 HF_HUB_OFFLINE=1 和 TRANSFORMERS_OFFLINE=1。加载器允许已知缺失的 bert.embeddings.position_ids 键。

Hashes哈希

Startup compares the weight files with wapr/det2d_assets.json. These are the measured SHA256 values. *.engine is skipped. The DINO engine is built on the Install page and belongs to that GPU.

启动时用 wapr/det2d_assets.json 比对权重。下面是实测 SHA256。*.engine 不比对。DINO 引擎在安装页编译,属于那块 GPU。

File文件SHA256
groundingdino_swinb_cogcoor.pth46270f7a822e6906b655b729c90613e48929d0f2bb8b9b76fd10a856f3ac6ab7
sam2.1_l.ptab7e1ac9cb9f6eb3bcf197ece044f06a707ec49129361a2b47e93e1db6989efd
dinov2_vitl14_pretrain.pthd5383ea8f4877b2472eb973e0fd72d557c7da5d3611bd527ceeb1d7162cbf428
bert-base-uncased/config.json7160e1553ad2ca51d8c1cb066be533db31826e12d173824c1bb0cb1a4f187d20
bert-base-uncased/pytorch_model.bin097417381d6c7230bd9e3557456d726de6e83245ec8b24f529f60198a67b203a
bert-base-uncased/tokenizer.jsonce64fce797c24f68df90b40a3f74f579b336a493db14bd583fd520ea0d8c9a98
bert-base-uncased/tokenizer_config.jsona025160ef0431f1a392f6f050c1310f4c5d9fb6f275932dbccba73c4d214bf10
bert-base-uncased/vocab.txt07eced375cec144d27c900241f3e339478dec958f92fddbc551f295c992038a3
groundingdino_swinb.json65c95dabc0e6b02ad0edd506d21e2fbf535e422a56a870813e9b1806f4537d71

The saved BOP2D accuracy results used MartinSmeyer/cocoapi v1.0. Its evaluation files are not required for detection or pose inference.

BOP2D 精度统计采用 MartinSmeyer/cocoapi v1.0 评测;检测与位姿推理无需安装其评测文件。

Troubleshooting故障排查

Common installation problems常见安装问题
Symptom现象Check and action检查与处理
nvcc, EGL, or OpenGL missingThe PyTorch[5] wheel does not supply the compiler or GL/EGL development files. For source compilation, install a CUDA toolkit supporting your GPU and your distribution's OpenGL/EGL development packages; set CUDA_HOME if the toolkit is outside /usr/local/cuda.PyTorch[5] wheel 不提供编译器或 GL/EGL 开发文件。源码编译时安装支持当前 GPU 的 CUDA toolkit 和发行版对应的 OpenGL/EGL 开发包;若 toolkit 不在 /usr/local/cuda,设置 CUDA_HOME。
torch.cuda.is_available() == FalseCheck the NVIDIA driver and GPU access inside the container. A CUDA-enabled wheel alone cannot make a GPU visible.检查 NVIDIA 驱动和容器内 GPU 映射;仅安装带 CUDA 的 wheel 无法让容器获得 GPU。
2D source check fails2D 源码检查失败Use the compatible third_party/ trees. Fresh GroundingDINO[1] clones need the PyTorch attention adaptation. For pose-only examples 02 and 03, set prepare_2d_detector = False.使用兼容的 third_party/ 源码;新克隆的 GroundingDINO[1] 需要纯 PyTorch 注意力适配。仅运行位姿示例 02、03 时可将 prepare_2d_detector 设为 False。
Asset download or hash fails资源下载或哈希失败Check disk space and the URLs in the file lists above. Pose packs come from SEU-WYL/WAPR; detector files are listed above. Fetch incomplete files again and retain SHA-256 verification. Run python -m wapr.download_assets to retry the configured sample/pose packs.检查磁盘空间及上方文件清单中的下载地址。位姿资源来自 SEU-WYL/WAPR,检测文件见上方清单。文件不完整时重新获取,保留 SHA-256 校验;可运行 python -m wapr.download_assets 重试当前配置的小样例与位姿数据包。
Engine fails after moving machines换机器后引擎无法加载Run python -m wapr.export_engines on the inference GPU to rebuild the four standard pose engines and the grouped scoring engine. Rerun python wapr/tools/install_wapr.py to prepare the 2D engine through the default installation. Build with the Python, TensorRT[6] version, and GPU used for inference.在推理显卡上运行 python -m wapr.export_engines,重建四份常规位姿引擎及批量计算评分引擎。2D 引擎通过运行默认安装 python wapr/tools/install_wapr.py 准备。编译环境应与推理所用的 Python、TensorRT[6] 版本和 GPU 对应。

References and licenses参考文献与许可

  1. GroundingDINO — Apache-2.0. Vendored detector source; original notices retained.随包检测源码;保留原始声明。 · GitHubGitHub · License/notice 1许可/声明 1 · License/notice 2许可/声明 2
    Liu et al. Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection. ECCV 2024. · Paper论文 ↩ ↩ ↩ ↩
  2. Ultralytics — AGPL-3.0. Vendored 8.3.70 loader; review AGPL obligations or obtain an upstream commercial license. Native Meta SAM 2 has separate terms.随包 8.3.70 加载器;须遵守 AGPL 条款或取得上游商业授权。Meta 原生 SAM 2 另有许可。 · GitHubGitHub · License/notice 1许可/声明 1 · License/notice 2许可/声明 2 ↩ ↩
  3. DINOv2 — Apache-2.0. Code and official DINOv2 weights; retain copyright and license.源码与官方 DINOv2 权重;保留版权和许可。 · GitHubGitHub · License/notice 1许可/声明 1
    Oquab et al. DINOv2: Learning Robust Visual Features without Supervision. arXiv:2304.07193, 2023. · Paper论文 ↩ ↩
  4. SAM 2 — Apache-2.0; cctorch: BSD-3-Clause. Native source and official checkpoints; cctorch carries an additional BSD notice. Ultralytics-converted files require checking their distributor terms.原生源码与官方权重;cctorch 另附 BSD 声明。Ultralytics 转换文件还需核对分发方条款。 · GitHubGitHub · License/notice 1许可/声明 1 · License/notice 2许可/声明 2
    Ravi et al. SAM 2: Segment Anything in Images and Videos. arXiv:2408.00714, 2024. · Paper论文 ↩ ↩
  5. torch — BSD License. Runtime dependency; dependencies and model/data assets retain their own licenses.运行依赖;依赖及模型/数据资源保留各自许可。 · Original source原始来源 · License/notice 1许可/声明 1 · License/notice 2许可/声明 2 ↩ ↩ ↩ ↩
  6. tensorrt-cu12 — Other/Proprietary License. Runtime dependency; dependencies and model/data assets retain their own licenses.运行依赖;依赖及模型/数据资源保留各自许可。 · GitHubGitHub · License/notice 1许可/声明 1 ↩ ↩ ↩ ↩
  7. BERT base uncased — Apache-2.0. Official selected text model; preserve its model-repository terms.所选官方文本模型;保留其模型仓库条款。 · Original source原始来源 · License/notice 1许可/声明 1
    Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL-HLT 2019. · Paper论文 ↩ ↩
  8. LM-O data — CC-BY-SA-4.0. Data and model excerpt, not the loader code.数据与模型摘录,不是读取代码。 · Original source原始来源 · GitHub: BOP toolkitGitHub:BOP 工具集
    Brachmann et al. Learning 6D Object Pose Estimation Using 3D Object Coordinates. ECCV 2014. · Paper论文 ↩ ↩

Model weights and dataset/task assets may have separate terms. WAPR's source license does not replace them.模型权重、数据集与任务资源可能有独立条款;WAPR 源码许可不替代这些许可。