Offloading ZED Depth Computation to Jetson DLA / PVA on AGX Orin

Hi everyone,

I am currently developing a Vision-Language-Action (VLA) pipeline deployed on an NVIDIA Jetson AGX Orin.

Because the downstream policy and vision models demand substantial GPU compute and VRAM, I am looking to maximize the AGX Orin’s dedicated hardware accelerators—specifically the DLA (Deep Learning Accelerator) and PVA (Programmable Vision Accelerator)—to relieve the GPU load.

I would like to ask:

  1. Does the ZED SDK currently support offloading depth estimation (e.g., NEURAL or ULTRA mode) to the DLA or PVA engines?

  2. If not natively exposed via InitParameters, is there any planned roadmap, experimental flag, or workaround to achieve this?

  3. Has anyone in the community successfully routed ZED stereo matching or depth processing to DLA/PVA?

Any insights, benchmarks, or recommended architectural practices would be greatly appreciated.

Environment:

  • Platform: Jetson AGX Orin (64GB)

  • Camera: ZED X

  • Capture card: ZED Link (Duo)

  • JetPack: 7.2

  • ZED SDK: ZED SDK for JetPack 7.2 (L4T 39.2) 5.5

Thanks in advance!

Hi @Chaejun-Lim,
Welcome to the StereoLabs community.

No, the ZED SDK depth processing runs on the GPU only; DLA and PVA are not supported.
Please note that ULTRA is an obsolete mode based on the old stereo matching pipeline, kept only for backward compatibility with ZED SDK v4.x applications; for new applications, use one of the AI depth modes: NEURAL_LIGHT, NEURAL, or NEURAL_PLUS.

There are no experimental flags or workarounds to route the ZED SDK depth processing to DLA or PVA, and I cannot share roadmap details on this topic.

This is not possible, because the depth engine is internal to the ZED SDK and is not exposed for external deployment.

To reduce the GPU load of the depth processing, you can:

  • set InitParameters::depth_precision = DEPTH_PRECISION::INT8 with the NEURAL depth mode; the neural network inference runs at lower precision, reducing GPU memory usage and increasing speed with a small accuracy cost, typically 1-2% (Depth Precision documentation)
  • use DEPTH_MODE::NEURAL_LIGHT, the fastest AI depth mode, if its ideal range (0.3 to 5 m) fits your application; you can compare the computational performance of each mode on embedded devices in the Depth Modes documentation
  • set InitParameters::depth_stabilization = 0 if you do not need temporal filtering, so the positional tracking module is not activated in the background (Depth Stabilization documentation)
  • lower the camera resolution and the grab frame rate to the values strictly required by your policy
  • set RuntimeParameters::enable_depth = false for the frames where depth is not needed

On the other side, you can deploy your own downstream vision models on the AGX Orin DLA engines with TensorRT, in INT8 or FP16 precision, leaving more GPU resources available to the ZED SDK and to the models that cannot run on DLA.