I am currently developing a Vision-Language-Action (VLA) pipeline deployed on an NVIDIA Jetson AGX Orin.
Because the downstream policy and vision models demand substantial GPU compute and VRAM, I am looking to maximize the AGX Orin’s dedicated hardware accelerators—specifically the DLA (Deep Learning Accelerator) and PVA (Programmable Vision Accelerator)—to relieve the GPU load.
I would like to ask:
Does the ZED SDK currently support offloading depth estimation (e.g., NEURAL or ULTRA mode) to the DLA or PVA engines?
If not natively exposed via InitParameters, is there any planned roadmap, experimental flag, or workaround to achieve this?
Has anyone in the community successfully routed ZED stereo matching or depth processing to DLA/PVA?
Any insights, benchmarks, or recommended architectural practices would be greatly appreciated.
No, the ZED SDK depth processing runs on the GPU only; DLA and PVA are not supported.
Please note that ULTRA is an obsolete mode based on the old stereo matching pipeline, kept only for backward compatibility with ZED SDK v4.x applications; for new applications, use one of the AI depth modes: NEURAL_LIGHT, NEURAL, or NEURAL_PLUS.
There are no experimental flags or workarounds to route the ZED SDK depth processing to DLA or PVA, and I cannot share roadmap details on this topic.
This is not possible, because the depth engine is internal to the ZED SDK and is not exposed for external deployment.
To reduce the GPU load of the depth processing, you can:
set InitParameters::depth_precision = DEPTH_PRECISION::INT8 with the NEURAL depth mode; the neural network inference runs at lower precision, reducing GPU memory usage and increasing speed with a small accuracy cost, typically 1-2% (Depth Precision documentation)
use DEPTH_MODE::NEURAL_LIGHT, the fastest AI depth mode, if its ideal range (0.3 to 5 m) fits your application; you can compare the computational performance of each mode on embedded devices in the Depth Modes documentation
set InitParameters::depth_stabilization = 0 if you do not need temporal filtering, so the positional tracking module is not activated in the background (Depth Stabilization documentation)
lower the camera resolution and the grab frame rate to the values strictly required by your policy
set RuntimeParameters::enable_depth = false for the frames where depth is not needed
On the other side, you can deploy your own downstream vision models on the AGX Orin DLA engines with TensorRT, in INT8 or FP16 precision, leaving more GPU resources available to the ZED SDK and to the models that cannot run on DLA.