JOBSEARCHER

Senior/Staff Software Engineer - Machine Learning & System Optimization

ZooxQuincy, MAL6 LeadSeptember 15th, 2026
Overview In this role, you will optimize and orchestrate on-vehicle AI workloads, focusing on efficient deployment of large models on edge hardware. You will drive cross-cutting initiatives to share and fuse models, improve compute utilization, and ensure real-time, deterministic inference for autonomous perception. You’ll build production-grade C++/CUDA components and edge deployment pipelines that push multi-modal models into production on SoCs. This is a high-impact, fast-moving opportunity to advance autonomous system intelligence at scale. Compensation / Benefitsbase salaryAmazon RSUsZoox Stock Appreciation Rightssign-on bonus may be offeredhealth insurancepaid time off ResponsibilitiesAllocate and distribute CPU/GPU/interconnect resources across models and inference engines on the robotLead initiatives to share/fuse models and improve scheduling for better compute utilizationOptimize large-scale models (Multi-Modal Sensor Fusion, LLMs, VLMs) using PTQ/QAT, pruning, mixed-precision, and LoRA-based fine-tuningDesign and implement model conversion and compilation pipelines with TensorRT for edge deploymentDevelop production-level, low-latency, memory-safe C++ and CUDA code for real-time edge inference on vehicle systems Key requirementsDeep experience in system and performance optimization for CPU/GPU with low latency or high throughputStrong expertise in real-time systems and constraints such as latency, memory usage, and memory bandwidthProficiency in model quantization (PTQ, QAT) and mixed-precision inference (INT8, FP8, FP4, BF16/FP16)Experience with developing and optimizing custom ML operators and TensorRT plugins with efficient CUDA kernelsProduction-level C++ (14/17/20) and Python, with ability to write concurrent, memory-safe real-time edge codeSystem and performance optimization on CPU/GPU for low latencyReal-time systems design and constraints managementModel quantization (PTQ, QAT) and mixed-precision inference (INT8, FP8, FP16/BF16)CUDA kernel development and low-level AI accelerator programmingTensorRT and TensorRT Plugins developmentC++ (14/17/20) and Python for concurrent, real-time edge coding