{"schemaVersion":"jobsearcher.job.v1","id":"8b2ee54b670f4a507e5416aa","url":"https://jobsearcher.com/jobs/8b2ee54b670f4a507e5416aa","canonicalUrl":"https://jobsearcher.com/jobs/8b2ee54b670f4a507e5416aa","title":"Senior/Staff Software Engineer - Machine Learning & System Optimization","description":"Overview\nIn this role, you will optimize and orchestrate on-vehicle AI workloads, focusing on efficient deployment of large models on edge hardware. You will drive cross-cutting initiatives to share and fuse models, improve compute utilization, and ensure real-time, deterministic inference for autonomous perception. You’ll build production-grade C++/CUDA components and edge deployment pipelines that push multi-modal models into production on SoCs. This is a high-impact, fast-moving opportunity to advance autonomous system intelligence at scale.\n\nCompensation / Benefitsbase salaryAmazon RSUsZoox Stock Appreciation Rightssign-on bonus may be offeredhealth insurancepaid time off\nResponsibilitiesAllocate and distribute CPU/GPU/interconnect resources across models and inference engines on the robotLead initiatives to share/fuse models and improve scheduling for better compute utilizationOptimize large-scale models (Multi-Modal Sensor Fusion, LLMs, VLMs) using PTQ/QAT, pruning, mixed-precision, and LoRA-based fine-tuningDesign and implement model conversion and compilation pipelines with TensorRT for edge deploymentDevelop production-level, low-latency, memory-safe C++ and CUDA code for real-time edge inference on vehicle systems\nKey requirementsDeep experience in system and performance optimization for CPU/GPU with low latency or high throughputStrong expertise in real-time systems and constraints such as latency, memory usage, and memory bandwidthProficiency in model quantization (PTQ, QAT) and mixed-precision inference (INT8, FP8, FP4, BF16/FP16)Experience with developing and optimizing custom ML operators and TensorRT plugins with efficient CUDA kernelsProduction-level C++ (14/17/20) and Python, with ability to write concurrent, memory-safe real-time edge codeSystem and performance optimization on CPU/GPU for low latencyReal-time systems design and constraints managementModel quantization (PTQ, QAT) and mixed-precision inference (INT8, FP8, FP16/BF16)CUDA kernel development and low-level AI accelerator programmingTensorRT and TensorRT Plugins developmentC++ (14/17/20) and Python for concurrent, real-time edge coding","company":"Zoox","rawCompany":"zoox","city":"Somerville","state":"MA","isRemote":false,"isActive":false,"createdAt":"2026-09-15T03:56:13.779Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior/Staff Software Engineer - Machine Learning & System Optimization","description":"Overview\nIn this role, you will optimize and orchestrate on-vehicle AI workloads, focusing on efficient deployment of large models on edge hardware. You will drive cross-cutting initiatives to share and fuse models, improve compute utilization, and ensure real-time, deterministic inference for autonomous perception. You’ll build production-grade C++/CUDA components and edge deployment pipelines that push multi-modal models into production on SoCs. This is a high-impact, fast-moving opportunity to advance autonomous system intelligence at scale.\n\nCompensation / Benefitsbase salaryAmazon RSUsZoox Stock Appreciation Rightssign-on bonus may be offeredhealth insurancepaid time off\nResponsibilitiesAllocate and distribute CPU/GPU/interconnect resources across models and inference engines on the robotLead initiatives to share/fuse models and improve scheduling for better compute utilizationOptimize large-scale models (Multi-Modal Sensor Fusion, LLMs, VLMs) using PTQ/QAT, pruning, mixed-precision, and LoRA-based fine-tuningDesign and implement model conversion and compilation pipelines with TensorRT for edge deploymentDevelop production-level, low-latency, memory-safe C++ and CUDA code for real-time edge inference on vehicle systems\nKey requirementsDeep experience in system and performance optimization for CPU/GPU with low latency or high throughputStrong expertise in real-time systems and constraints such as latency, memory usage, and memory bandwidthProficiency in model quantization (PTQ, QAT) and mixed-precision inference (INT8, FP8, FP4, BF16/FP16)Experience with developing and optimizing custom ML operators and TensorRT plugins with efficient CUDA kernelsProduction-level C++ (14/17/20) and Python, with ability to write concurrent, memory-safe real-time edge codeSystem and performance optimization on CPU/GPU for low latencyReal-time systems design and constraints managementModel quantization (PTQ, QAT) and mixed-precision inference (INT8, FP8, FP16/BF16)CUDA kernel development and low-level AI accelerator programmingTensorRT and TensorRT Plugins developmentC++ (14/17/20) and Python for concurrent, real-time edge coding","datePosted":"2026-09-15T03:56:13.779Z","dateModified":"2026-09-15T03:56:13.779Z","hiringOrganization":{"@type":"Organization","name":"Zoox","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Somerville","addressRegion":"MA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"8b2ee54b670f4a507e5416aa"},"url":"https://jobsearcher.com/jobs/8b2ee54b670f4a507e5416aa"}}