Perception Engineer
Company Description Sauron is a home operating system focused on creating a secure, private sanctuary for households through an autonomous perimeter security platform. Operating discreetly in the background, the system reliably identifies potential threats in varied environmental conditions while instantly recognizing trusted individuals. The company delivers a bespoke white-glove service, installing each client’s security platform with precision and minimal disruption to home comfort and aesthetics. Sauron’s technology is complemented by the IRIS Command Center, staffed 24/7 by highly trained security professionals with diverse protective-service backgrounds. This team collaborates closely with local law enforcement to help ensure rapid response to verified security incidents.Role Description The Perception Engineer is a full-time, on-site role based in San Francisco, CA, focused on advancing Sauron’s sensing and perception capabilities for residential perimeter security. In this role, the engineer designs, trains, and optimizes perception algorithms that fuse data from cameras, sensors, and other inputs to reliably detect, classify, and track people, vehicles, and potential threats in real-world environments. Day-to-day work includes building and maintaining data pipelines, developing computer vision and machine learning models, evaluating system performance across diverse lighting and weather conditions, and iterating based on field feedback. The engineer collaborates with hardware, software, and operations teams to integrate perception modules into production systems, improve real-time processing, and enhance safety, reliability, and user experience. The role also involves documenting methodologies, participating in code reviews, and contributing to long-term technical roadmaps for Sauron’s perception technology.The System You'd Own Perception is the foundation of the Sauron product. Everything the system decides - what to show a homeowner, what to ignore, what to escalate - rests on whether we correctly understood what happened outside their house. We are hiring the person who owns that surface. Our hardware operates around the home and has to complete its mission reliably in every environmental condition, at night, in weather, against occlusion and deliberate evasion. You would set the technical direction for how we get there: what we sense, what we infer, which problems are model problems and which are systems problems, and what "good enough to ship" means in numbers rather than impressions. This is a staff role in the real sense. The scope is a system and a team, not a backlog. You would own architecture and the accuracy, latency, and cost tradeoffs behind it; work closely with the hardware team to set sensing requirements across generations of the product; and be the person cloud, product, and hardware come to when they need to know what perception can and cannot do. You will write a great deal of code, on the hardest parts — but you will be judged on whether the perception surface as a whole gets better, faster, and more trustworthy. You Will Contribute By These are real open items in the codebase today, not hypotheticals.Evolving the pipeline architecture so cameras can come and go, and configuration can change, without disrupting live video or losing object identity. Setting the boundary between systems and models. You decide what belongs in the compiled service, what belongs in the inference graph, and where the performance ceiling really is. Owning perception quality end to end. Defining what good looks like in numbers: evaluation data, a regression harness, and a defensible story on model choice for our hardware. Extracting the maximum value from our sensors. Fusing every observation available while staying robust to occlusion, poor lighting, and deliberate evasion. Taking the service from "runs" to "trustworthy unattended." Failure detection, graceful degradation, and observability good enough that we know why something broke without a site visit. Closing the loop from the field. Using deployment data to find headroom, and building the dataset and evaluation infrastructure that makes that repeatable. Leading the work and the people. Setting direction across perception, reviewing the hard changes, and owning the interfaces perception exposes to the rest of the product. Your Background Includes Around 10+ years building production systems, with several at staff scope: owning a system's architecture, not just its tickets. Significant professional experience with perception or machine learning for hardware products in a safety-critical field — aerospace, robotics, medical devices, autonomous vehicles, or physical security. Deep modern C++. This is a C++20 codebase with Abseil, gRPC, and CMake; you should be comfortable owning lifetime, threading, and shutdown semantics in a long-running daemon. Real GStreamer or media-pipeline experience: pads, probes, caps negotiation, bus messages, and the specific pathology of a pipeline that is alive but not moving. DeepStream or another NVIDIA video stack is a strong plus. Applied computer vision you have shipped — multi-object tracking, re-identification, or multi-camera association — and the judgment to know when the answer is a better model versus better geometry versus better plumbing. A clear grasp of linear algebra, optimization, statistics, and algorithms, and the theory behind the techniques you reach for. Experience across the deep-learning lifecycle: PyTorch or an equivalent framework, custom layers and operations, optimizing networks for inference on edge compute, reproducibility, and honest evaluation. GPU inference in practice: TensorRT engines, batching, fp16, and reasoning about where latency actually goes. Edge instincts. You have debugged something that only fails on the device, after nine hours, on one customer's network, and you treat observability and failure classification as part of the feature. Python fluency. A meaningful share of the load-bearing logic is Python, and you will be the one deciding what it costs us. A generalist mindset — able to dive in wherever the bottleneck is, from cloud training infrastructure down to embedded systems. Excellent written and verbal communication, and the ability to set technical direction and disagree productively with adjacent teams. Nice to HaveNVIDIA Jetson in production — JetPack, L4T, Yocto images, or the joy of cross-building for aarch64. A multi-camera tracking or re-ID system that people depended on, and the ability to talk about where it broke. Video surveillance, VMS, or ONVIF/RTSP integrations, and knowing how cameras actually misbehave. Middleware frameworks such as ROS. GPU architecture and CUDA programming. Familiarity with VLMs and other multi-modal models for semantic scene understanding. Owning model evaluation: datasets, metrics, and the discipline to reject a model that benchmarks better but ships worse.