{"schemaVersion":"jobsearcher.job.v1","id":"fe35b6d06407cb412f802ee4","url":"https://jobsearcher.com/jobs/fe35b6d06407cb412f802ee4","canonicalUrl":"https://jobsearcher.com/jobs/fe35b6d06407cb412f802ee4","title":"Site Reliability Engineer","description":"Company Background\nSpecter's mission is to help automate the physical world.\n\nToday, we build video sensors with state‑of‑the‑art AI agents that answer any question, anywhere in their environments. Our systems can automatically detect and reason about any physical activity captured on camera, from security incidents (e.g. perimeter intrusion, theft, LPR), to safety monitoring (e.g. PPE detection, injured people), to operational efficiency (e.g. material tracking, congestion monitoring). We offer both long‑range wireless (1km range) and wired sensor variants to suit any deployment.\n\nOur co‑founders Xerxes and Philip are passionate about empowering our partners in the fast approaching world of physical AI and robotics. We are a small, fast growing team who hail from Anduril, Tesla, Uber, and the U.S. Special Forces.\n\nThe Role\nWe're hiring a Site Reliability Engineer to own the operational health of our connected sensor platform — spanning a live fleet of edge hardware deployed at customer sites and the cloud infrastructure behind it.\n\nThis is a high‑ownership role at the intersection of ops and platform engineering. You'll drive reliability across our sensor fleet — triaging issues in the field, building the systems that prevent them from recurring, and owning the observability that keeps us ahead of problems as we scale.\n\nYou set your own priorities across all three:\n\nResponsibilities\nReactive — Triage & Recovery\n\nDebug and triage issues across a live fleet of diverse Linux‑based sensor nodes and edge appliances deployed at customer sites.\n\nSSH into field hardware to diagnose, patch, and recover systems — often with limited remote access and incomplete information.\n\nOwn site bring‑ups end to end; be the person who gets things back online.\n\nSystems Builder — Close the Loop\n\nBuild and maintain fleet management systems: OTA update pipelines, device health tracking, remote diagnostics, and lifecycle tooling.\n\nIdentify repeat fires and eliminate them — build tooling, pre‑deployment checks, and root cause processes that prevent recurrence.\n\nAutomate toil relentlessly: if you're doing something twice, you should be scripting it.\n\nCollaborate with embedded systems, and platform teams to define reliability and deployment requirements.\n\nObservability Owner — Fleet Visibility\n\nDesign and implement observability (logging, metrics, alerting) across edge devices and cloud infrastructure (AWS).\n\nSurface and close telemetry gaps; build fleet‑wide visibility that enables data‑driven reliability decisions.\n\nDevelop runbooks, incident response procedures, and participate in on‑call rotations.\n\nQualifications\n\nStrong Linux systems administration — comfortable working over SSH in production, not just dev environments.\n\nExperience with edge or on‑prem hardware alongside cloud infrastructure.\n\nSolid networking fundamentals: DNS, firewalls, VPNs, subnets, secure remote access.\n\nScripting or programming in Python, Go, or Bash for operational tooling.\n\nFamiliarity with containerization (Docker, Kubernetes a plus).\n\nEmbedded systems experience — reading firmware logs, understanding hardware‑software boundaries, and reasoning about what's happening below the OS is a meaningful edge in this role.\n\nDeeper cloud experience (AWS infrastructure, IAM, networking, observability tooling) is a strong plus for owning the cloud side of the fleet.\n\nRust or C experience — we have firmware in both; being able to read and reason about low‑level code accelerates triage significantly.\n\n#J-18808-Ljbffr","company":"Specter Services","rawCompany":"specter services","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-08T03:24:38.006Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541519","title":"Other Computer Related Services","slug":"other-computer-related-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Site Reliability Engineer","description":"Company Background\nSpecter's mission is to help automate the physical world.\n\nToday, we build video sensors with state‑of‑the‑art AI agents that answer any question, anywhere in their environments. Our systems can automatically detect and reason about any physical activity captured on camera, from security incidents (e.g. perimeter intrusion, theft, LPR), to safety monitoring (e.g. PPE detection, injured people), to operational efficiency (e.g. material tracking, congestion monitoring). We offer both long‑range wireless (1km range) and wired sensor variants to suit any deployment.\n\nOur co‑founders Xerxes and Philip are passionate about empowering our partners in the fast approaching world of physical AI and robotics. We are a small, fast growing team who hail from Anduril, Tesla, Uber, and the U.S. Special Forces.\n\nThe Role\nWe're hiring a Site Reliability Engineer to own the operational health of our connected sensor platform — spanning a live fleet of edge hardware deployed at customer sites and the cloud infrastructure behind it.\n\nThis is a high‑ownership role at the intersection of ops and platform engineering. You'll drive reliability across our sensor fleet — triaging issues in the field, building the systems that prevent them from recurring, and owning the observability that keeps us ahead of problems as we scale.\n\nYou set your own priorities across all three:\n\nResponsibilities\nReactive — Triage & Recovery\n\nDebug and triage issues across a live fleet of diverse Linux‑based sensor nodes and edge appliances deployed at customer sites.\n\nSSH into field hardware to diagnose, patch, and recover systems — often with limited remote access and incomplete information.\n\nOwn site bring‑ups end to end; be the person who gets things back online.\n\nSystems Builder — Close the Loop\n\nBuild and maintain fleet management systems: OTA update pipelines, device health tracking, remote diagnostics, and lifecycle tooling.\n\nIdentify repeat fires and eliminate them — build tooling, pre‑deployment checks, and root cause processes that prevent recurrence.\n\nAutomate toil relentlessly: if you're doing something twice, you should be scripting it.\n\nCollaborate with embedded systems, and platform teams to define reliability and deployment requirements.\n\nObservability Owner — Fleet Visibility\n\nDesign and implement observability (logging, metrics, alerting) across edge devices and cloud infrastructure (AWS).\n\nSurface and close telemetry gaps; build fleet‑wide visibility that enables data‑driven reliability decisions.\n\nDevelop runbooks, incident response procedures, and participate in on‑call rotations.\n\nQualifications\n\nStrong Linux systems administration — comfortable working over SSH in production, not just dev environments.\n\nExperience with edge or on‑prem hardware alongside cloud infrastructure.\n\nSolid networking fundamentals: DNS, firewalls, VPNs, subnets, secure remote access.\n\nScripting or programming in Python, Go, or Bash for operational tooling.\n\nFamiliarity with containerization (Docker, Kubernetes a plus).\n\nEmbedded systems experience — reading firmware logs, understanding hardware‑software boundaries, and reasoning about what's happening below the OS is a meaningful edge in this role.\n\nDeeper cloud experience (AWS infrastructure, IAM, networking, observability tooling) is a strong plus for owning the cloud side of the fleet.\n\nRust or C experience — we have firmware in both; being able to read and reason about low‑level code accelerates triage significantly.\n\n#J-18808-Ljbffr","datePosted":"2026-07-08T03:24:38.006Z","dateModified":"2026-07-08T03:24:38.006Z","hiringOrganization":{"@type":"Organization","name":"Specter Services","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"fe35b6d06407cb412f802ee4"},"url":"https://jobsearcher.com/jobs/fe35b6d06407cb412f802ee4"}}