{"schemaVersion":"jobsearcher.job.v1","id":"27ab229dc1a1585ab099ab55","url":"https://jobsearcher.com/jobs/27ab229dc1a1585ab099ab55","canonicalUrl":"https://jobsearcher.com/jobs/27ab229dc1a1585ab099ab55","title":"DevOps / Site Reliability Engineer","description":"Description\nWe're looking for a DevOps / SRE engineer to own the reliability, delivery, and observability of our AI platform. You'll be the person who ensures models get from a developer's branch to production without anyone losing sleep — and when something does go wrong at 2am, you'll be the one who knows where to look.\n\nDepartment: Engineering\n\nLocation: San Francisco\n\nWe run production across multiple Kubernetes clusters, cloud providers, and regions. Our deployment pipeline is fully automated through CI/CD and GitOps, our infrastructure is managed as code, and our observability stack gives us full visibility across every service and GPU workload. This role is about making all of that faster, more reliable, and easier to operate as we scale.\n\nWhat You'll Do\n\nOwn and evolve our CI/CD pipelines: dynamic pipeline generation across a monorepo of Go services, Python model containers, and Helm charts\n\nOperate and improve our GitOps deployment lifecycle: Helm releases, Kustomizations, and image automation across multiple clusters\n\nBuild and maintain our observability stack: distributed tracing, metrics, dashboards, and alerting across all services and GPU workloads\n\nDefine and track SLOs for core platform services, including session latency, model cold start time, and streaming reliability\n\nRun incident response: triage production issues, write postmortems, build runbooks, and drive reliability improvements\n\nManage infrastructure-as-code across multiple cloud providers and regions: plan/apply workflows, state management, drift detection\n\nOperate secret management: encrypted secrets, external secret syncing, certificate automation\n\nImprove deployment safety: canary rollouts, health checks, startup probes, rollback automation\n\nManage authentication infrastructure: OIDC federation for CI, workload identity for cloud services, cross-cloud credential management\n\nParticipate in on-call rotation and build the tooling that makes on-call less painful\n\nWhat We're Looking For\n\nYou've run production Kubernetes clusters and been on-call for them. You've debugged node scheduling failures, OOM kills, and mysterious pod evictions at 3am\n\nStrong CI/CD experience: you've built and maintained pipelines for monorepos, not just single-service repos\n\nGitOps experience: you understand reconciliation loops, drift detection, and why image automation matters\n\nInfrastructure-as-code fluency with Terraform or similar across multiple environments and cloud accounts\n\nYou know observability beyond just \"set up dashboards\". You've defined SLOs, built alerting that doesn't page on noise, and used traces to debug cross-service latency issues.\n\nComfortable with secret management patterns (KMS, encrypted configs, external secret operators). You've thought about credential rotation and zero-trust.\n\nIncident response experience: you've triaged production outages, written postmortems that actually led to improvements, and built runbooks that other engineers could follow\n\nYou write code, not just YAML. Proficiency in Go, Python, or Bash for building tooling, automation, and pipeline scripts\n\nNice to Have\n\nExperience with GPU workloads on Kubernetes: device plugins, GPU-aware scheduling, GPU monitoring\n\nMulti-cloud operations beyond a single provider\n\nReal-time or streaming workloads: low-latency systems where p99 matters more than average\n\nExperience with Helm chart authoring and managing complex value layering across environments\n\nFamiliarity with real-time media or relay infrastructure\n\nFinOps experience: GPU cost optimization, spot/preemptible instance management\n\nWhat We're Not Looking For\n\nEngineers who treat infrastructure-as-code as \"click around in the console and import later\"\n\nSREs who've only monitored systems but never built the deployment pipelines that ship to them\n\nCandidates whose CI/CD experience is limited to GitHub Actions for a single-service repo\n\nPeople who write alerts that fire every day and then get ignored\n\nLogistics\nWe are based in-person in San Francisco. We are also hiring for this role in Europe for on-call coverage and timezone distribution.\n\nBenefits\n\nCompetitive salary and meaningful early equity\n\nVisa sponsorship and relocation support\n\nGenerous health, dental, and vision coverage\n\n#J-18808-Ljbffr","company":"Reactoram","rawCompany":"reactoram","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-16T03:59:49.138Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"DevOps / Site Reliability Engineer","description":"Description\nWe're looking for a DevOps / SRE engineer to own the reliability, delivery, and observability of our AI platform. You'll be the person who ensures models get from a developer's branch to production without anyone losing sleep — and when something does go wrong at 2am, you'll be the one who knows where to look.\n\nDepartment: Engineering\n\nLocation: San Francisco\n\nWe run production across multiple Kubernetes clusters, cloud providers, and regions. Our deployment pipeline is fully automated through CI/CD and GitOps, our infrastructure is managed as code, and our observability stack gives us full visibility across every service and GPU workload. This role is about making all of that faster, more reliable, and easier to operate as we scale.\n\nWhat You'll Do\n\nOwn and evolve our CI/CD pipelines: dynamic pipeline generation across a monorepo of Go services, Python model containers, and Helm charts\n\nOperate and improve our GitOps deployment lifecycle: Helm releases, Kustomizations, and image automation across multiple clusters\n\nBuild and maintain our observability stack: distributed tracing, metrics, dashboards, and alerting across all services and GPU workloads\n\nDefine and track SLOs for core platform services, including session latency, model cold start time, and streaming reliability\n\nRun incident response: triage production issues, write postmortems, build runbooks, and drive reliability improvements\n\nManage infrastructure-as-code across multiple cloud providers and regions: plan/apply workflows, state management, drift detection\n\nOperate secret management: encrypted secrets, external secret syncing, certificate automation\n\nImprove deployment safety: canary rollouts, health checks, startup probes, rollback automation\n\nManage authentication infrastructure: OIDC federation for CI, workload identity for cloud services, cross-cloud credential management\n\nParticipate in on-call rotation and build the tooling that makes on-call less painful\n\nWhat We're Looking For\n\nYou've run production Kubernetes clusters and been on-call for them. You've debugged node scheduling failures, OOM kills, and mysterious pod evictions at 3am\n\nStrong CI/CD experience: you've built and maintained pipelines for monorepos, not just single-service repos\n\nGitOps experience: you understand reconciliation loops, drift detection, and why image automation matters\n\nInfrastructure-as-code fluency with Terraform or similar across multiple environments and cloud accounts\n\nYou know observability beyond just \"set up dashboards\". You've defined SLOs, built alerting that doesn't page on noise, and used traces to debug cross-service latency issues.\n\nComfortable with secret management patterns (KMS, encrypted configs, external secret operators). You've thought about credential rotation and zero-trust.\n\nIncident response experience: you've triaged production outages, written postmortems that actually led to improvements, and built runbooks that other engineers could follow\n\nYou write code, not just YAML. Proficiency in Go, Python, or Bash for building tooling, automation, and pipeline scripts\n\nNice to Have\n\nExperience with GPU workloads on Kubernetes: device plugins, GPU-aware scheduling, GPU monitoring\n\nMulti-cloud operations beyond a single provider\n\nReal-time or streaming workloads: low-latency systems where p99 matters more than average\n\nExperience with Helm chart authoring and managing complex value layering across environments\n\nFamiliarity with real-time media or relay infrastructure\n\nFinOps experience: GPU cost optimization, spot/preemptible instance management\n\nWhat We're Not Looking For\n\nEngineers who treat infrastructure-as-code as \"click around in the console and import later\"\n\nSREs who've only monitored systems but never built the deployment pipelines that ship to them\n\nCandidates whose CI/CD experience is limited to GitHub Actions for a single-service repo\n\nPeople who write alerts that fire every day and then get ignored\n\nLogistics\nWe are based in-person in San Francisco. We are also hiring for this role in Europe for on-call coverage and timezone distribution.\n\nBenefits\n\nCompetitive salary and meaningful early equity\n\nVisa sponsorship and relocation support\n\nGenerous health, dental, and vision coverage\n\n#J-18808-Ljbffr","datePosted":"2026-07-16T03:59:49.138Z","dateModified":"2026-07-16T03:59:49.138Z","hiringOrganization":{"@type":"Organization","name":"Reactoram","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"27ab229dc1a1585ab099ab55"},"url":"https://jobsearcher.com/jobs/27ab229dc1a1585ab099ab55"}}