{"schemaVersion":"jobsearcher.job.v1","id":"b41a56edf734effd2e9e83dd","url":"https://jobsearcher.com/jobs/b41a56edf734effd2e9e83dd","canonicalUrl":"https://jobsearcher.com/jobs/b41a56edf734effd2e9e83dd","title":"Staff Site Reliability Engineer","description":"This is not a ticket-taking SRE role.\nYou will define how mission-critical machine learning and real-time analytics systems operate in production — influencing reliability strategy, deployment standards, and infrastructure architecture across engineering.\nThis team operates in a highly collaborative, in-person engineering environment in SOMA. Infrastructure, ML, and engineering leaders work side by side to design, build, and operate complex systems in real time. The pace is fast, the feedback loops are tight, and decisions happen quickly.\nIf you’ve grown from Linux systems DevOps Staff-level SRE, and you now think in terms of systemic risk, scalability, and long-term reliability strategy — this role gives you direct influence and visibility.\nThis role is intentionally in-person because:\nReliability decisions happen at architectural depth — not over Slack threads\n\nML, data, and infrastructure teams collaborate continuously in real time\n\nPost-incident reviews, system design debates, and performance tuning sessions are hands-on and high impact\n\nYou will have direct access to engineering leadership and decision-makers\n\nThe infrastructure you’re operating is mission-critical and evolving quickly\n\nIf you value deep technical collaboration, tight feedback loops, and being at the center of high-scale ML systems — this environment is built for that.\n\nWhat You’ll Own\nProduction reliability for ML and real-time analytics workloads\n\nCI/CD strategy, deployment automation, and rollback design\n\nObservability frameworks (SLOs, alerting, monitoring, incident response)\n\nInfrastructure-as-Code and Kubernetes environments\n\nCapacity planning and performance optimization\n\nPost-incident reviews that drive measurable, long-term reliability improvements\n\nReliability standards across teams — not just within a single service\n\nYou’ll partner directly with engineering and data science teams to ensure ML workloads are production-ready and reliable by design.\n\nWhat We’re Looking For\nDeep experience operating Linux infrastructure and networking in production environments\n\nProven impact as a Staff SRE, Senior SRE, or senior-level DevOps/Platform Engineer supporting distributed systems\n\nExperience supporting complex, data-intensive or ML-driven systems in production\n\nStrong hands-on experience with Docker and Kubernetes\n\nInfrastructure-as-Code expertise\n\nStrong scripting ability (Bash and/or Python)\n\nCI/CD ownership experience (GitHub Actions, ArgoCD, or similar)\n\nExperience with modern observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry)\n\nAbility to debug systemic failures across infrastructure, deployments, and workloads\n\nClear communicator who works effectively across engineering and data teams\n\nEngineers who have evolved from infrastructure foundations into strategic reliability leaders will thrive here.\n\nThese Skills Are a Plus\nExperience operating ML platforms at scale (training + inference)\n\nAWS or cloud-managed services experience\n\nExposure to data platforms such as Spark, Airflow, or Kafka\n\nExperience in SOC 2 or regulated environments\n\nWhy This Opportunity\nStaff-level ownership of mission-critical ML infrastructure\n\nDirect influence over reliability standards across engineering\n\nHigh-visibility role with architectural impact\n\nCollaborative engineering culture designed for speed and depth\n\nCompetitive base compensation ($210K–$250K)\n\nIf you're a Staff-level reliability engineer who wants real ownership and architectural influence — let’s start the conversation.\nStratITech is partnering with our San Francisco client to build the next generation of high-scale ML infrastructure.","company":"Stratitech Services","rawCompany":"stratitech services","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-05T13:01:30.707Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Staff Site Reliability Engineer","description":"This is not a ticket-taking SRE role.\nYou will define how mission-critical machine learning and real-time analytics systems operate in production — influencing reliability strategy, deployment standards, and infrastructure architecture across engineering.\nThis team operates in a highly collaborative, in-person engineering environment in SOMA. Infrastructure, ML, and engineering leaders work side by side to design, build, and operate complex systems in real time. The pace is fast, the feedback loops are tight, and decisions happen quickly.\nIf you’ve grown from Linux systems DevOps Staff-level SRE, and you now think in terms of systemic risk, scalability, and long-term reliability strategy — this role gives you direct influence and visibility.\nThis role is intentionally in-person because:\nReliability decisions happen at architectural depth — not over Slack threads\n\nML, data, and infrastructure teams collaborate continuously in real time\n\nPost-incident reviews, system design debates, and performance tuning sessions are hands-on and high impact\n\nYou will have direct access to engineering leadership and decision-makers\n\nThe infrastructure you’re operating is mission-critical and evolving quickly\n\nIf you value deep technical collaboration, tight feedback loops, and being at the center of high-scale ML systems — this environment is built for that.\n\nWhat You’ll Own\nProduction reliability for ML and real-time analytics workloads\n\nCI/CD strategy, deployment automation, and rollback design\n\nObservability frameworks (SLOs, alerting, monitoring, incident response)\n\nInfrastructure-as-Code and Kubernetes environments\n\nCapacity planning and performance optimization\n\nPost-incident reviews that drive measurable, long-term reliability improvements\n\nReliability standards across teams — not just within a single service\n\nYou’ll partner directly with engineering and data science teams to ensure ML workloads are production-ready and reliable by design.\n\nWhat We’re Looking For\nDeep experience operating Linux infrastructure and networking in production environments\n\nProven impact as a Staff SRE, Senior SRE, or senior-level DevOps/Platform Engineer supporting distributed systems\n\nExperience supporting complex, data-intensive or ML-driven systems in production\n\nStrong hands-on experience with Docker and Kubernetes\n\nInfrastructure-as-Code expertise\n\nStrong scripting ability (Bash and/or Python)\n\nCI/CD ownership experience (GitHub Actions, ArgoCD, or similar)\n\nExperience with modern observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry)\n\nAbility to debug systemic failures across infrastructure, deployments, and workloads\n\nClear communicator who works effectively across engineering and data teams\n\nEngineers who have evolved from infrastructure foundations into strategic reliability leaders will thrive here.\n\nThese Skills Are a Plus\nExperience operating ML platforms at scale (training + inference)\n\nAWS or cloud-managed services experience\n\nExposure to data platforms such as Spark, Airflow, or Kafka\n\nExperience in SOC 2 or regulated environments\n\nWhy This Opportunity\nStaff-level ownership of mission-critical ML infrastructure\n\nDirect influence over reliability standards across engineering\n\nHigh-visibility role with architectural impact\n\nCollaborative engineering culture designed for speed and depth\n\nCompetitive base compensation ($210K–$250K)\n\nIf you're a Staff-level reliability engineer who wants real ownership and architectural influence — let’s start the conversation.\nStratITech is partnering with our San Francisco client to build the next generation of high-scale ML infrastructure.","datePosted":"2026-08-05T13:01:30.707Z","dateModified":"2026-08-05T13:01:30.707Z","hiringOrganization":{"@type":"Organization","name":"Stratitech Services","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"b41a56edf734effd2e9e83dd"},"url":"https://jobsearcher.com/jobs/b41a56edf734effd2e9e83dd"}}