{"schemaVersion":"jobsearcher.job.v1","id":"d8ee884cd64a6a8036bf76f8","url":"https://jobsearcher.com/jobs/d8ee884cd64a6a8036bf76f8","canonicalUrl":"https://jobsearcher.com/jobs/d8ee884cd64a6a8036bf76f8","title":"Senior Forward Deployed Engineer (DevOps/SRE)","description":"Experience: Senior Level\nSalary: $300,000 - $350,000 per year\n\nJob Details\n-\n\nResponsibilities:\nImplement and optimize an AI-powered Site Reliability Engineering (SRE) platform to meet customer needs across production and pre-production environments.\nProactively monitor customer deployments to ensure customers maximize value from the platform.\nIdentify latent reliability issues such as misconfigurations, deployment regressions, and scaling challenges within customer environments.\nRecommend best practices for implementing AI-powered SRE solutions.\nPlan, design, build, and maintain highly scalable, reliable, and efficient cloud infrastructure.\nServe as the customer's technical advocate with internal engineering and product teams.\nConduct post-incident reviews to identify root causes and implement preventative measures.\nEnsure security best practices are integrated into customer deployments.\nTrain customer SRE, Operations, and Platform Engineering teams on platform usage and best practices.\nLead enterprise migrations from legacy alerting, AIOps, and incident management platforms, including correlation rule migration, phased cutovers, and production go-live execution.\nDesign, build, and optimize alert normalization and correlation policies using conditions, regular expressions, field extraction, and customized workflows.\nIntegrate the platform with customer operational systems, including ITSM, collaboration, observability, source control, and documentation platforms.\nValidate and continuously improve AI investigation quality by tuning enrichment, root cause analysis accuracy, and investigation workflows.\nBuild proactive monitoring for customer deployments to identify issues before they impact customers.\nOwn customer-facing project communications, including executive status updates, SLA documentation, escalation management, and implementation tracking.\nDevelop long-term technical relationships with senior engineering leadership.\nOwn customer implementations from technical discovery through solution design, implementation, user acceptance testing, production go-live, stabilization, and ongoing optimization.\nTranslate ambiguous customer requirements into clear technical designs, milestones, acceptance criteria, and execution plans.\nDesign and implement AI-powered investigation and automation workflows with appropriate guardrails, governance, deterministic fallbacks, and human oversight.\nDevelop reusable deployment modules, reference architectures, implementation guides, and operational runbooks to accelerate future deployments.\nDefine customer success metrics, establish baselines, measure operational improvements, and demonstrate business value through KPIs such as MTTR reduction and operational efficiency.\nCapture customer feedback and recurring implementation learnings to influence future product development.\nFoster a culture of continuous improvement and technical excellence.\n\nQualifications:\nCustomer-focused with deep empathy for SRE, DevOps, Platform Engineering, and IT Operations teams.\nBachelor's degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience).\n6+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or similar infrastructure-focused roles, including technical leadership or end-to-end customer delivery.\nExperience in Forward Deployed Engineering, Solutions Engineering, Technical Customer Success, or Professional Services is highly preferred.\nStrong programming experience in at least one language such as Python, Go, or Java.\nHands-on experience with public cloud platforms (AWS, Azure, or Google Cloud Platform).\nStrong knowledge of Kubernetes, Infrastructure as Code (Terraform, CloudFormation, Ansible), and CI/CD pipelines.\nPractical experience using Generative AI and machine learning technologies to improve engineering productivity.\nExperience with observability platforms, ITSM systems, and incident management tools, including systems integration and data mapping.\nStrong troubleshooting, analytical, and debugging skills, including alert correlation, normalization, and regular expression development.\nExcellent written and verbal communication skills.\nDemonstrated ownership of enterprise software implementations from discovery through production deployment.\nStrong integration experience with APIs, webhooks, event-driven architectures, authentication (SSO/SAML), data transformations, synchronization, and enterprise application integrations.\nExperience designing and deploying production-grade AI or automation workflows with governance and evaluation frameworks.\nUnderstanding of enterprise security concepts including RBAC, encryption, identity management, auditing, and secure networking.\nAbility to operate effectively in ambiguous, fast-paced customer environments while balancing architecture with execution.\nSelf-motivated, adaptable, and capable of managing shifting priorities while driving successful customer outcomes.\n\nPreferred Qualifications:\nExperience supporting customers operating AI infrastructure or AI-enabled platforms.\nExperience migrating customers from legacy alerting, AIOps, or incident management platforms.\nExperience building internal automation and tooling using Python, Node.js, Bash, or similar scripting languages.\n\nA bit about us:\n-\n\nBacked by over $21M in capital from leading investors, they are building a next-generation AI product designed to transform how reliability engineering is done.\n\nThe founding team includes senior leaders and technical pioneers from industry giants like AWS, Cisco, VMware, and Gigamon - holding dozens of patents and having built critical systems at some of the most respected tech companies in the industry. This is a rare opportunity to join an early-stage team that’s solving tough technical problems in distributed systems, observability, and automation — all while shaping a product from the ground up.\n\nWhy join us?\n-\n\nBenefits: Comprehensive medical, vision, and dental benefits. 401 (k) plans and commuter benefits. Free lunches, snacks, and top-of-the-line espressos!\n\nEquity that could change your life.\n\nHigh-impact role with plenty of mentorship opportunities from founders and other coworkers\n\nCollaborative coworkers with high IQ and high EQ. No politics. No bureaucracy. Open door policy.\n\n#techservices #python #aws #ai #java #azure #splunk #datadog #prometheus #devops #gcp #sre #servicenow #dynatrace #oberservability #claude-code #bigpanda #moogsoft #tier3","company":"Leoforce","rawCompany":"leoforce","city":"Pleasanton","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-06T15:07:41.095Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior Forward Deployed Engineer (DevOps/SRE)","description":"Experience: Senior Level\nSalary: $300,000 - $350,000 per year\n\nJob Details\n-\n\nResponsibilities:\nImplement and optimize an AI-powered Site Reliability Engineering (SRE) platform to meet customer needs across production and pre-production environments.\nProactively monitor customer deployments to ensure customers maximize value from the platform.\nIdentify latent reliability issues such as misconfigurations, deployment regressions, and scaling challenges within customer environments.\nRecommend best practices for implementing AI-powered SRE solutions.\nPlan, design, build, and maintain highly scalable, reliable, and efficient cloud infrastructure.\nServe as the customer's technical advocate with internal engineering and product teams.\nConduct post-incident reviews to identify root causes and implement preventative measures.\nEnsure security best practices are integrated into customer deployments.\nTrain customer SRE, Operations, and Platform Engineering teams on platform usage and best practices.\nLead enterprise migrations from legacy alerting, AIOps, and incident management platforms, including correlation rule migration, phased cutovers, and production go-live execution.\nDesign, build, and optimize alert normalization and correlation policies using conditions, regular expressions, field extraction, and customized workflows.\nIntegrate the platform with customer operational systems, including ITSM, collaboration, observability, source control, and documentation platforms.\nValidate and continuously improve AI investigation quality by tuning enrichment, root cause analysis accuracy, and investigation workflows.\nBuild proactive monitoring for customer deployments to identify issues before they impact customers.\nOwn customer-facing project communications, including executive status updates, SLA documentation, escalation management, and implementation tracking.\nDevelop long-term technical relationships with senior engineering leadership.\nOwn customer implementations from technical discovery through solution design, implementation, user acceptance testing, production go-live, stabilization, and ongoing optimization.\nTranslate ambiguous customer requirements into clear technical designs, milestones, acceptance criteria, and execution plans.\nDesign and implement AI-powered investigation and automation workflows with appropriate guardrails, governance, deterministic fallbacks, and human oversight.\nDevelop reusable deployment modules, reference architectures, implementation guides, and operational runbooks to accelerate future deployments.\nDefine customer success metrics, establish baselines, measure operational improvements, and demonstrate business value through KPIs such as MTTR reduction and operational efficiency.\nCapture customer feedback and recurring implementation learnings to influence future product development.\nFoster a culture of continuous improvement and technical excellence.\n\nQualifications:\nCustomer-focused with deep empathy for SRE, DevOps, Platform Engineering, and IT Operations teams.\nBachelor's degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience).\n6+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or similar infrastructure-focused roles, including technical leadership or end-to-end customer delivery.\nExperience in Forward Deployed Engineering, Solutions Engineering, Technical Customer Success, or Professional Services is highly preferred.\nStrong programming experience in at least one language such as Python, Go, or Java.\nHands-on experience with public cloud platforms (AWS, Azure, or Google Cloud Platform).\nStrong knowledge of Kubernetes, Infrastructure as Code (Terraform, CloudFormation, Ansible), and CI/CD pipelines.\nPractical experience using Generative AI and machine learning technologies to improve engineering productivity.\nExperience with observability platforms, ITSM systems, and incident management tools, including systems integration and data mapping.\nStrong troubleshooting, analytical, and debugging skills, including alert correlation, normalization, and regular expression development.\nExcellent written and verbal communication skills.\nDemonstrated ownership of enterprise software implementations from discovery through production deployment.\nStrong integration experience with APIs, webhooks, event-driven architectures, authentication (SSO/SAML), data transformations, synchronization, and enterprise application integrations.\nExperience designing and deploying production-grade AI or automation workflows with governance and evaluation frameworks.\nUnderstanding of enterprise security concepts including RBAC, encryption, identity management, auditing, and secure networking.\nAbility to operate effectively in ambiguous, fast-paced customer environments while balancing architecture with execution.\nSelf-motivated, adaptable, and capable of managing shifting priorities while driving successful customer outcomes.\n\nPreferred Qualifications:\nExperience supporting customers operating AI infrastructure or AI-enabled platforms.\nExperience migrating customers from legacy alerting, AIOps, or incident management platforms.\nExperience building internal automation and tooling using Python, Node.js, Bash, or similar scripting languages.\n\nA bit about us:\n-\n\nBacked by over $21M in capital from leading investors, they are building a next-generation AI product designed to transform how reliability engineering is done.\n\nThe founding team includes senior leaders and technical pioneers from industry giants like AWS, Cisco, VMware, and Gigamon - holding dozens of patents and having built critical systems at some of the most respected tech companies in the industry. This is a rare opportunity to join an early-stage team that’s solving tough technical problems in distributed systems, observability, and automation — all while shaping a product from the ground up.\n\nWhy join us?\n-\n\nBenefits: Comprehensive medical, vision, and dental benefits. 401 (k) plans and commuter benefits. Free lunches, snacks, and top-of-the-line espressos!\n\nEquity that could change your life.\n\nHigh-impact role with plenty of mentorship opportunities from founders and other coworkers\n\nCollaborative coworkers with high IQ and high EQ. No politics. No bureaucracy. Open door policy.\n\n#techservices #python #aws #ai #java #azure #splunk #datadog #prometheus #devops #gcp #sre #servicenow #dynatrace #oberservability #claude-code #bigpanda #moogsoft #tier3","datePosted":"2026-08-06T15:07:41.095Z","dateModified":"2026-08-06T15:07:41.095Z","hiringOrganization":{"@type":"Organization","name":"Leoforce","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Pleasanton","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"d8ee884cd64a6a8036bf76f8"},"url":"https://jobsearcher.com/jobs/d8ee884cd64a6a8036bf76f8"}}