{"schemaVersion":"jobsearcher.job.v1","id":"ec0df3a074847eeef1763173","url":"https://jobsearcher.com/jobs/ec0df3a074847eeef1763173","canonicalUrl":"https://jobsearcher.com/jobs/ec0df3a074847eeef1763173","title":"Site Reliability Engineer","description":"Basic Qualifications :\nBachelor's degree in Software Engineering, or related Science, Technology, Engineering or Mathematics field, plus a minimum of 8 years of relevant experience; or Master's degree, plus 6 years relevant experience.\n\nCLEARANCE REQUIREMENTS:: Department of Defense Secret security clearance is required at time of hire. Applicants selected will be subject to a U.S. Government security investigation and must meet eligibility requirements for access to classified information. Due to the nature of work performed within our facilities, U.S. citizenship is required.\nResponsibilities for this Position:\nWhat You'll Own\nSLOs and reliability metrics. Define service level objectives for every AI service that goes to production. Establish error budgets and use them to drive engineering decisions — not just measure uptime.\nMonitoring and observability. Build and maintain monitoring, logging, and alerting infrastructure for AI services. You will know when something is degrading before users do.\nIncident response. Establish incident management procedures, lead post-incident reviews, and drive corrective actions. When something breaks, you coordinate the response and ensure it doesn't break the same way again.\nOperational readiness reviews. Before any AI service goes live, you validate that it meets reliability, security, and operational standards. You are the gate between \"it works in dev\" and \"it's ready for production.\"\nCapacity planning and cost monitoring. Track resource consumption, forecast capacity needs, and monitor costs — tokens, compute, storage. You ensure the platform scales without surprises.\nToil elimination. Identify and automate repetitive operational tasks. If a human is doing something a script could do, you fix that.\nWhat You Won't Own\nApplication development or AI model building — you ensure what they build is operable, you don't build it\nInfrastructure provisioning — IT provides the infrastructure; you define what's needed and validate it works\nBusiness process decisions or backlog prioritization\nWhat Makes This Role Different\nAI services have failure modes that traditional applications don't — model drift, token budget exhaustion, prompt injection, upstream data quality degradation. You will build monitoring for problems that most SRE teams have never encountered.\nYou are applying SRE principles from scratch. There is no existing SRE practice to inherit — you will define it for the platform.\nYour operational readiness reviews directly determine whether AI services go live. You have real authority to say \"not ready.\"\nRequired Qualifications\nBachelor’s degree in Computer Science, Software Engineering, or a related field, plus 8 years of experience; or Master’s degree plus 6 years of experience\nProduction SRE or DevOps experience — you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines\nHands-on experience with monitoring and observability tools — Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that caught real problems.\nStrong scripting and automation skills — Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar)\nExperience with containerized environments — Docker, Kubernetes, container orchestration at scale\nExperience defining and managing SLOs, error budgets, and incident response procedures in production\nU.S. citizenship required. Department of Defense Secret security clearance is required at time of hire.\nPreferred Qualifications\nExperience with AI/ML production systems — model serving, inference monitoring, token cost tracking, or similar\nMulti-cloud experience (AWS, Azure, GCP) including cloud-native monitoring and logging services\nExperience building operational readiness review processes or production launch checklists\nFamiliarity with Google SRE principles — you have read the book and applied the concepts, not just referenced them in interviews\nExperience in environments where reliability has compliance or safety implications — defense, healthcare, finance, or critical infrastructure\nWhat Sets You Apart\nYou think about failure before you think about features. Your first question about any new system is \"how does this break?\"\nYou automate yourself out of toil. If you're doing the same thing twice, you write a script.\nYou have said \"not ready\" to a team that wanted to ship, and you were right.\nYou build monitoring that tells you what's wrong, not just that something is wrong.\nYou write post-incident reviews that actually change how systems are built, not just how incidents are documented.\nDetails\nRemote — 100% telework\n9/80 schedule\nDefense industry experience is not required\nSalary Note: This estimate represents the typical salary range for this position based on experience and other factors (geographic location, etc.). Actual pay may vary. This job posting will remain open until the position is filled. Combined Salary Range: USD $142,696.00 - USD $158,303.00 /Yr. Company Overview:\nGeneral Dynamics Mission Systems (GDMS) engineers a diverse portfolio of high technology solutions, products and services that enable customers to successfully execute missions across all domains of operation. With a global team of 12,000+ top professionals, we partner with the best in industry to expand the bounds of innovation in the defense and scientific arenas. Given the nature of our work and who we are, we value trust, honesty, alignment and transparency. We offer highly competitive benefits and pride ourselves in being a great place to work with a shared sense of purpose. You will also enjoy a flexible work environment where contributions are recognized and rewarded. If who we are and what we do resonates with you, we invite you to join our high-performance team!\n\nEqual Opportunity Employer / Individuals with Disabilities / Protected Veterans","company":"General Dynamics Mission Systems","rawCompany":"general dynamics mission systems","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-08-06T13:41:17.872Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541519","title":"Other Computer Related Services","slug":"other-computer-related-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Site Reliability Engineer","description":"Basic Qualifications :\nBachelor's degree in Software Engineering, or related Science, Technology, Engineering or Mathematics field, plus a minimum of 8 years of relevant experience; or Master's degree, plus 6 years relevant experience.\n\nCLEARANCE REQUIREMENTS:: Department of Defense Secret security clearance is required at time of hire. Applicants selected will be subject to a U.S. Government security investigation and must meet eligibility requirements for access to classified information. Due to the nature of work performed within our facilities, U.S. citizenship is required.\nResponsibilities for this Position:\nWhat You'll Own\nSLOs and reliability metrics. Define service level objectives for every AI service that goes to production. Establish error budgets and use them to drive engineering decisions — not just measure uptime.\nMonitoring and observability. Build and maintain monitoring, logging, and alerting infrastructure for AI services. You will know when something is degrading before users do.\nIncident response. Establish incident management procedures, lead post-incident reviews, and drive corrective actions. When something breaks, you coordinate the response and ensure it doesn't break the same way again.\nOperational readiness reviews. Before any AI service goes live, you validate that it meets reliability, security, and operational standards. You are the gate between \"it works in dev\" and \"it's ready for production.\"\nCapacity planning and cost monitoring. Track resource consumption, forecast capacity needs, and monitor costs — tokens, compute, storage. You ensure the platform scales without surprises.\nToil elimination. Identify and automate repetitive operational tasks. If a human is doing something a script could do, you fix that.\nWhat You Won't Own\nApplication development or AI model building — you ensure what they build is operable, you don't build it\nInfrastructure provisioning — IT provides the infrastructure; you define what's needed and validate it works\nBusiness process decisions or backlog prioritization\nWhat Makes This Role Different\nAI services have failure modes that traditional applications don't — model drift, token budget exhaustion, prompt injection, upstream data quality degradation. You will build monitoring for problems that most SRE teams have never encountered.\nYou are applying SRE principles from scratch. There is no existing SRE practice to inherit — you will define it for the platform.\nYour operational readiness reviews directly determine whether AI services go live. You have real authority to say \"not ready.\"\nRequired Qualifications\nBachelor’s degree in Computer Science, Software Engineering, or a related field, plus 8 years of experience; or Master’s degree plus 6 years of experience\nProduction SRE or DevOps experience — you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines\nHands-on experience with monitoring and observability tools — Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that caught real problems.\nStrong scripting and automation skills — Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar)\nExperience with containerized environments — Docker, Kubernetes, container orchestration at scale\nExperience defining and managing SLOs, error budgets, and incident response procedures in production\nU.S. citizenship required. Department of Defense Secret security clearance is required at time of hire.\nPreferred Qualifications\nExperience with AI/ML production systems — model serving, inference monitoring, token cost tracking, or similar\nMulti-cloud experience (AWS, Azure, GCP) including cloud-native monitoring and logging services\nExperience building operational readiness review processes or production launch checklists\nFamiliarity with Google SRE principles — you have read the book and applied the concepts, not just referenced them in interviews\nExperience in environments where reliability has compliance or safety implications — defense, healthcare, finance, or critical infrastructure\nWhat Sets You Apart\nYou think about failure before you think about features. Your first question about any new system is \"how does this break?\"\nYou automate yourself out of toil. If you're doing the same thing twice, you write a script.\nYou have said \"not ready\" to a team that wanted to ship, and you were right.\nYou build monitoring that tells you what's wrong, not just that something is wrong.\nYou write post-incident reviews that actually change how systems are built, not just how incidents are documented.\nDetails\nRemote — 100% telework\n9/80 schedule\nDefense industry experience is not required\nSalary Note: This estimate represents the typical salary range for this position based on experience and other factors (geographic location, etc.). Actual pay may vary. This job posting will remain open until the position is filled. Combined Salary Range: USD $142,696.00 - USD $158,303.00 /Yr. Company Overview:\nGeneral Dynamics Mission Systems (GDMS) engineers a diverse portfolio of high technology solutions, products and services that enable customers to successfully execute missions across all domains of operation. With a global team of 12,000+ top professionals, we partner with the best in industry to expand the bounds of innovation in the defense and scientific arenas. Given the nature of our work and who we are, we value trust, honesty, alignment and transparency. We offer highly competitive benefits and pride ourselves in being a great place to work with a shared sense of purpose. You will also enjoy a flexible work environment where contributions are recognized and rewarded. If who we are and what we do resonates with you, we invite you to join our high-performance team!\n\nEqual Opportunity Employer / Individuals with Disabilities / Protected Veterans","datePosted":"2026-08-06T13:41:17.872Z","dateModified":"2026-08-06T13:41:17.872Z","hiringOrganization":{"@type":"Organization","name":"General Dynamics Mission Systems","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"ec0df3a074847eeef1763173"},"url":"https://jobsearcher.com/jobs/ec0df3a074847eeef1763173"}}