{"schemaVersion":"jobsearcher.job.v1","id":"b1547b024f20d6438826c055","url":"https://jobsearcher.com/jobs/b1547b024f20d6438826c055","canonicalUrl":"https://jobsearcher.com/jobs/b1547b024f20d6438826c055","title":"DevOps Site Reliability Engineer (SRE)","description":"DevOps Site Reliability Engineer (SRE)\nLocation: Washington, D.C.\nClearance Required: TS/SCI\n\nPosition Overview\nIT Veterans is seeking a DevOps Site Reliability Engineer (SRE) to support the reliability, performance, and operational stability of a mission-critical enterprise platform. This position plays a vital role in ensuring continuous availability across multi-cloud environments while supporting software deployments, infrastructure monitoring, incident response, and system automation.\nThe ideal candidate will help bridge development and operations by implementing reliable deployment practices, building robust monitoring capabilities, and rapidly responding to production issues to maintain the required 99.9% platform availability for critical Department of Defense (DoD) systems.\n\nKey Responsibilities:\nMonitor the health and performance of enterprise infrastructure through continuous system monitoring and automated telemetry to support the required 99.9% platform uptime.\nParticipate in an on-call rotation and respond to major incidents or platform outages within one hour of notification, executing rapid troubleshooting and system stabilization activities.\nDevelop, maintain, and enhance automation scripts and internal tools that streamline diagnostics, health checks, and routine operational tasks.\nDesign and maintain dashboards that provide real-time visibility into platform health, including uptime, API performance, incident status, and other key operational metrics.\nCoordinate directly with Cloud Service Providers (CSPs) during infrastructure outages or service disruptions to expedite issue resolution.\nContinuously assess system reliability, logging, monitoring, and overall architecture, providing recommendations that improve scalability, resiliency, and operational efficiency.\nRequired Qualifications:\nTS/SCI security clearance.\nStrong understanding of Site Reliability Engineering (SRE) principles and best practices.\nHands-on experience deploying and managing containerized applications using Kubernetes.\nExperience administering and troubleshooting multi-cloud environments, including Google Cloud Platform (GCP), Microsoft Azure, and Amazon Web Services (AWS).\nExperience implementing and maintaining enterprise monitoring, logging, and automated alerting solutions.\nProficiency with scripting and automation using languages such as Python, Bash, or similar technologies.\nPreferred Qualifications:\nPassion for building and maintaining highly reliable, mission-critical systems with demanding uptime requirements.\nAbility to remain composed and methodical while responding to high-priority production incidents.\nStrong troubleshooting, root cause analysis, and diagnostic skills with a focus on rapid issue resolution.\nA continuous improvement mindset with an emphasis on automation and eliminating repetitive operational tasks.\nExperience supporting secure, cloud-native environments within government or defense organizations is a plus.\nAt IT Veterans LLC, we are committed to providing an environment of mutual respect where equal employment opportunities are available to all applicants and teammates without regard to race, color, religion, sex, pregnancy, national origin, age, physical and mental disability, marital status, sexual orientation, gender identity, gender expression, genetic information, military and veteran status, and any other characteristic protected by applicable law. We believe that diversity and inclusion among our teammates is critical to our success.","company":"Itveterans","rawCompany":"itveterans","city":"Washington","state":"DC","isRemote":false,"isActive":false,"createdAt":"2026-08-04T22:07:56.944Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541519","title":"Other Computer Related Services","slug":"other-computer-related-services"},{"code":"541513","title":"Computer Facilities Management Services","slug":"computer-facilities-management-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"DevOps Site Reliability Engineer (SRE)","description":"DevOps Site Reliability Engineer (SRE)\nLocation: Washington, D.C.\nClearance Required: TS/SCI\n\nPosition Overview\nIT Veterans is seeking a DevOps Site Reliability Engineer (SRE) to support the reliability, performance, and operational stability of a mission-critical enterprise platform. This position plays a vital role in ensuring continuous availability across multi-cloud environments while supporting software deployments, infrastructure monitoring, incident response, and system automation.\nThe ideal candidate will help bridge development and operations by implementing reliable deployment practices, building robust monitoring capabilities, and rapidly responding to production issues to maintain the required 99.9% platform availability for critical Department of Defense (DoD) systems.\n\nKey Responsibilities:\nMonitor the health and performance of enterprise infrastructure through continuous system monitoring and automated telemetry to support the required 99.9% platform uptime.\nParticipate in an on-call rotation and respond to major incidents or platform outages within one hour of notification, executing rapid troubleshooting and system stabilization activities.\nDevelop, maintain, and enhance automation scripts and internal tools that streamline diagnostics, health checks, and routine operational tasks.\nDesign and maintain dashboards that provide real-time visibility into platform health, including uptime, API performance, incident status, and other key operational metrics.\nCoordinate directly with Cloud Service Providers (CSPs) during infrastructure outages or service disruptions to expedite issue resolution.\nContinuously assess system reliability, logging, monitoring, and overall architecture, providing recommendations that improve scalability, resiliency, and operational efficiency.\nRequired Qualifications:\nTS/SCI security clearance.\nStrong understanding of Site Reliability Engineering (SRE) principles and best practices.\nHands-on experience deploying and managing containerized applications using Kubernetes.\nExperience administering and troubleshooting multi-cloud environments, including Google Cloud Platform (GCP), Microsoft Azure, and Amazon Web Services (AWS).\nExperience implementing and maintaining enterprise monitoring, logging, and automated alerting solutions.\nProficiency with scripting and automation using languages such as Python, Bash, or similar technologies.\nPreferred Qualifications:\nPassion for building and maintaining highly reliable, mission-critical systems with demanding uptime requirements.\nAbility to remain composed and methodical while responding to high-priority production incidents.\nStrong troubleshooting, root cause analysis, and diagnostic skills with a focus on rapid issue resolution.\nA continuous improvement mindset with an emphasis on automation and eliminating repetitive operational tasks.\nExperience supporting secure, cloud-native environments within government or defense organizations is a plus.\nAt IT Veterans LLC, we are committed to providing an environment of mutual respect where equal employment opportunities are available to all applicants and teammates without regard to race, color, religion, sex, pregnancy, national origin, age, physical and mental disability, marital status, sexual orientation, gender identity, gender expression, genetic information, military and veteran status, and any other characteristic protected by applicable law. We believe that diversity and inclusion among our teammates is critical to our success.","datePosted":"2026-08-04T22:07:56.944Z","dateModified":"2026-08-04T22:07:56.944Z","hiringOrganization":{"@type":"Organization","name":"Itveterans","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Washington","addressRegion":"DC","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"b1547b024f20d6438826c055"},"url":"https://jobsearcher.com/jobs/b1547b024f20d6438826c055"}}