JOBSEARCHER

Cleared DevOps Site Reliability Engineer (SRE)

Job Summary We are seeking a highly skilled and TS/SCI security-cleared DevOps Site Reliability Engineer (SRE) to join our dynamic team. In this role, you will be responsible for designing, implementing, and maintaining resilient, scalable, and secure cloud infrastructure and IT systems. Your expertise will ensure high availability, optimal performance, and robust disaster recovery solutions across enterprise environments. The ideal candidate will possess a strong background in DevOps practices, cloud computing, system administration, and software troubleshooting within a security-sensitive context. This position offers an exciting opportunity to contribute to mission-critical projects that demand precision, innovation, and operational excellence. Key Responsibilities: Monitor the health and performance of enterprise infrastructure through continuous system monitoring and automated telemetry to support the required 99.9% platform uptime. Participate in an on-call rotation and respond to major incidents or platform outages within one hour of notification, executing rapid troubleshooting and system stabilization activities. Develop, maintain, and enhance automation scripts and internal tools that streamline diagnostics, health checks, and routine operational tasks. Design and maintain dashboards that provide real-time visibility into platform health, including uptime, API performance, incident status, and other key operational metrics. Coordinate directly with Cloud Service Providers (CSPs) during infrastructure outages or service disruptions to expedite issue resolution. Continuously assess system reliability, logging, monitoring, and overall architecture, providing recommendations that improve scalability, resiliency, and operational efficiency. Required Qualifications: TS/SCI security clearance. Strong understanding of Site Reliability Engineering (SRE) principles and best practices. Hands-on experience deploying and managing containerized applications using Kubernetes. Experience administering and troubleshooting multi-cloud environments, including Google Cloud Platform (GCP), Microsoft Azure, and Amazon Web Services (AWS). Experience implementing and maintaining enterprise monitoring, logging, and automated alerting solutions. Proficiency with scripting and automation using languages such as Python, Bash, or similar technologies. Preferred Qualifications: Passion for building and maintaining highly reliable, mission-critical systems with demanding uptime requirements. Ability to remain composed and methodical while responding to high-priority production incidents. Strong troubleshooting, root cause analysis, and diagnostic skills with a focus on rapid issue resolution. A continuous improvement mindset with an emphasis on automation and eliminating repetitive operational tasks. Experience supporting secure, cloud-native environments within government or defense organizations is a plus. Pay: $100,000.00 - $150,000.00 per year Benefits: Health insurance Paid time off Retirement plan Vision insurance Application Question(s): Is your TS/SCI clearance still active and valid? Education: Bachelor's (Required) Work Location: In person