{"schemaVersion":"jobsearcher.job.v1","id":"747516075018c87f3de91cae","url":"https://jobsearcher.com/jobs/747516075018c87f3de91cae","canonicalUrl":"https://jobsearcher.com/jobs/747516075018c87f3de91cae","title":"Senior Site Reliability Engineer","description":"At Technatomy, we deliver innovative solutions through the efforts of our diverse and talented people who are dedicated to our customer’s success. We provide solutions to agencies and entities including the Department of Veterans Affairs, Department of Defense, Defense Logistics Agency, National Institute of Health, and more. Everything we do is built on a commitment to do the right thing for our customers, our people, and our community. Our Mission, Vision, and Values guide the way we do business.\nIf this sounds like an environment where you can thrive, keep reading!\nWe are seeking an experienced Senior Site Reliability Engineer to serve as a key technical contributor supporting the Technical Director in advancing reliability engineering, cloud operations, automation, and resilient service delivery for Department of Veterans Affairs enterprise healthcare platforms and applications. This role partners with platform, development, operations, monitoring, incident-management, security, and VA stakeholder teams to improve availability, performance, scalability, and operational excellence across mission-critical environments. The Senior Site Reliability Engineer applies software engineering principles to operations while aligning solutions with Federal security and governance requirements.\nDUTIES AND RESPONSIBILITIES:\nPartner with the Technical Director to implement and mature Site Reliability Engineering practices across platform services and hosted applications.\nImprove the full service lifecycle from design and deployment through operation and continuous refinement, with a focus on availability, latency, performance, efficiency, and capacity.\nDefine, track, and report service-level indicators, service-level objectives, error budgets, and service health measures that guide engineering decisions.\nBuild, enhance, and maintain CI/CD pipelines that enable secure, automated, repeatable application and infrastructure delivery.\nDevelop and support Infrastructure as Code and configuration automation using Terraform, Ansible, and comparable technologies.\nIntegrate automated testing, validation, security checks, rollback, and operational readiness controls into delivery workflows.\nDesign and improve monitoring, logging, tracing, alerting, and dashboards to strengthen observability and accelerate issue detection and response.\nAnalyze system behavior, performance trends, capacity, failure patterns, and operational data to improve reliability, scalability, and efficiency.\nReduce operational toil by automating repetitive tasks, improving runbooks, and engineering durable solutions for recurring issues.\nSupport AWS infrastructure and Kubernetes, EKS, ECS, Docker, or comparable container platforms with an emphasis on resilience, scalability, and security.\nContribute to platform modernization, capacity planning, reliability reviews, deployment-pattern improvements, and operational readiness for cloud-native services.\nImplement reliability practices that align with Federal security requirements, including secure configuration, least privilege, vulnerability remediation, and policy-based controls.\nCollaborate with development, platform, operations, monitoring, incident-management, architecture, and cybersecurity teams to improve service and deployment outcomes.\nParticipate in incident response, service restoration, root cause analysis, and blameless post-incident reviews for critical systems and services.\nIdentify recurring issues, reliability gaps, and failure patterns and drive corrective actions through automation, architecture improvements, and process refinement.\nStrengthen on-call readiness, operational documentation, escalation procedures, and continuous improvement practices that reduce mean time to recovery.\nKNOWLEDGE AND SKILLS REQUIRED:\n6+ years of experience in Site Reliability Engineering, DevOps, platform engineering, cloud operations, or related roles supporting enterprise or mission-critical environments.\nHands-on experience supporting AWS or comparable cloud platforms, Linux-based environments, distributed systems, and production services at scale.\nStrong experience with Infrastructure as Code and configuration automation using Terraform, Ansible, or comparable technologies.\nExperience with Kubernetes, EKS, ECS, Docker, or another container and orchestration platform in production environments.\nExperience building or maintaining CI/CD pipelines and deployment automation for secure, reliable software and infrastructure delivery.\nStrong understanding of monitoring, logging, tracing, observability, incident response, root cause analysis, capacity planning, and performance optimization.\nProficiency with one or more scripting or programming languages such as Python, Go, Bash, or PowerShell.\nDemonstrated ability to troubleshoot complex systems, automate operational tasks, reduce toil, and implement durable reliability improvements.\nExperience collaborating across software engineering, platform, operations, cybersecurity, architecture, and customer stakeholder teams.\nStrong written and verbal communication skills, including the ability to document technical standards, incidents, risks, operational decisions, and improvement plans.\nKNOWLEDGE AND SKILLS DESIRED:\nExperience supporting the Department of Veterans Affairs, another Federal agency, or a regulated enterprise environment with significant security and compliance requirements.\nExperience defining and operationalizing service-level indicators, service-level objectives, error budgets, and production service health metrics.\nAdvanced experience with Prometheus, Grafana, CloudWatch, Elasticsearch, Kibana, Splunk, OpenTelemetry, or comparable observability platforms.\nExperience with FedRAMP, NIST, Zero Trust, or other Federal security frameworks relevant to cloud and platform operations.\nExperience supporting healthcare platforms, high-availability enterprise services, or large-scale modernization initiatives.\nRelevant certification such as AWS Certified DevOps Engineer – Professional, AWS Certified Solutions Architect, Certified Kubernetes Administrator, HashiCorp Terraform Associate, or an SRE/DevOps credential.\nEDUCATION:\nBachelor’s degree in Computer Science, Engineering, Information Technology, or a related technical field, or equivalent practical experience.\nCLEARANCE:\nMust be able to obtain and maintain a Public Trust clearance.\nWORK LOCATION:\nRemote\nAs part of the application process, you are expected to be on camera during interviews and assessments. We reserve the right to take your picture to verify your identity and prevent fraud.\nThis position requires U.S. citizenship or Greencard.\nThis position is contingent upon contract award.\nTechnatomy Corporation is an Equal Opportunity Employer. It is the policy of Technatomy Corporation to afford equal employment opportunity regardless of race, color, religion, national origin, sex, age, marital status, disability or veteran status, or any other status protected by applicable law.","company":"Technatomy","rawCompany":"technatomy","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-08-03T19:34:58.076Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541519","title":"Other Computer Related Services","slug":"other-computer-related-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior Site Reliability Engineer","description":"At Technatomy, we deliver innovative solutions through the efforts of our diverse and talented people who are dedicated to our customer’s success. We provide solutions to agencies and entities including the Department of Veterans Affairs, Department of Defense, Defense Logistics Agency, National Institute of Health, and more. Everything we do is built on a commitment to do the right thing for our customers, our people, and our community. Our Mission, Vision, and Values guide the way we do business.\nIf this sounds like an environment where you can thrive, keep reading!\nWe are seeking an experienced Senior Site Reliability Engineer to serve as a key technical contributor supporting the Technical Director in advancing reliability engineering, cloud operations, automation, and resilient service delivery for Department of Veterans Affairs enterprise healthcare platforms and applications. This role partners with platform, development, operations, monitoring, incident-management, security, and VA stakeholder teams to improve availability, performance, scalability, and operational excellence across mission-critical environments. The Senior Site Reliability Engineer applies software engineering principles to operations while aligning solutions with Federal security and governance requirements.\nDUTIES AND RESPONSIBILITIES:\nPartner with the Technical Director to implement and mature Site Reliability Engineering practices across platform services and hosted applications.\nImprove the full service lifecycle from design and deployment through operation and continuous refinement, with a focus on availability, latency, performance, efficiency, and capacity.\nDefine, track, and report service-level indicators, service-level objectives, error budgets, and service health measures that guide engineering decisions.\nBuild, enhance, and maintain CI/CD pipelines that enable secure, automated, repeatable application and infrastructure delivery.\nDevelop and support Infrastructure as Code and configuration automation using Terraform, Ansible, and comparable technologies.\nIntegrate automated testing, validation, security checks, rollback, and operational readiness controls into delivery workflows.\nDesign and improve monitoring, logging, tracing, alerting, and dashboards to strengthen observability and accelerate issue detection and response.\nAnalyze system behavior, performance trends, capacity, failure patterns, and operational data to improve reliability, scalability, and efficiency.\nReduce operational toil by automating repetitive tasks, improving runbooks, and engineering durable solutions for recurring issues.\nSupport AWS infrastructure and Kubernetes, EKS, ECS, Docker, or comparable container platforms with an emphasis on resilience, scalability, and security.\nContribute to platform modernization, capacity planning, reliability reviews, deployment-pattern improvements, and operational readiness for cloud-native services.\nImplement reliability practices that align with Federal security requirements, including secure configuration, least privilege, vulnerability remediation, and policy-based controls.\nCollaborate with development, platform, operations, monitoring, incident-management, architecture, and cybersecurity teams to improve service and deployment outcomes.\nParticipate in incident response, service restoration, root cause analysis, and blameless post-incident reviews for critical systems and services.\nIdentify recurring issues, reliability gaps, and failure patterns and drive corrective actions through automation, architecture improvements, and process refinement.\nStrengthen on-call readiness, operational documentation, escalation procedures, and continuous improvement practices that reduce mean time to recovery.\nKNOWLEDGE AND SKILLS REQUIRED:\n6+ years of experience in Site Reliability Engineering, DevOps, platform engineering, cloud operations, or related roles supporting enterprise or mission-critical environments.\nHands-on experience supporting AWS or comparable cloud platforms, Linux-based environments, distributed systems, and production services at scale.\nStrong experience with Infrastructure as Code and configuration automation using Terraform, Ansible, or comparable technologies.\nExperience with Kubernetes, EKS, ECS, Docker, or another container and orchestration platform in production environments.\nExperience building or maintaining CI/CD pipelines and deployment automation for secure, reliable software and infrastructure delivery.\nStrong understanding of monitoring, logging, tracing, observability, incident response, root cause analysis, capacity planning, and performance optimization.\nProficiency with one or more scripting or programming languages such as Python, Go, Bash, or PowerShell.\nDemonstrated ability to troubleshoot complex systems, automate operational tasks, reduce toil, and implement durable reliability improvements.\nExperience collaborating across software engineering, platform, operations, cybersecurity, architecture, and customer stakeholder teams.\nStrong written and verbal communication skills, including the ability to document technical standards, incidents, risks, operational decisions, and improvement plans.\nKNOWLEDGE AND SKILLS DESIRED:\nExperience supporting the Department of Veterans Affairs, another Federal agency, or a regulated enterprise environment with significant security and compliance requirements.\nExperience defining and operationalizing service-level indicators, service-level objectives, error budgets, and production service health metrics.\nAdvanced experience with Prometheus, Grafana, CloudWatch, Elasticsearch, Kibana, Splunk, OpenTelemetry, or comparable observability platforms.\nExperience with FedRAMP, NIST, Zero Trust, or other Federal security frameworks relevant to cloud and platform operations.\nExperience supporting healthcare platforms, high-availability enterprise services, or large-scale modernization initiatives.\nRelevant certification such as AWS Certified DevOps Engineer – Professional, AWS Certified Solutions Architect, Certified Kubernetes Administrator, HashiCorp Terraform Associate, or an SRE/DevOps credential.\nEDUCATION:\nBachelor’s degree in Computer Science, Engineering, Information Technology, or a related technical field, or equivalent practical experience.\nCLEARANCE:\nMust be able to obtain and maintain a Public Trust clearance.\nWORK LOCATION:\nRemote\nAs part of the application process, you are expected to be on camera during interviews and assessments. We reserve the right to take your picture to verify your identity and prevent fraud.\nThis position requires U.S. citizenship or Greencard.\nThis position is contingent upon contract award.\nTechnatomy Corporation is an Equal Opportunity Employer. It is the policy of Technatomy Corporation to afford equal employment opportunity regardless of race, color, religion, national origin, sex, age, marital status, disability or veteran status, or any other status protected by applicable law.","datePosted":"2026-08-03T19:34:58.076Z","dateModified":"2026-08-03T19:34:58.076Z","hiringOrganization":{"@type":"Organization","name":"Technatomy","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"747516075018c87f3de91cae"},"url":"https://jobsearcher.com/jobs/747516075018c87f3de91cae"}}