{"schemaVersion":"jobsearcher.job.v1","id":"a7fcc02546acb2ffdf7776ef","url":"https://jobsearcher.com/jobs/a7fcc02546acb2ffdf7776ef","canonicalUrl":"https://jobsearcher.com/jobs/a7fcc02546acb2ffdf7776ef","title":"Site Reliability Engineer","description":"At Technatomy, we deliver innovative solutions through the efforts of our diverse and talented people who are dedicated to our customer’s success. We provide solutions to agencies and entities including the Department of Veterans Affairs, Department of Defense, Defense Logistics Agency, National Institute of Health, and more. Everything we do is built on a commitment to do the right thing for our customers, our people, and our community. Our Mission, Vision, and Values guide the way we do business.\nIf this sounds like an environment where you can thrive, keep reading!\nWe are seeking a motivated and detail-oriented Site Reliability Engineer to support the Technical Director’s team in advancing reliability engineering, cloud operations, automation, and resilient service delivery for Department of Veterans Affairs enterprise healthcare platforms and applications. This role works with senior engineers, platform and operations teams, and VA stakeholders to support the availability, performance, and operational stability of mission-critical environments. The Site Reliability Engineer applies foundational software engineering and operational practices to improve monitoring, automation, incident response, and service reliability.\nDUTIES AND RESPONSIBILITIES:\nSupport day-to-day Site Reliability Engineering activities across platform services, hosted applications, and cloud environments.\nHelp maintain service reliability, availability, and performance by following established operational procedures, runbooks, and engineering standards.\nGather and review operational metrics, alerts, logs, and system health information to identify issues and support service improvements.\nMaintain monitoring, logging, alerting, and dashboard configurations that improve visibility into infrastructure and application performance.\nParticipate in incident response, service restoration, escalation, and post-incident follow-up under the guidance of senior team members.\nDocument incidents, recurring issues, operational procedures, configuration details, and troubleshooting guidance.\nDevelop simple scripts and automation that reduce manual effort, improve consistency, and address recurring operational tasks.\nSupport CI/CD processes and environment maintenance for application and infrastructure delivery across development, test, and production environments.\nAssist with Infrastructure as Code, configuration changes, and environment updates using approved tools, templates, and team guidance.\nPerform routine operational checks and support activities for AWS and container-based platforms.\nMaintain service inventory, configuration records, operational documentation, and other artifacts used by the reliability team.\nAssist with validation, testing, deployment readiness, and operational acceptance activities for releases and environment changes.\nFollow established security, access, change, and operational procedures that support Federal compliance and secure administration.\nCollaborate with software, infrastructure, platform, monitoring, incident-management, and support teams to resolve issues and improve reliable service delivery.\nKNOWLEDGE AND SKILLS REQUIRED:\n1–3 years of experience in Site Reliability Engineering, DevOps, systems administration, cloud operations, platform support, software engineering, or a related technical role.\nFoundational understanding of Linux systems, cloud infrastructure concepts, enterprise application support, and basic networking.\nExposure to scripting or programming using Python, Bash, PowerShell, or a similar language.\nFamiliarity with monitoring, logging, alerting, troubleshooting, incident response, and service restoration concepts.\nBasic knowledge of CI/CD, version control, automation, configuration management, or Infrastructure as Code concepts.\nAbility to follow technical procedures, document work accurately, analyze operational information, and escalate issues appropriately.\nStrong attention to detail and the ability to learn new cloud, platform, observability, and automation tools quickly.\nAbility to work effectively in a collaborative, remote team environment with engineers, operations personnel, and customer stakeholders.\nKNOWLEDGE AND SKILLS DESIRED:\nInternship, academic, lab, or hands-on experience with AWS, Microsoft Azure, Google Cloud, or another cloud platform.\nFamiliarity with Docker, Kubernetes, EKS, ECS, or another container and orchestration technology.\nExposure to CloudWatch, Grafana, Prometheus, Elasticsearch, Kibana, Splunk, OpenTelemetry, or similar observability tools.\nExperience with Git-based workflows, pipeline tooling, or automation through coursework, labs, internships, or professional experience.\nUnderstanding of Federal security, compliance, healthcare technology, or other regulated enterprise environments.\nRelevant foundational certification such as AWS Certified Cloud Practitioner, AWS Certified Developer – Associate, CompTIA Linux+, Security+, or HashiCorp Terraform Associate.\nEDUCATION:\nBachelor’s degree in Computer Science, Information Technology, Engineering, or a related technical field, or equivalent practical experience.\nCLEARANCE:\nMust be able to obtain and maintain a Public Trust clearance.\nWORK LOCATION:\nRemote\nAs part of the application process, you are expected to be on camera during interviews and assessments. We reserve the right to take your picture to verify your identity and prevent fraud.\nThis position requires U.S. citizenship or Greencard.\nThis position is contingent upon contract award.\nTechnatomy Corporation is an Equal Opportunity Employer. It is the policy of Technatomy Corporation to afford equal employment opportunity regardless of race, color, religion, national origin, sex, age, marital status, disability or veteran status, or any other status protected by applicable law.","company":"Technatomy","rawCompany":"technatomy","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-08-03T20:13:58.350Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541519","title":"Other Computer Related Services","slug":"other-computer-related-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Site Reliability Engineer","description":"At Technatomy, we deliver innovative solutions through the efforts of our diverse and talented people who are dedicated to our customer’s success. We provide solutions to agencies and entities including the Department of Veterans Affairs, Department of Defense, Defense Logistics Agency, National Institute of Health, and more. Everything we do is built on a commitment to do the right thing for our customers, our people, and our community. Our Mission, Vision, and Values guide the way we do business.\nIf this sounds like an environment where you can thrive, keep reading!\nWe are seeking a motivated and detail-oriented Site Reliability Engineer to support the Technical Director’s team in advancing reliability engineering, cloud operations, automation, and resilient service delivery for Department of Veterans Affairs enterprise healthcare platforms and applications. This role works with senior engineers, platform and operations teams, and VA stakeholders to support the availability, performance, and operational stability of mission-critical environments. The Site Reliability Engineer applies foundational software engineering and operational practices to improve monitoring, automation, incident response, and service reliability.\nDUTIES AND RESPONSIBILITIES:\nSupport day-to-day Site Reliability Engineering activities across platform services, hosted applications, and cloud environments.\nHelp maintain service reliability, availability, and performance by following established operational procedures, runbooks, and engineering standards.\nGather and review operational metrics, alerts, logs, and system health information to identify issues and support service improvements.\nMaintain monitoring, logging, alerting, and dashboard configurations that improve visibility into infrastructure and application performance.\nParticipate in incident response, service restoration, escalation, and post-incident follow-up under the guidance of senior team members.\nDocument incidents, recurring issues, operational procedures, configuration details, and troubleshooting guidance.\nDevelop simple scripts and automation that reduce manual effort, improve consistency, and address recurring operational tasks.\nSupport CI/CD processes and environment maintenance for application and infrastructure delivery across development, test, and production environments.\nAssist with Infrastructure as Code, configuration changes, and environment updates using approved tools, templates, and team guidance.\nPerform routine operational checks and support activities for AWS and container-based platforms.\nMaintain service inventory, configuration records, operational documentation, and other artifacts used by the reliability team.\nAssist with validation, testing, deployment readiness, and operational acceptance activities for releases and environment changes.\nFollow established security, access, change, and operational procedures that support Federal compliance and secure administration.\nCollaborate with software, infrastructure, platform, monitoring, incident-management, and support teams to resolve issues and improve reliable service delivery.\nKNOWLEDGE AND SKILLS REQUIRED:\n1–3 years of experience in Site Reliability Engineering, DevOps, systems administration, cloud operations, platform support, software engineering, or a related technical role.\nFoundational understanding of Linux systems, cloud infrastructure concepts, enterprise application support, and basic networking.\nExposure to scripting or programming using Python, Bash, PowerShell, or a similar language.\nFamiliarity with monitoring, logging, alerting, troubleshooting, incident response, and service restoration concepts.\nBasic knowledge of CI/CD, version control, automation, configuration management, or Infrastructure as Code concepts.\nAbility to follow technical procedures, document work accurately, analyze operational information, and escalate issues appropriately.\nStrong attention to detail and the ability to learn new cloud, platform, observability, and automation tools quickly.\nAbility to work effectively in a collaborative, remote team environment with engineers, operations personnel, and customer stakeholders.\nKNOWLEDGE AND SKILLS DESIRED:\nInternship, academic, lab, or hands-on experience with AWS, Microsoft Azure, Google Cloud, or another cloud platform.\nFamiliarity with Docker, Kubernetes, EKS, ECS, or another container and orchestration technology.\nExposure to CloudWatch, Grafana, Prometheus, Elasticsearch, Kibana, Splunk, OpenTelemetry, or similar observability tools.\nExperience with Git-based workflows, pipeline tooling, or automation through coursework, labs, internships, or professional experience.\nUnderstanding of Federal security, compliance, healthcare technology, or other regulated enterprise environments.\nRelevant foundational certification such as AWS Certified Cloud Practitioner, AWS Certified Developer – Associate, CompTIA Linux+, Security+, or HashiCorp Terraform Associate.\nEDUCATION:\nBachelor’s degree in Computer Science, Information Technology, Engineering, or a related technical field, or equivalent practical experience.\nCLEARANCE:\nMust be able to obtain and maintain a Public Trust clearance.\nWORK LOCATION:\nRemote\nAs part of the application process, you are expected to be on camera during interviews and assessments. We reserve the right to take your picture to verify your identity and prevent fraud.\nThis position requires U.S. citizenship or Greencard.\nThis position is contingent upon contract award.\nTechnatomy Corporation is an Equal Opportunity Employer. It is the policy of Technatomy Corporation to afford equal employment opportunity regardless of race, color, religion, national origin, sex, age, marital status, disability or veteran status, or any other status protected by applicable law.","datePosted":"2026-08-03T20:13:58.350Z","dateModified":"2026-08-03T20:13:58.350Z","hiringOrganization":{"@type":"Organization","name":"Technatomy","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"a7fcc02546acb2ffdf7776ef"},"url":"https://jobsearcher.com/jobs/a7fcc02546acb2ffdf7776ef"}}