{"schemaVersion":"jobsearcher.job.v1","id":"77624ff4bf05e2c0a7151236","url":"https://jobsearcher.com/jobs/77624ff4bf05e2c0a7151236","canonicalUrl":"https://jobsearcher.com/jobs/77624ff4bf05e2c0a7151236","title":"Site Reliability Engineer - DevOps","description":"We are looking for a technical expert in building out and refining the DevOps discipline within an enterprise SaaS environment, someone who is focused on continuous integration, continuous deployment and promoting the productivity of all of Engineering. As a Site Reliability Engineer, you will be building, evolving, and operating the infrastructure automation platform used to power our Services. You will need to ensure that our production environment is operating and performing optimized and efficiently; and that software is released and deployed in an efficient and streamlined manner, from development all the way to production. This is a hands-on operational role with a balanced amount of tool and infrastructure development, including advanced scripting and automation. You will be supporting our systems/ tools/ processes infrastructure – on premises or external cloud, and support the entire stack for our service offering.Responsibilities:Champion the Site Reliability (DevOps) needs for continuous integration and continuous deployment while maintaining focus on Quality of Service.Work with Infrastructure architects to deploy and operate cloud services and related projects from development to productionLead the efforts to improve the existing server and configuration management automation, identify opportunities to improve overall productivity and investigate tools that might speed up the process or make us more efficient in continuous integration continuous deployment. Assist in the roll-out and deployment of new product features and services to facilitate our rapid iteration and constant growth. making it faster and easier to create and deploy software.Bridge Engineering and core shared operations servicesMaintain consistent system performance. This means being up and available, as well as fast and reliable. Participate in troubleshooting, capacity planning and analysis, performance analysis, infrastructure improvement activitiesTroubleshoot issues across the whole stack - hardware, software, applications and network.Take part in a 24x7 on-call rotation.Required Qualifications:Assertiveness & creative ideas are mandatory.Excellent interpersonal skills suitable for user support, including the ability to lead projects with peer-level engineers/managers.Exceptional communication skills – both written and oral (one-on-one and group).Strong analytical, problem-solving, and decision-making skills.Must have self-starting personality, unafraid to display initiative and innovation on the job.Solid understanding and experience working with high availability, high performance, multi-data center systems.Three or more years' experience in supporting internet fronted infrastructure with Cloud, building and running large-scale web production systems, troubleshooting problems as well as improving the reliability of systems.Experience with scaling services out horizontally as well as vertically. Build, monitor, troubleshoot and manage production, testing and development environments.(VMWare ESX or other equivalent hypervisor)Knowledge on Cloud-based services - preferably AWS including but not limited to: EC2, S3, Cloudfront, EBS, SQS, etc.Extensive knowledge of Windows and Unix/Linux systems including hardware, software and applications.Working knowledge of configuration management tools to help you manage software and system changes repeatedly and predictably (Puppet, Chef, Docker, Salt, Automated Testing).Experience with system deployment and automation with scripting like Python, Shell, or Perl.Some experience in analysis and building metrics gathering systems (Splunk and AppDynamics)Knowledge of build/continuous integration tools (Hudson, Jenkins).BS/BA (4 yr) or higher in Computer Science or a related field.Desired Qualifications:Ability to manage multiple projects with competing priorities.A track record of maintaining and improving skills in existing and emerging open source technologies through training or self-research.Love to learn, enjoy troubleshooting and thinking through complicated problems.Ability to train others on new technologies and processes.Ability to manage time effectively in a fast-paced, customer-focused, changing environment.Comfortable with collaboration, open communication and reaching across functional borders.Willingness and ability to maintain a positive, quality-oriented, reliable, and flexible attitude.Willingness and ability to do what it takes to achieve objectives, including off-hours support or tasks.","company":"Stride Search","rawCompany":"stride search","city":"Sunnyvale","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-27T13:37:05.849Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541519","title":"Other Computer Related Services","slug":"other-computer-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Site Reliability Engineer - DevOps","description":"We are looking for a technical expert in building out and refining the DevOps discipline within an enterprise SaaS environment, someone who is focused on continuous integration, continuous deployment and promoting the productivity of all of Engineering. As a Site Reliability Engineer, you will be building, evolving, and operating the infrastructure automation platform used to power our Services. You will need to ensure that our production environment is operating and performing optimized and efficiently; and that software is released and deployed in an efficient and streamlined manner, from development all the way to production. This is a hands-on operational role with a balanced amount of tool and infrastructure development, including advanced scripting and automation. You will be supporting our systems/ tools/ processes infrastructure – on premises or external cloud, and support the entire stack for our service offering.Responsibilities:Champion the Site Reliability (DevOps) needs for continuous integration and continuous deployment while maintaining focus on Quality of Service.Work with Infrastructure architects to deploy and operate cloud services and related projects from development to productionLead the efforts to improve the existing server and configuration management automation, identify opportunities to improve overall productivity and investigate tools that might speed up the process or make us more efficient in continuous integration continuous deployment. Assist in the roll-out and deployment of new product features and services to facilitate our rapid iteration and constant growth. making it faster and easier to create and deploy software.Bridge Engineering and core shared operations servicesMaintain consistent system performance. This means being up and available, as well as fast and reliable. Participate in troubleshooting, capacity planning and analysis, performance analysis, infrastructure improvement activitiesTroubleshoot issues across the whole stack - hardware, software, applications and network.Take part in a 24x7 on-call rotation.Required Qualifications:Assertiveness & creative ideas are mandatory.Excellent interpersonal skills suitable for user support, including the ability to lead projects with peer-level engineers/managers.Exceptional communication skills – both written and oral (one-on-one and group).Strong analytical, problem-solving, and decision-making skills.Must have self-starting personality, unafraid to display initiative and innovation on the job.Solid understanding and experience working with high availability, high performance, multi-data center systems.Three or more years' experience in supporting internet fronted infrastructure with Cloud, building and running large-scale web production systems, troubleshooting problems as well as improving the reliability of systems.Experience with scaling services out horizontally as well as vertically. Build, monitor, troubleshoot and manage production, testing and development environments.(VMWare ESX or other equivalent hypervisor)Knowledge on Cloud-based services - preferably AWS including but not limited to: EC2, S3, Cloudfront, EBS, SQS, etc.Extensive knowledge of Windows and Unix/Linux systems including hardware, software and applications.Working knowledge of configuration management tools to help you manage software and system changes repeatedly and predictably (Puppet, Chef, Docker, Salt, Automated Testing).Experience with system deployment and automation with scripting like Python, Shell, or Perl.Some experience in analysis and building metrics gathering systems (Splunk and AppDynamics)Knowledge of build/continuous integration tools (Hudson, Jenkins).BS/BA (4 yr) or higher in Computer Science or a related field.Desired Qualifications:Ability to manage multiple projects with competing priorities.A track record of maintaining and improving skills in existing and emerging open source technologies through training or self-research.Love to learn, enjoy troubleshooting and thinking through complicated problems.Ability to train others on new technologies and processes.Ability to manage time effectively in a fast-paced, customer-focused, changing environment.Comfortable with collaboration, open communication and reaching across functional borders.Willingness and ability to maintain a positive, quality-oriented, reliable, and flexible attitude.Willingness and ability to do what it takes to achieve objectives, including off-hours support or tasks.","datePosted":"2026-08-27T13:37:05.849Z","dateModified":"2026-08-27T13:37:05.849Z","hiringOrganization":{"@type":"Organization","name":"Stride Search","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Sunnyvale","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"77624ff4bf05e2c0a7151236"},"url":"https://jobsearcher.com/jobs/77624ff4bf05e2c0a7151236"}}