JOBSEARCHER

Staff Site Reliability Engineer

About Our ClientThe organization operates in the AI marketing platform sector, focusing on 1:1 personalization to strengthen brand and customer connections. It provides a unified platform integrating SMS, RCS, email, and push notifications, using AI and real-time behavioral data to deliver personalized customer experiences.Serving more than 8,000 customers across over 70 industries, the organization facilitates billions of customer interactions and supports leading global brands. With a distributed workforce and multiple employee hubs, the company has been recognized for both business performance and workplace culture.About the OpportunityThe Staff Site Reliability Engineer plays a strategic role in improving the reliability, scalability, and performance of the organization’s platform infrastructure. This position designs and implements solutions that strengthen observability, traceability, incident management, and platform scalability.The role also provides technical leadership across teams, mentors engineers, establishes production standards, and influences the technical roadmap to help engineering teams deliver reliable and secure solutions efficiently.Responsibilities• Design and implement systems that improve reliability, observability, traceability, and incident management• Lead strategic cross-team projects and provide technical leadership• Collaborate with AI/ML, Data, Platform, and Product teams to develop advanced services• Define and enforce production standards, processes, and tools• Establish and implement reliability metrics, including SLIs and SLOs• Mentor and guide team members to support technical growth and development• Drive continuous improvement by introducing innovative solutions and challenging existing practicesRequirements• 7+ years of experience in Production Engineering, Backend Engineering, SRE, DevOps, or a related field• Strong technical vision and ability to plan for future platform needs• Proficiency in at least one programming language, such as Golang, Python, Java, or TypeScript• Proven experience delivering medium- to large-scale projects that improve platform reliability and scalability• Deep understanding of production reliability concepts, including SLIs, SLOs, and incident management• Excellent communication skills with the ability to collaborate across technical and non-technical teams• Experience working in dynamic, reliability-focused production environments preferredPay Range and Compensation Package• US base salary range of $180,000 to $240,000 annually• Equity compensation and benefits included• Compensation may vary based on role, level, location, and other relevant factorsBenefits & Perks• Health and wellness benefits• Equity compensationEqual Opportunity Statement:Our client is an equal opportunity employer. They celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, or national origin.Note: RemoteHunter is not the Employer of Record (EOR) for this role. Our purpose in this opportunity is to connect exceptional candidates with leading employers. We help job seekers worldwide discover roles that match their goals and guide them to complete their full application directly through the hiring company’s career page or ATS.