JOBSEARCHER

Site Reliability Engineer

ARCHIVED

We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.

As a Site Reliability Engineer (SRE) at Together, you are responsible for keeping all user-facing services and production systems running smoothly. You are a blend of a pragmatic operator and a software engineer that applies sound engineering principles, operational discipline, and mature automation to our operating environments and codebase.Read on to fully understand what this job requires in terms of skills and experience If you are a good match, make an application.You specialize in systems (operating systems, storage subsystems, networking), while implementing best practices for availability, reliability and scalability, with varied interests in algorithms and distributed systems.Requirements2+ years of professional SRE or related experienceBachelor's degree in Computer Science or a related field or equivalent work experienceKnowledge of Ansible (roles, playbooks), Terraform, and KubernetesProficiency in programming/scripting languagesDirect experience in monitoring and observability practicesKnowledge of cloud servicesAbility to thrive in a collaborative environment involving different stakeholders and subject matter expertsResponsibilitiesParticipate in on-call rotation (Pagerduty) to respond to production incidentsBuild and run our infrastructure with Ansible, Terraform, and Kubernetes to enable scaling to a massive number of concurrent usersBuild monitoring systems to ensure the highest quality service for our customersDesign and implement operational processes (such as deployments and upgrades)Debug production issues across all services and levels of the stackIdentify improvements for the product architecture from the reliability, performance and availability perspectivesPlan the growth of Together AI's infrastructureAbout Together AITogether AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers and engineers in our journey in building the next generation AI infrastructure.CompensationWe offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $150,000 - $200,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.Equal OpportunityTogether AI is an Equal Opportunity Employer and is proud to offer employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. xmcpwfuInterested in building your career at Together AI? Get future opportunities sent straight to your email.