Site Reliability Engineer
About RunloopRunloop.ai is pioneering the next generation of infrastructure and orchestration to power the Agentic Web/age of AI Agents. Our platform empowers developers to deploy agents that write code, browse the web, and use computers the way a human would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet, dedicated to solving the complex challenges of productionizing AI for software engineering at scale.The RoleWe're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible for the reliability, observability, performance, and security of our core platform. You'll work closely with our engineering team to develop and maintain the systems that power our code sandboxes, ensuring a seamless and stable experience for our customers. This is a critical role that blends a deep understanding of distributed systems with a software engineering mindset.ResponsibilitiesDesign and maintain our production infrastructure on cloud platforms like AWS, GCP, Azure, and emergent Neo-CloudsMonitor and respond to system alerts and incidents using Grafana and Prometheus, ensuring high availability and a secure environment for our usersCollaborate with developers to ensure new features and services are designed with scalability and reliability in mindTroubleshoot and resolve complex issues related to our infrastructure, networking, and the sandbox environmentParticipate in an on-call rotation to support our production systemsDefine and track SLIs/SLOs, manage error budgets, and proactively monitor distributed systems with logging and tracingAutomate deployments, scaling, provisioning, and recovery tasks to reduce toil and build self-healing systemsLead incident response, conduct root-cause analysis, and facilitate blameless post-mortems to drive continual improvementCollaborate cross-functionally with product, engineering, and developer relations to ensure reliable releases and an outstanding developer experiencePlan for capacity growth, forecast system usage, and contribute to safe release and change management processesQualificationsStrong computer science fundamentals, backed by a degree from a top-tier CS/EE program, or equivalent experience5+ years of experience in software engineering, with at least 3 years focused explicitly on site reliability, DevOps, or infrastructure operationsStrong programming skills in languages like Python or GoDeep expertise in containerization technologies such as Docker and KubernetesExperience with cloud infrastructure and tools like Terraform and/or PulumiFamiliarity with monitoring and alerting tools like Prometheus, Grafana, or DatadogA solid understanding of networking, security, and Linux systems administrationExperience designing, scaling, and maintaining distributed systems (backend platforms, APIs, or front-end infrastructure)Proficiency in implementing observability frameworks (metrics, logging, tracing) and aligning reliability goals with developer velocityHands-on experience managing incidents, running on-call operations, and producing actionable post-mortemsAbility to mentor engineers and influence reliability practices across teams, especially for front-end infrastructure and performanceBonus PointsExperience with chaos engineering techniques, front-end observability tools (e.g., Sentry, RUM, synthetic monitoring), or building CI/CD pipelines for front-end deliveryBenefitsCompetitive salary and equityComprehensive health, dental, and vision insurance for employee and dependentsOpportunity to work on cutting-edge technology and make a real impact on the future of software engineeringDaily catered lunch for all employees and a fridge full of your favorite snacks and drinksLocation:Onsite 4 days a week in San Francisco; Optional 1 day a week remoteJoin Us! If you're excited about shaping the future of AI-driven software engineering and empowering developers to build the next generation of AI powered coding tools, we want to hear from you. Join the Runloop team and be at the forefront of the AI revolution in software development.Runloop AI is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veteran status, sexual orientation, gender identity, or any other characteristic protected by law.