JOBSEARCHER

Senior CloudOps Engineer

CloudzeroQuincy, MAL6 LeadSeptember 15th, 2026
Overview As a Senior Site Reliability Engineer at CloudZero, you own the reliability, performance, and observability of our real-time ingestion path, enabling cross-team shipping of features that reveal and optimize cloud spend. You’ll work within a serverless, scale-focused platform handling billions of events daily, without traditional clusters, to ensure fast, accurate cost data delivery. Your work directly impacts customers’ planning and financial decisions by making systems predictable and instrumented. This role offers meaningful ownership and real impact at scale. Compensation / BenefitsEquityHybrid work location (Boston or San Francisco)Competitive compensation ($130K-$190K)Full-time employmentOpportunities to impact AI ROI at scaleGrowth and ownership in a fast-moving team ResponsibilitiesOwn reliability practice for real-time ingestion path across teams, including SLOs and failure mode managementSign off on critical paths and pause on elevated error budgetsInstrument systems to surface failures with data-driven debuggingBuild and integrate observability into all components; develop instrumentation libraries and toolingDevelop reliability tooling (load generators, fault injection, deployment safety checks) and production Python across libraries and servicesDesign and maintain CloudFormation and SAM modules for reliable cloud resourcesOwn end-to-end infrastructure without manual console workAutomate deployments, scaling, backups, and change limits; reduce manual toilEvaluate autonomous agents in production and iteratively improve or discard themEnhance service visibility for AI tooling and developer portalsCollaborate with product engineering to design resilient services and safe shipping pipelinesEmbed SLOs and instrumentation into shared templates for cross-team adoptionDrive cost-performance optimization and demonstrate efficient cloud usage Key requirementsStrong production Python experience, maintained at scaleDesigned or defined an SLO with deliberate alerting decisionsExperience with asynchronous, event-driven systems and back-pressure concepts (Kafka, Kinesis, SQS, Pulsar, or Step Functions; MSK experience emphasized)Proven influence on reliability changes across teams5+ years building and operating distributed systems in AWSInfrastructure as Code in practice (CloudFormation and SAM, or Terraform/Pulumi)Hands-on monitoring experience (Sumo Logic, Datadog, Prometheus, or Splunk)Ability to debug production issues under pressureInterest in frontier AI models (e.g., Claude, Codex, Gemini)Thoughtful, non-reactive system designStrong documentation for long-term clarityAbility to explain technical issues to non-technical stakeholdersBreadth-oriented: manage multiple domains, deep dive when neededBonus: chaos engineering or load testing, internal developer portal experience (Cortex, Backstage), test automation or ephemeral environments, GitHub Actions at scale, LLM-backed toolingThoughtful and reliable system designClear communication with non-technical stakeholdersStrong documentation habitsPython production-grade developmentSLO design and instrumentationEvent-driven architectures and Kafka/MSK experience