{"schemaVersion":"jobsearcher.job.v1","id":"6a885dfba46d2cff1cd27daf","url":"https://jobsearcher.com/jobs/6a885dfba46d2cff1cd27daf","canonicalUrl":"https://jobsearcher.com/jobs/6a885dfba46d2cff1cd27daf","title":"Senior Platform Engineer","description":"About the Role\n\nCloudZero is growing fast. Our customer base is expanding, the data challenges we're solving are getting more complex, and the platform is scaling to match. We're standing up real-time ingestion on Kafka right now, spanning several engineering teams, and it's the most operationally demanding thing we've built. Nobody owns the reliability of that path end to end today. That's the first thing you'd own. As a Senior Site Reliability Engineer you'll be a force multiplier for our engineering organization, owning the reliability, performance, and observability of the systems every team depends on, and empowering teams to ship features that help customers understand and optimize their cloud spend.\n\nThis is real infrastructure work at real scale, not a ticket-closing role and not a console-clicking job. CloudZero processes billions of events daily across AWS, Azure, and GCP. Our customers rely on real-time, accurate cost data to make business-critical decisions, and any instability in our system impacts their planning. Built entirely on a unique serverless architecture with no EC2s and no containers, our platform demands infrastructure that scales gracefully, fails predictably, and recovers automatically. There are no Kubernetes clusters or broker fleets to tune here, so the reliability work is engineering, not firefighting. On-call is light: Nimbus carries a weekly rotation for shared infrastructure, and feature teams respond for their own services.\n\nIf you thrive on hard operational problems, care deeply about reliability and performance, and want to see your work matter to customers in direct and measurable ways, this role was built for you.\n\nWhat You'll Do\n\nReliability and Observability\n\nOwn the reliability practice for CloudZero's real-time ingestion path: SLOs that span team boundaries, the failure modes nobody owns because they live in the seams, and the architectural changes that come out of what you learn\n\nSign off on shared critical paths before they go live, and call the pause conversation when an error budget burns\n\nInstrument systems so that failures surface quickly and debugging happens with data, not guesswork\n\nBuild observability into everything so you know about problems before customers do\n\nPython and Infrastructure as Code\n\nBuild the reliability tooling, not just the recommendations: load generators, fault-injection harnesses, SLO instrumentation libraries, deployment safety checks\n\nWrite production Python across shared libraries, internal services, automation, and agents, and set the standards others build against\n\nDesign and maintain CloudFormation and SAM modules that provision reliable, cost-efficient cloud resources\n\nOwn infrastructure end to end with no clicking through consoles\n\nAutomation\n\nAutomate deployments, scaling, backups, and limit changes; if humans are doing it repeatedly, build a system to do it instead\n\nBalance automation intelligently, building solutions to real problems rather than automating for its own sake\n\nEvaluate the autonomous agents we already run in production honestly, make the good ones excellent, and throw away the ones that aren't worth it\n\nMake our systems legible to AI tooling as well as to people, starting with the service and ownership metadata in our developer portal\n\nPartner with Product Engineering\n\nHelp teams design resilient services, review architectures for operational complexity, and build deployment pipelines that enable safe and fast shipping\n\nBake SLOs and instrumentation into shared templates so teams inherit good practice instead of reinventing it\n\nDrive adoption across 40+ engineers by building the case, not by mandate\n\nOptimize for cost and performance; CloudZero's business is helping others optimize cloud costs, and we should be exemplars of efficient cloud usage ourselves\n\nWhat You Bring\n\nStrong production Python as your primary language, owned, tested, and maintained at scale\n\nAn SLO you defined yourself, including what you deliberately chose not to alert on\n\nExperience operating asynchronous, event-driven systems and reasoning about back-pressure, consumer lag, replay, poison messages, and partial failure (Kafka, Kinesis, SQS, Pulsar, or Step Functions; we run MSK, so failure reasoning matters more than broker tuning)\n\nOne reliability or platform change you drove through a team that didn't report to you and hadn't asked for it\n\nEnough production time to have owned reliability outcomes rather than reliability tasks, which usually means 5+ years building and operating distributed systems in AWS\n\nInfrastructure as Code in practice, using CloudFormation and SAM or transferable depth in Terraform or Pulumi\n\nHands-on experience instrumenting systems in monitoring tools such as Sumo Logic, Datadog, Prometheus, or Splunk, not just reading dashboards\n\nProven ability to debug production issues under pressure\n\nAn appetite for frontier AI models such as Claude, Codex, or Gemini\n\nValues thoughtful, reliable system design over reactive hero efforts\n\nStrong documentation habits to support long-term team clarity and system stability\n\nAbility to clearly explain complex technical issues to non-technical stakeholders\n\nEnergized by breadth: comfortable holding several areas at once, going deep in whichever one is currently on fire, then handing it back as a standard and moving on\n\nBonus: chaos engineering or load testing as a practice you built, internal developer portal experience (Cortex, Backstage), test automation or ephemeral test environments, GitHub Actions at scale, LLM-backed tooling real engineers used daily\n\nAbout CloudZero\nCloudZero is the AI ROI Company. We built the financial control plane for AI: the system finance, IT, and engineering use to connect every AI dollar to the outcome it produced. Across every provider. In real time.\nAI spend is the fastest-growing line on enterprise P&Ls and the least understood. Only 14% of CFOs can prove AI ROI today. CloudZero answers the question no one else can: what did it cost to produce this outcome, for this customer, on this model.\nThe largest cloud spenders on the planet already run on CloudZero, including Coinbase, Duolingo, DoorDash, and Shutterstock. We processed 14 trillion billing events in the last twelve months. We're the first listed partner on Anthropic's cost and usage API. We've raised over $119 million, including a $56 million Series C backed by leading venture capital firms.\n\nWhy Join Our Team?\nAt CloudZero, you’ll find a collaborative, fast-moving environment where your work makes a direct impact. We’re a team that values ownership, creativity, and curiosity — and we’re tackling some of the most complex challenges in the cloud space. If you’re excited by working with cutting-edge technology, driving meaningful outcomes, and growing with a company that’s scaling fast, we’d love to hear from you!\n\nCompensation Range: $130K - $190K","company":"Cloudzero","rawCompany":"cloudzero","city":"Millbrae","state":"CA","isRemote":false,"isActive":true,"createdAt":"2026-10-02T08:10:06.273Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1211.00","title":"Computer Systems Analysts","slug":"computer-systems-analysts"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior Platform Engineer","description":"About the Role\n\nCloudZero is growing fast. Our customer base is expanding, the data challenges we're solving are getting more complex, and the platform is scaling to match. We're standing up real-time ingestion on Kafka right now, spanning several engineering teams, and it's the most operationally demanding thing we've built. Nobody owns the reliability of that path end to end today. That's the first thing you'd own. As a Senior Site Reliability Engineer you'll be a force multiplier for our engineering organization, owning the reliability, performance, and observability of the systems every team depends on, and empowering teams to ship features that help customers understand and optimize their cloud spend.\n\nThis is real infrastructure work at real scale, not a ticket-closing role and not a console-clicking job. CloudZero processes billions of events daily across AWS, Azure, and GCP. Our customers rely on real-time, accurate cost data to make business-critical decisions, and any instability in our system impacts their planning. Built entirely on a unique serverless architecture with no EC2s and no containers, our platform demands infrastructure that scales gracefully, fails predictably, and recovers automatically. There are no Kubernetes clusters or broker fleets to tune here, so the reliability work is engineering, not firefighting. On-call is light: Nimbus carries a weekly rotation for shared infrastructure, and feature teams respond for their own services.\n\nIf you thrive on hard operational problems, care deeply about reliability and performance, and want to see your work matter to customers in direct and measurable ways, this role was built for you.\n\nWhat You'll Do\n\nReliability and Observability\n\nOwn the reliability practice for CloudZero's real-time ingestion path: SLOs that span team boundaries, the failure modes nobody owns because they live in the seams, and the architectural changes that come out of what you learn\n\nSign off on shared critical paths before they go live, and call the pause conversation when an error budget burns\n\nInstrument systems so that failures surface quickly and debugging happens with data, not guesswork\n\nBuild observability into everything so you know about problems before customers do\n\nPython and Infrastructure as Code\n\nBuild the reliability tooling, not just the recommendations: load generators, fault-injection harnesses, SLO instrumentation libraries, deployment safety checks\n\nWrite production Python across shared libraries, internal services, automation, and agents, and set the standards others build against\n\nDesign and maintain CloudFormation and SAM modules that provision reliable, cost-efficient cloud resources\n\nOwn infrastructure end to end with no clicking through consoles\n\nAutomation\n\nAutomate deployments, scaling, backups, and limit changes; if humans are doing it repeatedly, build a system to do it instead\n\nBalance automation intelligently, building solutions to real problems rather than automating for its own sake\n\nEvaluate the autonomous agents we already run in production honestly, make the good ones excellent, and throw away the ones that aren't worth it\n\nMake our systems legible to AI tooling as well as to people, starting with the service and ownership metadata in our developer portal\n\nPartner with Product Engineering\n\nHelp teams design resilient services, review architectures for operational complexity, and build deployment pipelines that enable safe and fast shipping\n\nBake SLOs and instrumentation into shared templates so teams inherit good practice instead of reinventing it\n\nDrive adoption across 40+ engineers by building the case, not by mandate\n\nOptimize for cost and performance; CloudZero's business is helping others optimize cloud costs, and we should be exemplars of efficient cloud usage ourselves\n\nWhat You Bring\n\nStrong production Python as your primary language, owned, tested, and maintained at scale\n\nAn SLO you defined yourself, including what you deliberately chose not to alert on\n\nExperience operating asynchronous, event-driven systems and reasoning about back-pressure, consumer lag, replay, poison messages, and partial failure (Kafka, Kinesis, SQS, Pulsar, or Step Functions; we run MSK, so failure reasoning matters more than broker tuning)\n\nOne reliability or platform change you drove through a team that didn't report to you and hadn't asked for it\n\nEnough production time to have owned reliability outcomes rather than reliability tasks, which usually means 5+ years building and operating distributed systems in AWS\n\nInfrastructure as Code in practice, using CloudFormation and SAM or transferable depth in Terraform or Pulumi\n\nHands-on experience instrumenting systems in monitoring tools such as Sumo Logic, Datadog, Prometheus, or Splunk, not just reading dashboards\n\nProven ability to debug production issues under pressure\n\nAn appetite for frontier AI models such as Claude, Codex, or Gemini\n\nValues thoughtful, reliable system design over reactive hero efforts\n\nStrong documentation habits to support long-term team clarity and system stability\n\nAbility to clearly explain complex technical issues to non-technical stakeholders\n\nEnergized by breadth: comfortable holding several areas at once, going deep in whichever one is currently on fire, then handing it back as a standard and moving on\n\nBonus: chaos engineering or load testing as a practice you built, internal developer portal experience (Cortex, Backstage), test automation or ephemeral test environments, GitHub Actions at scale, LLM-backed tooling real engineers used daily\n\nAbout CloudZero\nCloudZero is the AI ROI Company. We built the financial control plane for AI: the system finance, IT, and engineering use to connect every AI dollar to the outcome it produced. Across every provider. In real time.\nAI spend is the fastest-growing line on enterprise P&Ls and the least understood. Only 14% of CFOs can prove AI ROI today. CloudZero answers the question no one else can: what did it cost to produce this outcome, for this customer, on this model.\nThe largest cloud spenders on the planet already run on CloudZero, including Coinbase, Duolingo, DoorDash, and Shutterstock. We processed 14 trillion billing events in the last twelve months. We're the first listed partner on Anthropic's cost and usage API. We've raised over $119 million, including a $56 million Series C backed by leading venture capital firms.\n\nWhy Join Our Team?\nAt CloudZero, you’ll find a collaborative, fast-moving environment where your work makes a direct impact. We’re a team that values ownership, creativity, and curiosity — and we’re tackling some of the most complex challenges in the cloud space. If you’re excited by working with cutting-edge technology, driving meaningful outcomes, and growing with a company that’s scaling fast, we’d love to hear from you!\n\nCompensation Range: $130K - $190K","datePosted":"2026-10-02T08:10:06.273Z","dateModified":"2026-10-02T08:10:06.273Z","hiringOrganization":{"@type":"Organization","name":"Cloudzero","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"6a885dfba46d2cff1cd27daf"},"url":"https://jobsearcher.com/jobs/6a885dfba46d2cff1cd27daf"}}