{"schemaVersion":"jobsearcher.job.v1","id":"f3341e1c5e9aa6fd45e2ccd0","url":"https://jobsearcher.com/jobs/f3341e1c5e9aa6fd45e2ccd0","canonicalUrl":"https://jobsearcher.com/jobs/f3341e1c5e9aa6fd45e2ccd0","title":"Senior Site Reliability Engineer","description":"Location San Francisco, CA\nEmployment Type Full time\nDepartment Engineering\nWho We Are Hyperbolic Labs is on a mission to democratize AI by breaking down the barriers to computing power with our Open-Access AI Cloud. By aggregating computing resources across the globe, we offer an innovative GPU marketplace and AI inference service that promise affordability and accessibility for all. As pioneers at the intersection of AI and open-source technology, we believe in an open future where AI innovation is limited only by imagination, not by access to resources. We\\'re looking for forward-thinking individuals who share our passion for making AI universally accessible, secure, and affordable. Join us in building a platform that empowers innovators everywhere to turn their visionary AI projects into reality.\nAs we prepare for growth after our Series A, our team — led by co-founders with PhDs in AI, Math, and Computer Science — is poised to redefine computing.\nAbout the Role We\\'re seeking a Site Reliability Engineer to ensure Hyperbolic\\'s GPU marketplace and AI infrastructure operate with exceptional reliability, performance, and security. As an aggregator of compute resources from hundreds of global suppliers, our SLOs, trust, and economic efficiency are product-critical. You\\'ll be responsible for defining and maintaining service level objectives for job success rates, building robust incident response systems, managing capacity across our distributed GPU network, and implementing secure rollout and rollback mechanisms that keep our platform running smoothly 24/7.\nIn this role, you\\'ll establish the reliability standards that define customer trust in our platform, design monitoring and alerting systems that provide deep visibility into our infrastructure, build automation for capacity management and resource allocation, lead incident response and post-mortem processes, and work closely with engineering teams to improve system resilience. You\\'ll also focus on security and infrastructure hardening, ensuring strong isolation between tenants and suppliers, implementing key management systems, and building compliance frameworks. This is a high-impact position where your work directly influences our ability to deliver on our promise of affordable, accessible AI compute at scale.\nWho You Are Expert in site reliability engineering with proven experience defining, monitoring, and maintaining SLOs and SLAs for production systems\n\nStrong background in capacity planning and management, including forecasting, resource allocation, and cost optimization for distributed systems\n\nExperienced in incident response, on-call rotations, and post-mortem processes with a track record of reducing MTTR and improving system resilience\n\nDeep knowledge of deployment systems including progressive rollouts, canary deployments, feature flags, and automated rollback mechanisms\n\nProficient in observability tools and practices including metrics, logging, tracing, and alerting systems (Prometheus, Grafana, ELK stack, or similar)\n\nStrong understanding of infrastructure security including tenant isolation, workload isolation, network segmentation, and security hardening\n\nExperience with secrets management, key management systems (KMS), certificate management, and secure credential rotation\n\nKnowledge of compliance frameworks and security best practices for cloud platforms (SOC 2, ISO 27001, or similar)\n\nExcellent problem-solving skills with ability to debug complex distributed systems issues under pressure\n\nStrong automation mindset with experience using infrastructure-as-code, configuration management, and CI/CD pipelines\n\nPreferred Qualifications Experience operating GPU infrastructure, AI/ML platforms, or compute marketplaces at scale\n\nBackground in distributed systems, peer-to-peer networks, or decentralized infrastructure\n\nKnowledge of multi-tenancy security patterns, container security, and runtime security tools\n\nExperience with chaos engineering, fault injection, and resilience testing\n\nFamiliarity with cost optimization strategies for cloud infrastructure and GPU resources\n\nExperience building and operating systems with demanding uptime requirements (99.9%+ SLAs)\n\nBackground at companies like AWS, Google Cloud, Azure, or fast-growing infrastructure startups\n\nContributions to open-source reliability, observability, or security tools\n\nHyperbolic is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.\n\n#J-18808-Ljbffr","company":"Hyperbolic","rawCompany":"hyperbolic","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-16T03:59:47.392Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior Site Reliability Engineer","description":"Location San Francisco, CA\nEmployment Type Full time\nDepartment Engineering\nWho We Are Hyperbolic Labs is on a mission to democratize AI by breaking down the barriers to computing power with our Open-Access AI Cloud. By aggregating computing resources across the globe, we offer an innovative GPU marketplace and AI inference service that promise affordability and accessibility for all. As pioneers at the intersection of AI and open-source technology, we believe in an open future where AI innovation is limited only by imagination, not by access to resources. We\\'re looking for forward-thinking individuals who share our passion for making AI universally accessible, secure, and affordable. Join us in building a platform that empowers innovators everywhere to turn their visionary AI projects into reality.\nAs we prepare for growth after our Series A, our team — led by co-founders with PhDs in AI, Math, and Computer Science — is poised to redefine computing.\nAbout the Role We\\'re seeking a Site Reliability Engineer to ensure Hyperbolic\\'s GPU marketplace and AI infrastructure operate with exceptional reliability, performance, and security. As an aggregator of compute resources from hundreds of global suppliers, our SLOs, trust, and economic efficiency are product-critical. You\\'ll be responsible for defining and maintaining service level objectives for job success rates, building robust incident response systems, managing capacity across our distributed GPU network, and implementing secure rollout and rollback mechanisms that keep our platform running smoothly 24/7.\nIn this role, you\\'ll establish the reliability standards that define customer trust in our platform, design monitoring and alerting systems that provide deep visibility into our infrastructure, build automation for capacity management and resource allocation, lead incident response and post-mortem processes, and work closely with engineering teams to improve system resilience. You\\'ll also focus on security and infrastructure hardening, ensuring strong isolation between tenants and suppliers, implementing key management systems, and building compliance frameworks. This is a high-impact position where your work directly influences our ability to deliver on our promise of affordable, accessible AI compute at scale.\nWho You Are Expert in site reliability engineering with proven experience defining, monitoring, and maintaining SLOs and SLAs for production systems\n\nStrong background in capacity planning and management, including forecasting, resource allocation, and cost optimization for distributed systems\n\nExperienced in incident response, on-call rotations, and post-mortem processes with a track record of reducing MTTR and improving system resilience\n\nDeep knowledge of deployment systems including progressive rollouts, canary deployments, feature flags, and automated rollback mechanisms\n\nProficient in observability tools and practices including metrics, logging, tracing, and alerting systems (Prometheus, Grafana, ELK stack, or similar)\n\nStrong understanding of infrastructure security including tenant isolation, workload isolation, network segmentation, and security hardening\n\nExperience with secrets management, key management systems (KMS), certificate management, and secure credential rotation\n\nKnowledge of compliance frameworks and security best practices for cloud platforms (SOC 2, ISO 27001, or similar)\n\nExcellent problem-solving skills with ability to debug complex distributed systems issues under pressure\n\nStrong automation mindset with experience using infrastructure-as-code, configuration management, and CI/CD pipelines\n\nPreferred Qualifications Experience operating GPU infrastructure, AI/ML platforms, or compute marketplaces at scale\n\nBackground in distributed systems, peer-to-peer networks, or decentralized infrastructure\n\nKnowledge of multi-tenancy security patterns, container security, and runtime security tools\n\nExperience with chaos engineering, fault injection, and resilience testing\n\nFamiliarity with cost optimization strategies for cloud infrastructure and GPU resources\n\nExperience building and operating systems with demanding uptime requirements (99.9%+ SLAs)\n\nBackground at companies like AWS, Google Cloud, Azure, or fast-growing infrastructure startups\n\nContributions to open-source reliability, observability, or security tools\n\nHyperbolic is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.\n\n#J-18808-Ljbffr","datePosted":"2026-07-16T03:59:47.392Z","dateModified":"2026-07-16T03:59:47.392Z","hiringOrganization":{"@type":"Organization","name":"Hyperbolic","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"f3341e1c5e9aa6fd45e2ccd0"},"url":"https://jobsearcher.com/jobs/f3341e1c5e9aa6fd45e2ccd0"}}