{"schemaVersion":"jobsearcher.job.v1","id":"dfe319f765ba5a8f45a0ca49","url":"https://jobsearcher.com/jobs/dfe319f765ba5a8f45a0ca49","canonicalUrl":"https://jobsearcher.com/jobs/dfe319f765ba5a8f45a0ca49","title":"Sr Cloud Reliability Engineer (26861)","description":"Sr Cloud Reliability Engineer (26861)\n\nDate: Mar 26, 2026\nLocation: San Jose, California, United States\nCompany: Super Micro Computer\nAbout Supermicro:\nSupermicro® is a Top Tier provider of advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing company among the Silicon Valley Top 50 technology firms. Our unprecedented global expansion has provided us with the opportunity to offer a large number of new positions to the technology community. We seek talented, passionate, and committed engineers, technologists, and business leaders to join us.\n\nJob Req ID: 26861\nJob Summary:\nAs a Sr Cloud Reliability Engineer for our Linux-based AI cloud platforms, you will help us deploy, scale, automate, and ensure high availability, performance, scalability, and security across GPU-accelerated compute clusters, Kubernetes workloads, and supporting storage/network infrastructure. You will bridge Development and Operations by automating infrastructure deployment, enhancing observability, and applying SRE best practices to support reliable AI development environments, AI developer tools, and MLOps platforms in our on-premises and hybrid cloud environments.\nThis role will also support the operational enablement of AI development by helping manage AI platform services, developer access, API key governance, MCP servers, and other shared tooling required for secure and reliable use of modern AI workflows.\nEssential Duties and Responsibilities:\nIncludes the following essential duties and responsibilities (other duties may also be assigned):\n\nCloud Infra Automation: Design and provision cloud infrastructure using Infrastructure as Code (Terraform, Ansible, or Helm) on bare metal or cloud platforms. Develop custom automation and tooling in Python or Go to extend deployment workflows and streamline operations.\nPlatform Reliability: Deploy, scale, maintain, and optimize uptime for AI cloud services including GPU clusters, Kubernetes (K8s), and storage systems (e.g., Ceph, Vast, DDN, or Weka). Understand the tools required to benchmark and assure consistent application performance.\nAI Platform Operations: Support reliable operation of shared AI platform services used by development teams, including model-serving infrastructure, AI development environments, inference services, and internal AI tooling deployed on GPU-based on-premises infrastructure.\nAI Developer Tooling Support: Help enable and support secure enterprise use of AI developer tools and services, including API-based AI platforms, developer integrations, and related service configurations used in software development workflows.\nMCP / Tool Integration Services: Deploy, configure, secure, and support MCP servers and related service endpoints that enable controlled integration between AI tools, development environments, internal systems, and approved data/services.\nAccess and API Key Governance: Establish and help operate secure processes for AI service credentials, API keys, tokens, secrets rotation, usage controls, and environment-based access management to support development, testing, and production governance.\nMonitoring & Alerting: Implement observability tools (e.g., Prometheus, Grafana, ELK, Loki, Fluentd) to monitor system health and alert on anomalies, service degradation, or abnormal GPU / AI platform behavior.\nCapacity Planning: Analyze usage trends and forecast infrastructure needs to support AI workloads, large-scale model training/inference, and shared developer platform demand.\nIncident Management: Lead root cause analysis and resolution for system outages or degraded performance. Define and maintain service level objectives (SLOs), indicators (SLIs), and agreements (SLAs) aligned with uptime and performance goals.\nCI/CD Integration: Collaborate with DevOps and MLOps teams to ensure reliable delivery pipelines using GitLab CI/CD, ArgoCD, or similar tools.\nSecurity & Compliance: Harden Linux systems, manage TLS certificates, and enforce secure access controls via Role-Based Access Control (RBAC), LDAP-integrated SSO, TLS, secrets management, and network segmentation policies.\nDocumentation & Playbooks: Maintain clear, version-controlled documentation, including architecture diagrams, runbooks, API key management procedures, MCP service standards, and incident response playbooks to support cross-team knowledge transfer and rapid onboarding.\nQualifications:\nBachelor’s degree in Computer Science, Engineering, or a related field—or equivalent experience and 8 years of experience in the areas below\nProficiency in Linux (Ubuntu, RHEL/CentOS), containers (Docker, Podman), and orchestration (Kubernetes)\nExperience managing GPU compute clusters (NVIDIA / CUDA, AMD / ROCm)\nHands-on experience with observability tools (Prometheus, Grafana, Loki, ELK, etc.)\nStrong scripting and coding skills (Bash, Python, or Go)\nExperience supporting shared platform services for developers in production environments\nFamiliarity with AI development workflows and the operational needs of teams building with AI tools, APIs, and model services\nExperience with secrets, credentials, certificates, or API key management in enterprise environments\nExposure to secure multi-tenant environments and zero trust architecture\nFamiliarity with network protocols, DNS, DHCP, BGP, RoCEv2, and InfiniBand or high-throughput Ethernet fabrics\nExcellent collaboration and communication skills for cross-team, partner, and customer initiatives\nPreferred Qualifications:\nUnderstanding of AI/ML reference architectures and experience with workflows, MLflow, or Kubeflow\nFamiliarity with AI developer platforms, model APIs, inference services, and secure integration patterns for enterprise AI use cases\nExperience deploying or supporting MCP servers or similar integration services for AI tool ecosystems\nFamiliarity with API key lifecycle management, secrets vaults, token-based authentication, and environment-based access controls\nFamiliarity with storage backends optimized for AI\nPrior experience in bare-metal provisioning via PXE, Ironic, or Foreman\nUnderstanding of NVIDIA GPU telemetry and NCCL testing for performance benchmarking\nFamiliarity with ITIL processes or structured change management in production systems is a plus\nCertifications: CKA, CKAD, Linux+, or related credentials\nSalary Range\n$145,000 - $165,000\nThe salary offered will depend on several factors, including your location, level, education, training, specific skills, years of experience, and comparison to other employees already in this role. In addition to a comprehensive benefits package, candidates may be eligible for other forms of compensation, such as participation in bonus and equity award programs.\nEEO Statement\nSupermicro is an Equal Opportunity Employer and embraces diversity in our employee population. It is the policy of Supermicro to provide equal opportunity to all qualified applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, protected veteran status or special disabled veteran, marital status, pregnancy, genetic information, or any other legally protected status.","company":"Super Micro Computer","rawCompany":"super micro computer","city":"San Jose","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-15T11:50:40.977Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"11-3021.00","title":"Computer and Information Systems Managers","slug":"computer-and-information-systems-managers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Sr Cloud Reliability Engineer (26861)","description":"Sr Cloud Reliability Engineer (26861)\n\nDate: Mar 26, 2026\nLocation: San Jose, California, United States\nCompany: Super Micro Computer\nAbout Supermicro:\nSupermicro® is a Top Tier provider of advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing company among the Silicon Valley Top 50 technology firms. Our unprecedented global expansion has provided us with the opportunity to offer a large number of new positions to the technology community. We seek talented, passionate, and committed engineers, technologists, and business leaders to join us.\n\nJob Req ID: 26861\nJob Summary:\nAs a Sr Cloud Reliability Engineer for our Linux-based AI cloud platforms, you will help us deploy, scale, automate, and ensure high availability, performance, scalability, and security across GPU-accelerated compute clusters, Kubernetes workloads, and supporting storage/network infrastructure. You will bridge Development and Operations by automating infrastructure deployment, enhancing observability, and applying SRE best practices to support reliable AI development environments, AI developer tools, and MLOps platforms in our on-premises and hybrid cloud environments.\nThis role will also support the operational enablement of AI development by helping manage AI platform services, developer access, API key governance, MCP servers, and other shared tooling required for secure and reliable use of modern AI workflows.\nEssential Duties and Responsibilities:\nIncludes the following essential duties and responsibilities (other duties may also be assigned):\n\nCloud Infra Automation: Design and provision cloud infrastructure using Infrastructure as Code (Terraform, Ansible, or Helm) on bare metal or cloud platforms. Develop custom automation and tooling in Python or Go to extend deployment workflows and streamline operations.\nPlatform Reliability: Deploy, scale, maintain, and optimize uptime for AI cloud services including GPU clusters, Kubernetes (K8s), and storage systems (e.g., Ceph, Vast, DDN, or Weka). Understand the tools required to benchmark and assure consistent application performance.\nAI Platform Operations: Support reliable operation of shared AI platform services used by development teams, including model-serving infrastructure, AI development environments, inference services, and internal AI tooling deployed on GPU-based on-premises infrastructure.\nAI Developer Tooling Support: Help enable and support secure enterprise use of AI developer tools and services, including API-based AI platforms, developer integrations, and related service configurations used in software development workflows.\nMCP / Tool Integration Services: Deploy, configure, secure, and support MCP servers and related service endpoints that enable controlled integration between AI tools, development environments, internal systems, and approved data/services.\nAccess and API Key Governance: Establish and help operate secure processes for AI service credentials, API keys, tokens, secrets rotation, usage controls, and environment-based access management to support development, testing, and production governance.\nMonitoring & Alerting: Implement observability tools (e.g., Prometheus, Grafana, ELK, Loki, Fluentd) to monitor system health and alert on anomalies, service degradation, or abnormal GPU / AI platform behavior.\nCapacity Planning: Analyze usage trends and forecast infrastructure needs to support AI workloads, large-scale model training/inference, and shared developer platform demand.\nIncident Management: Lead root cause analysis and resolution for system outages or degraded performance. Define and maintain service level objectives (SLOs), indicators (SLIs), and agreements (SLAs) aligned with uptime and performance goals.\nCI/CD Integration: Collaborate with DevOps and MLOps teams to ensure reliable delivery pipelines using GitLab CI/CD, ArgoCD, or similar tools.\nSecurity & Compliance: Harden Linux systems, manage TLS certificates, and enforce secure access controls via Role-Based Access Control (RBAC), LDAP-integrated SSO, TLS, secrets management, and network segmentation policies.\nDocumentation & Playbooks: Maintain clear, version-controlled documentation, including architecture diagrams, runbooks, API key management procedures, MCP service standards, and incident response playbooks to support cross-team knowledge transfer and rapid onboarding.\nQualifications:\nBachelor’s degree in Computer Science, Engineering, or a related field—or equivalent experience and 8 years of experience in the areas below\nProficiency in Linux (Ubuntu, RHEL/CentOS), containers (Docker, Podman), and orchestration (Kubernetes)\nExperience managing GPU compute clusters (NVIDIA / CUDA, AMD / ROCm)\nHands-on experience with observability tools (Prometheus, Grafana, Loki, ELK, etc.)\nStrong scripting and coding skills (Bash, Python, or Go)\nExperience supporting shared platform services for developers in production environments\nFamiliarity with AI development workflows and the operational needs of teams building with AI tools, APIs, and model services\nExperience with secrets, credentials, certificates, or API key management in enterprise environments\nExposure to secure multi-tenant environments and zero trust architecture\nFamiliarity with network protocols, DNS, DHCP, BGP, RoCEv2, and InfiniBand or high-throughput Ethernet fabrics\nExcellent collaboration and communication skills for cross-team, partner, and customer initiatives\nPreferred Qualifications:\nUnderstanding of AI/ML reference architectures and experience with workflows, MLflow, or Kubeflow\nFamiliarity with AI developer platforms, model APIs, inference services, and secure integration patterns for enterprise AI use cases\nExperience deploying or supporting MCP servers or similar integration services for AI tool ecosystems\nFamiliarity with API key lifecycle management, secrets vaults, token-based authentication, and environment-based access controls\nFamiliarity with storage backends optimized for AI\nPrior experience in bare-metal provisioning via PXE, Ironic, or Foreman\nUnderstanding of NVIDIA GPU telemetry and NCCL testing for performance benchmarking\nFamiliarity with ITIL processes or structured change management in production systems is a plus\nCertifications: CKA, CKAD, Linux+, or related credentials\nSalary Range\n$145,000 - $165,000\nThe salary offered will depend on several factors, including your location, level, education, training, specific skills, years of experience, and comparison to other employees already in this role. In addition to a comprehensive benefits package, candidates may be eligible for other forms of compensation, such as participation in bonus and equity award programs.\nEEO Statement\nSupermicro is an Equal Opportunity Employer and embraces diversity in our employee population. It is the policy of Supermicro to provide equal opportunity to all qualified applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, protected veteran status or special disabled veteran, marital status, pregnancy, genetic information, or any other legally protected status.","datePosted":"2026-07-15T11:50:40.977Z","dateModified":"2026-07-15T11:50:40.977Z","hiringOrganization":{"@type":"Organization","name":"Super Micro Computer","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"San Jose","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"dfe319f765ba5a8f45a0ca49"},"url":"https://jobsearcher.com/jobs/dfe319f765ba5a8f45a0ca49"}}