{"schemaVersion":"jobsearcher.job.v1","id":"c0a1c7f4a2f943bdd3ca7dd0","url":"https://jobsearcher.com/jobs/c0a1c7f4a2f943bdd3ca7dd0","canonicalUrl":"https://jobsearcher.com/jobs/c0a1c7f4a2f943bdd3ca7dd0","title":"4 Remote Nvidia Engineers","description":"NVIDIA AI Infrastructure & Kubernetes Platform Engineer (DGX Systems) RemoteNVIDIA Certification required or no interview6 months to 1+ yrs$openUSC or GC reqAlternate titles depending on context:AI Platform Architect – DGX & SuperPODAI Infrastructure DevOps Engineer – NVIDIA DGX StackSenior AI Systems Engineer – DGX | Kubernetes | InfiniBandJob Description:We are seeking a highly skilled AI Infrastructure & Kubernetes Platform Engineer with a proven track record in deploying and managing NVIDIA DGX-based AI clusters, orchestrating containerized AI workloads using Kubernetes, and ensuring secure, high-throughput operations across InfiniBand-powered networks. The ideal candidate will hold a combination of Kubernetes certifications (CKA, CKAD, CKS) and NVIDIA certifications (NCA-AIIO, NCP-AIO, NCP-AII, NCP-AIN), coupled with hands-on training in DGX, BlueField, and high-speed network operations.This position plays a key role in supporting AI/ML infrastructure at scale, enabling efficient training and inference for complex models, and integrating NVIDIA's cutting-edge compute, storage, and fabric solutions with modern DevOps practices.AI Infrastructure OperationsDeploy and manage NVIDIA DGX BasePODs and SuperPODs for high-performance AI workloads.Oversee DGX system lifecycle operations including provisioning, monitoring, firmware upgrades, and capacity planning.Operate Base Command Manager to manage GPU clusters, schedule workloads, and integrate with MLOps tools.Perform DGX node health validation, NCCL interconnect testing, and NVLink topology verification following new deployments or hardware changes.Kubernetes Platform EngineeringArchitect secure and scalable Kubernetes clusters optimized for GPU-accelerated workloads using NVIDIA GPU Operator.Leverage expertise from CKA/CKAD/CKS to develop, deploy, and secure AI applications on Kubernetes.Implement CI/CD pipelines and GitOps methodologies for deploying and managing ML workflows.High-Performance Networking & DPUsAdminister InfiniBand networks and BlueField DPUs using Unified Fabric Manager (UFM).Enable NVLink/NVSwitch performance across GPU nodes and tune fabric configurations for minimal latency and maximum throughput.Use BlueField for offloading storage, firewalling, and telemetry, enhancing AI workload security and performance.Security & ComplianceApply best practices from the CKS certification to secure containerized AI environments.Configure runtime security, secrets management, network segmentation, and auditing using DPU-enhanced Kubernetes deployments.Support zero-trust architecture initiatives by enforcing workload identity, RBAC policies, and supply chain integrity across AI container images and model artifacts.Monitor GPU, CPU, and I/O performance using NVIDIA DCGM, Prometheus, Grafana, and Base Command APIs.Tune system performance and model training pipelines for cost-efficiency and throughput.Build and maintain operational runbooks, incident response playbooks, and SLA reporting dashboards covering GPU utilization, thermal thresholds, and fabric health.Expertise With:DGX System, BasePOD, and SuperPOD AdministrationBlueField DPU Configuration & OperationsInfiniBand Fabric and UFM ManagementBase Command Manager for workload orchestrationThis is a remote position.J-18808-Ljbffr","company":"It Search","rawCompany":"it search","isRemote":true,"isActive":false,"createdAt":"2026-05-12T08:45:24.021Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541513","title":"Computer Facilities Management Services","slug":"computer-facilities-management-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"4 Remote Nvidia Engineers","description":"NVIDIA AI Infrastructure & Kubernetes Platform Engineer (DGX Systems) RemoteNVIDIA Certification required or no interview6 months to 1+ yrs$openUSC or GC reqAlternate titles depending on context:AI Platform Architect – DGX & SuperPODAI Infrastructure DevOps Engineer – NVIDIA DGX StackSenior AI Systems Engineer – DGX | Kubernetes | InfiniBandJob Description:We are seeking a highly skilled AI Infrastructure & Kubernetes Platform Engineer with a proven track record in deploying and managing NVIDIA DGX-based AI clusters, orchestrating containerized AI workloads using Kubernetes, and ensuring secure, high-throughput operations across InfiniBand-powered networks. The ideal candidate will hold a combination of Kubernetes certifications (CKA, CKAD, CKS) and NVIDIA certifications (NCA-AIIO, NCP-AIO, NCP-AII, NCP-AIN), coupled with hands-on training in DGX, BlueField, and high-speed network operations.This position plays a key role in supporting AI/ML infrastructure at scale, enabling efficient training and inference for complex models, and integrating NVIDIA's cutting-edge compute, storage, and fabric solutions with modern DevOps practices.AI Infrastructure OperationsDeploy and manage NVIDIA DGX BasePODs and SuperPODs for high-performance AI workloads.Oversee DGX system lifecycle operations including provisioning, monitoring, firmware upgrades, and capacity planning.Operate Base Command Manager to manage GPU clusters, schedule workloads, and integrate with MLOps tools.Perform DGX node health validation, NCCL interconnect testing, and NVLink topology verification following new deployments or hardware changes.Kubernetes Platform EngineeringArchitect secure and scalable Kubernetes clusters optimized for GPU-accelerated workloads using NVIDIA GPU Operator.Leverage expertise from CKA/CKAD/CKS to develop, deploy, and secure AI applications on Kubernetes.Implement CI/CD pipelines and GitOps methodologies for deploying and managing ML workflows.High-Performance Networking & DPUsAdminister InfiniBand networks and BlueField DPUs using Unified Fabric Manager (UFM).Enable NVLink/NVSwitch performance across GPU nodes and tune fabric configurations for minimal latency and maximum throughput.Use BlueField for offloading storage, firewalling, and telemetry, enhancing AI workload security and performance.Security & ComplianceApply best practices from the CKS certification to secure containerized AI environments.Configure runtime security, secrets management, network segmentation, and auditing using DPU-enhanced Kubernetes deployments.Support zero-trust architecture initiatives by enforcing workload identity, RBAC policies, and supply chain integrity across AI container images and model artifacts.Monitor GPU, CPU, and I/O performance using NVIDIA DCGM, Prometheus, Grafana, and Base Command APIs.Tune system performance and model training pipelines for cost-efficiency and throughput.Build and maintain operational runbooks, incident response playbooks, and SLA reporting dashboards covering GPU utilization, thermal thresholds, and fabric health.Expertise With:DGX System, BasePOD, and SuperPOD AdministrationBlueField DPU Configuration & OperationsInfiniBand Fabric and UFM ManagementBase Command Manager for workload orchestrationThis is a remote position.J-18808-Ljbffr","datePosted":"2026-05-12T08:45:24.021Z","dateModified":"2026-05-12T08:45:24.021Z","hiringOrganization":{"@type":"Organization","name":"It Search","sameAs":"https://jobsearcher.com"},"jobLocationType":"TELECOMMUTE","applicantLocationRequirements":{"@type":"Country","name":"US"},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"c0a1c7f4a2f943bdd3ca7dd0"},"url":"https://jobsearcher.com/jobs/c0a1c7f4a2f943bdd3ca7dd0"}}