{"schemaVersion":"jobsearcher.job.v1","id":"3bd1024f9d52f916172dbb62","url":"https://jobsearcher.com/jobs/3bd1024f9d52f916172dbb62","canonicalUrl":"https://jobsearcher.com/jobs/3bd1024f9d52f916172dbb62","title":"Data Center Operations Engineer","description":"Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.\n\nHeadquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.\n\nTo learn more, visit https://ir.bitdeer.com/\n\nKey Responsibilities\n\nResponsible for the daily operation and maintenance of the Data Center infrastructure to ensure high availability and stable service operation.\nPerform installation, rack and stack, cabling, commissioning, maintenance, and troubleshooting of AI/HPC cluster infrastructure, including:\nNVIDIA B300 Cluster\nGPU Servers\nx86 Servers\nStorage Servers\nEthernet and InfiniBand Switches\nDAC, AOC, Optical Fiber, and related cabling infrastructure\nMonitor and maintain the health status of cluster systems, including servers, GPUs, storage, networking devices, and associated infrastructure.\nConduct hardware replacement and maintenance activities, including FRU replacement, BIOS/BMC/Firmware upgrades, and hardware diagnostics.\nSupport server provisioning, operating system installation, cluster expansion, network validation, and burn-in testing.\nTroubleshoot hardware and infrastructure issues, including server failures, GPU errors, storage issues, network connectivity problems, switch failures, and cabling faults.\nPerform routine inspections, preventive maintenance, and maintain accurate operational records and maintenance logs.\nExecute incident response procedures and provide timely escalation and resolution according to operational standards.\nPrepare shift handover reports and maintain operation documents, SOPs, and incident reports.\nWork closely with engineering, network, and infrastructure teams to support new deployments and ongoing operation improvements.\nParticipate in a three-shift rotation schedule, including night shifts, weekends, and holidays as required.\n\nQualifications\n\nEducation\n\nBachelor's degree or above in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or related disciplines.\n\nTechnical Skills\n\nBasic understanding of Data Center infrastructure and server hardware architecture.\nFamiliarity with one or more of the following systems:\nNVIDIA B300 Cluster\nGPU Servers\nx86 Servers\nStorage Servers\nEthernet and InfiniBand Networks\nKnowledge of server hardware components, including CPU, memory, storage, GPU, BMC/IPMI, and firmware management.\nFamiliarity with network concepts, including TCP/IP, Ethernet, VLAN, Link Aggregation (LACP), and high-speed interconnect technologies such as InfiniBand or RoCE.\nUnderstanding of structured cabling systems, including DAC, AOC, optical fiber, MPO, and LC connectors.\n\nLinux Skills\n\nBasic Linux administration skills, including:\nSystem monitoring and troubleshooting\nService management using systemctl\nLog analysis using journalctl and dmesg\nNetwork troubleshooting tools such as ip and ethtool\nBasic shell scripting\n\nPreferred Qualifications\n\nExperience in Data Center operations or hardware maintenance is preferred.\nExperience supporting AI/HPC infrastructure or GPU clusters is a plus.\nFamiliarity with NVIDIA AI infrastructure, including GB200 and GB300 systems, is highly desirable.\nExperience with large-scale cluster environments and high-speed networking technologies is a plus.\nFamiliarity with monitoring and orchestration tools such as Slurm, Kubernetes, Prometheus, or Grafana is an advantage.\n\nPersonal Attributes\n\nWillingness to work in a 24x7 shift rotation schedule, including night shifts.\nStrong sense of responsibility and ownership.\nGood teamwork and communication skills.\nAbility to work under pressure and respond effectively to operational incidents.\nDetail-oriented with strong adherence to operational procedures and safety standards.\nSelf-motivated with a proactive attitude toward learning and problem-solving.\n\nThis position is ideal for candidates who are interested in building and operating next-generation AI Data Center infrastructure supporting large-scale NVIDIA B300 clusters.\n\nEqual Opportunity Employer\nBitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.","company":"Bitdeer Technologies Group","rawCompany":"bitdeer technologies group","city":"Needham","state":"MA","isRemote":false,"isActive":false,"createdAt":"2026-09-17T10:50:44.587Z","occupations":[{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1231.00","title":"Computer Network Support Specialists","slug":"computer-network-support-specialists"}],"industries":[{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541513","title":"Computer Facilities Management Services","slug":"computer-facilities-management-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Data Center Operations Engineer","description":"Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.\n\nHeadquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.\n\nTo learn more, visit https://ir.bitdeer.com/\n\nKey Responsibilities\n\nResponsible for the daily operation and maintenance of the Data Center infrastructure to ensure high availability and stable service operation.\nPerform installation, rack and stack, cabling, commissioning, maintenance, and troubleshooting of AI/HPC cluster infrastructure, including:\nNVIDIA B300 Cluster\nGPU Servers\nx86 Servers\nStorage Servers\nEthernet and InfiniBand Switches\nDAC, AOC, Optical Fiber, and related cabling infrastructure\nMonitor and maintain the health status of cluster systems, including servers, GPUs, storage, networking devices, and associated infrastructure.\nConduct hardware replacement and maintenance activities, including FRU replacement, BIOS/BMC/Firmware upgrades, and hardware diagnostics.\nSupport server provisioning, operating system installation, cluster expansion, network validation, and burn-in testing.\nTroubleshoot hardware and infrastructure issues, including server failures, GPU errors, storage issues, network connectivity problems, switch failures, and cabling faults.\nPerform routine inspections, preventive maintenance, and maintain accurate operational records and maintenance logs.\nExecute incident response procedures and provide timely escalation and resolution according to operational standards.\nPrepare shift handover reports and maintain operation documents, SOPs, and incident reports.\nWork closely with engineering, network, and infrastructure teams to support new deployments and ongoing operation improvements.\nParticipate in a three-shift rotation schedule, including night shifts, weekends, and holidays as required.\n\nQualifications\n\nEducation\n\nBachelor's degree or above in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or related disciplines.\n\nTechnical Skills\n\nBasic understanding of Data Center infrastructure and server hardware architecture.\nFamiliarity with one or more of the following systems:\nNVIDIA B300 Cluster\nGPU Servers\nx86 Servers\nStorage Servers\nEthernet and InfiniBand Networks\nKnowledge of server hardware components, including CPU, memory, storage, GPU, BMC/IPMI, and firmware management.\nFamiliarity with network concepts, including TCP/IP, Ethernet, VLAN, Link Aggregation (LACP), and high-speed interconnect technologies such as InfiniBand or RoCE.\nUnderstanding of structured cabling systems, including DAC, AOC, optical fiber, MPO, and LC connectors.\n\nLinux Skills\n\nBasic Linux administration skills, including:\nSystem monitoring and troubleshooting\nService management using systemctl\nLog analysis using journalctl and dmesg\nNetwork troubleshooting tools such as ip and ethtool\nBasic shell scripting\n\nPreferred Qualifications\n\nExperience in Data Center operations or hardware maintenance is preferred.\nExperience supporting AI/HPC infrastructure or GPU clusters is a plus.\nFamiliarity with NVIDIA AI infrastructure, including GB200 and GB300 systems, is highly desirable.\nExperience with large-scale cluster environments and high-speed networking technologies is a plus.\nFamiliarity with monitoring and orchestration tools such as Slurm, Kubernetes, Prometheus, or Grafana is an advantage.\n\nPersonal Attributes\n\nWillingness to work in a 24x7 shift rotation schedule, including night shifts.\nStrong sense of responsibility and ownership.\nGood teamwork and communication skills.\nAbility to work under pressure and respond effectively to operational incidents.\nDetail-oriented with strong adherence to operational procedures and safety standards.\nSelf-motivated with a proactive attitude toward learning and problem-solving.\n\nThis position is ideal for candidates who are interested in building and operating next-generation AI Data Center infrastructure supporting large-scale NVIDIA B300 clusters.\n\nEqual Opportunity Employer\nBitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.","datePosted":"2026-09-17T10:50:44.587Z","dateModified":"2026-09-17T10:50:44.587Z","hiringOrganization":{"@type":"Organization","name":"Bitdeer Technologies Group","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Needham","addressRegion":"MA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"3bd1024f9d52f916172dbb62"},"url":"https://jobsearcher.com/jobs/3bd1024f9d52f916172dbb62"}}