Datacenter Operations Manager
Job DescriptionData Center Site Lead – AI Infrastructure (Datacenter Operations Manager)\n\nLocation: Santa Clara, California\n\nWorking model: Full-time, on-site\n\nEmployment type: Permanent\n\nAbout the opportunity\nWe are supporting a fast-growing AI infrastructure company that designs, deploys, and operates large-scale GPU compute environments.\nThe company is expanding its data center operations in Santa Clara and is looking for a hands-on Data Center Site Lead to take ownership of the site’s day-to-day operation, technical reliability, and future growth.\nThis is not a purely managerial position. You will be expected to understand the site in detail, including its power, cooling, networking, server infrastructure, dependencies, capacity constraints, and operational risks. You will act as the senior technical presence on-site, lead other technicians and vendors, and take ownership when incidents or equipment failures occur.\nThe environment supports demanding AI and high-performance computing workloads where uptime, response speed, and disciplined execution are critical.\n\nKey responsibilities\n\nTake day-to-day operational ownership of the Santa Clara data center site.\nAct as the senior technical lead for on-site technicians, contractors, vendors, and remote-hands teams.\nInstall, configure, troubleshoot, and maintain GPU servers, storage systems, networking equipment, cabling, and supporting infrastructure.\nMonitor site conditions, including power, cooling, temperature, humidity, capacity, alarms, and equipment health.\nEnsure the availability and reliability of the site within a 24/7 operational environment.\nLead the response to hardware failures, environmental alarms, connectivity issues, and other critical incidents.\nOwn incident reporting from initial detection through root-cause analysis, corrective action, and final closure.\nTrack equipment downtime, identify recurring failure patterns, and introduce preventive measures.\nCoordinate escalations with hardware manufacturers, colocation providers, network teams, and other technical partners.\nPlan and schedule server installations, rack deployments, maintenance activities, upgrades, and hardware refreshes.\nSupport data hall expansions, cluster deployments, migrations, and new capacity coming online.\nMaintain accurate records covering site assets, installations, incidents, maintenance work, capacity, and operational risks.\nEnsure all work follows the company’s safety, security, access-control, and change-management procedures.\nHelp develop site operating procedures, escalation processes, maintenance schedules, and reliability standards.\nMentor junior technicians and support the recruitment and development of the on-site team as the facility grows.\nProvide regular updates to leadership on uptime, incidents, staffing, capacity, operational risks, and planned work.\nParticipate in an on-call rotation and provide escalation support for critical incidents outside normal working hours.\n\n\nWhat we are looking for\n\nAt least five years of experience within data center operations, critical infrastructure, cloud infrastructure, or another mission-critical technical environment.\nStrong hands-on experience installing and supporting enterprise servers, storage, networking hardware, and structured cabling.\nPrevious experience acting as a site lead, senior technician, shift lead, operations manager, or technical escalation point.\nPractical understanding of data center power, cooling, networking, rack layouts, environmental monitoring, and common infrastructure failure modes.\nExperience managing incidents in a structured manner, including escalation, root-cause analysis, documentation, and preventive action.\nAbility to work independently and make sound operational decisions without requiring constant supervision.\nExperience coordinating technicians, contractors, vendors, and remote engineering teams.\nStrong planning and organisational skills, particularly around installations, maintenance windows, upgrades, and capacity expansion.\nClear written and verbal communication skills.\nWillingness to work on-site full-time in Santa Clara and participate in on-call coverage.\n\n\nParticularly relevant experience\n\nSupporting enterprise GPU platforms using NVIDIA Ampere, Hopper, Blackwell, GB200, GB300, or similar systems.\nOperating high-density AI, HPC, hyperscale, or cloud infrastructure.\nDirect liquid cooling, coolant distribution units, liquid loops, or other advanced cooling technologies.\nLarge GPU cluster deployments, server bring-up, burn-in, firmware updates, and hardware validation.\nData center build-outs, new data hall openings, migrations, expansions, or infrastructure refresh programmes.\nDCIM, CMMS, monitoring, alerting, ticketing, and maintenance-management platforms.\nManaging 24/7 shift coverage or supporting teams operating across multiple shifts.\nWorking within environments governed by strict SLAs, security controls, and safety procedures.\nData center qualifications such as CDCP, CDCS, or an equivalent certification.\n\n