JOBSEARCHER

System Software Engineer

System Software Engineer — Node & Cluster ManagementBay Area, CAWe’re working with an early-stage AI hardware company building a new generation of high-performance accelerator systems for large-scale AI workloads.They’re looking for a System Software Engineer to help build the management and observability layer that makes new AI hardware usable at scale — from individual nodes through racks and clusters.This is a hands-on systems role for someone who can work across infrastructure APIs, Linux services, hardware telemetry, BMC interfaces, and lower-level device software. You’ll be building the tooling operators rely on while also being comfortable dropping into drivers, firmware interfaces, and raw device access when something breaks.What You’ll Work OnDesign and build the node-level management plane for AI accelerator systemsExpose system health, inventory, telemetry, diagnostics, and control through HTTP/REST APIsBuild cluster management, failover, recovery, and availability mechanismsDevelop CLI tools used for diagnostics, firmware updates, device recovery, and system managementCreate unified management across host-side software and BMC/out-of-band interfacesBuild fleet-wide health aggregation, inventory, alerting, and external management integrationsDevelop management and telemetry daemons running on Linux hostsWork with lower-level driver interfaces and direct hardware-access utilities when debugging or prototypingBuild tooling for provisioning, lab automation, test orchestration, and regression monitoringDefine interfaces between Linux services, BMC firmware, device software, and management layersDebug issues spanning APIs, userspace daemons, kernel drivers, firmware, and hardwareHelp define how a brand-new AI platform is managed from a single server through full-scale clustersWhat We’re Looking For8+ years of experience in systems software, infrastructure software, platform software, or hardware managementStrong Linux systems programming experienceExperience building low-level userspace software and Linux daemonsStrong C skills plus experience with Go, Rust, C++, and/or PythonExperience building REST APIs and CLI tooling for hardware or infrastructure systemsSolid understanding of Linux systems and the hardware underneath them, including:device driversPCIe deviceshardware telemetryfirmware interfacesBMC-managed subsystemsComfortable reading and debugging kernel and systems-level codeExperience troubleshooting issues across application, daemon, kernel, firmware, and hardware boundariesAble to work closely with firmware and hardware engineers to create common management interfacesComfortable operating in an early-stage environment where specifications and platform capabilities are still evolvingParticularly Relevant Experience - Experience in one or more of the following would be especially valuable:Redfish, OpenBMC, IPMI, or gNMIGPU, accelerator, HPC, or datacenter fleet managementNode and cluster managementHardware health monitoring and telemetryFirmware update and recovery workflowsHardware bring-up and lab infrastructureProvisioning and test automationSecure boot or device attestationBMC and out-of-band managementLarge-scale Linux infrastructureKeywords: Systems Software, Linux, Node Management, Cluster Management, Fleet Management, Infrastructure Software, Platform Software, AI Infrastructure, GPU Infrastructure, Accelerator Infrastructure, Datacenter Systems, Redfish, OpenBMC, BMC, IPMI, gNMI, REST API, Linux Daemons, Telemetry, Observability, Hardware Management, Firmware Management, PCIe, Hardware Bring-Up, Lab Automation, Provisioning, Failover, Device Management, C, C++, Rust, Go, Python