{"schemaVersion":"jobsearcher.job.v1","id":"49333c6d1fe9ff42d2f30b0e","url":"https://jobsearcher.com/jobs/49333c6d1fe9ff42d2f30b0e","canonicalUrl":"https://jobsearcher.com/jobs/49333c6d1fe9ff42d2f30b0e","title":"System Software Engineer, Node & Cluster Management","description":"Overview\nIn this role you will help build the node-level management and cluster software for MatX AI systems, bridging on-node daemons, kernel drivers, and BMC firmware into a unified management plane. You will work across the stack from low-level drivers to HTTP/REST interfaces, shaping scalable, observable hardware platforms. The work enables reliable, high-performance GenAI deployment at scale and across clusters. You’ll collaborate with firmware and hardware teams to deliver cohesive, maintainable management capabilities.\n\nCompensation / Benefits4 weeks PTO12 company holidaysremote work up to 3 weekshealth insurance for employees and dependents401K with company contributionProfessional development budget ($1500/year)\" ,\"team meals\",\"commute reimbursement\",\"AI resources up to $20K/month\"],\nResponsibilitiesDesign and implement the node-level management plane with health, inventory, telemetry, and control exposed via APIsDevelop cluster management and failover logic to minimize downtimeCreate management CLI tools to query state, run diagnostics, update firmware, and recover devicesCoordinate with BMC firmware engineers for unified in-band/out-of-band management viewsExtend node capabilities to fleet-level health, inventory, and integration points for customer fleet-managementPrototype and debug across the low-level stack (telemetry, drivers, device access utilities)Build automation for lab bring-up, provisioning, testing, and regression monitoringDefine software contracts between host daemons, BMC stack, and management layerDebug issues spanning APIs, daemons, kernel drivers, firmware, and hardwareShape scalable management for single node to rack-scale deployments\nKey requirements8+ years in systems softwareStrong Linux systems development experience including low-level userspace and kernel driver/daemon debuggingProficiency in C and at least one systems language (Go, Rust, C++, and/or Python)Experience designing and building HTTP/REST APIs and CLI tools for infrastructure managementSolid understanding of device drivers, telemetry paths, PCIe, and BMC subsystemsProven ability to debug across API, daemon, kernel, firmware, and hardware boundariesExperience aligning host-side and BMC-side capabilities with common interfacesSelf-driven, pragmatic, able to ship against new hardware with minimal specSelf-drivenPragmatic problem-solverCollaborative in cross-team environmentsC and systems programmingGo, Rust, C++, or Python for services/toolsLinux kernel drivers and userspace daemons","company":"Matx","rawCompany":"matx","city":"San Jose","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-09-15T04:24:51.990Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"System Software Engineer, Node & Cluster Management","description":"Overview\nIn this role you will help build the node-level management and cluster software for MatX AI systems, bridging on-node daemons, kernel drivers, and BMC firmware into a unified management plane. You will work across the stack from low-level drivers to HTTP/REST interfaces, shaping scalable, observable hardware platforms. The work enables reliable, high-performance GenAI deployment at scale and across clusters. You’ll collaborate with firmware and hardware teams to deliver cohesive, maintainable management capabilities.\n\nCompensation / Benefits4 weeks PTO12 company holidaysremote work up to 3 weekshealth insurance for employees and dependents401K with company contributionProfessional development budget ($1500/year)\" ,\"team meals\",\"commute reimbursement\",\"AI resources up to $20K/month\"],\nResponsibilitiesDesign and implement the node-level management plane with health, inventory, telemetry, and control exposed via APIsDevelop cluster management and failover logic to minimize downtimeCreate management CLI tools to query state, run diagnostics, update firmware, and recover devicesCoordinate with BMC firmware engineers for unified in-band/out-of-band management viewsExtend node capabilities to fleet-level health, inventory, and integration points for customer fleet-managementPrototype and debug across the low-level stack (telemetry, drivers, device access utilities)Build automation for lab bring-up, provisioning, testing, and regression monitoringDefine software contracts between host daemons, BMC stack, and management layerDebug issues spanning APIs, daemons, kernel drivers, firmware, and hardwareShape scalable management for single node to rack-scale deployments\nKey requirements8+ years in systems softwareStrong Linux systems development experience including low-level userspace and kernel driver/daemon debuggingProficiency in C and at least one systems language (Go, Rust, C++, and/or Python)Experience designing and building HTTP/REST APIs and CLI tools for infrastructure managementSolid understanding of device drivers, telemetry paths, PCIe, and BMC subsystemsProven ability to debug across API, daemon, kernel, firmware, and hardware boundariesExperience aligning host-side and BMC-side capabilities with common interfacesSelf-driven, pragmatic, able to ship against new hardware with minimal specSelf-drivenPragmatic problem-solverCollaborative in cross-team environmentsC and systems programmingGo, Rust, C++, or Python for services/toolsLinux kernel drivers and userspace daemons","datePosted":"2026-09-15T04:24:51.990Z","dateModified":"2026-09-15T04:24:51.990Z","hiringOrganization":{"@type":"Organization","name":"Matx","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"San Jose","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"49333c6d1fe9ff42d2f30b0e"},"url":"https://jobsearcher.com/jobs/49333c6d1fe9ff42d2f30b0e"}}