{"schemaVersion":"jobsearcher.job.v1","id":"c028f313f1f21b480dcdcd3f","url":"https://jobsearcher.com/jobs/c028f313f1f21b480dcdcd3f","canonicalUrl":"https://jobsearcher.com/jobs/c028f313f1f21b480dcdcd3f","title":"Platform Engineer","description":"Role 1 - Core Platform Engineer (L1, breadth-first)The engineering first line of defense: incident response, triage, reliability, and automation across the full infrastructure stack.Day-to-day:Runs incident response drills, post-mortems, and root cause analysis; learns from past incidents to prevent recurrence.Starts the day reviewing overnight alerts and system performance metrics, triaging anomalies.Participates in team stand-ups on projects, incidents, and daily priorities.Automates routine processes, analyzes system logs, and builds tools to strengthen monitoring.Works alongside software engineers advising on resilient-code best practices and reviewing changes pre-deployment.Maintains high SLIs/SLOs; documents work and shares insights with a customer-centric mindset.Must have:Architecture, design patterns, reliability, and scaling of new and existing systems.Incident command experience - driving RCA, coordinating cross-functional teams, ensuring corrective-action follow-through.Observability built from the ground up - defining SLOs/SLIs, closing monitoring gaps, alerting strategies that catch failures before customers do.Linux kernel internals - scheduler, memory allocation, driver subsystems.High-quality code in at least one language (Python, Go, or similar).System-level debugging - kdump, kernel panic analysis.IaC (Ansible, Terraform, Kubernetes) and CI/CD (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure.TCP/IP and network programming.Distributed storage systems - object, block, and/or file storage paradigms.Strong communication skills.Nice to have:Hardware and GPU troubleshooting.OVN/OVS-based networking stack exposure.Sourcing note: This is a deep SRE profile, not a pure generalist. The kernel-internals and system-level debugging bar is real and higher than a typical \"L1\" label implies - screen for genuine engineering depth, not helpdesk/NOC-tier breadth.Role 2 - Platform Engineer (L2, depth-first)Specialized domain expert embedded in a single foundation team: Storage, Compute, or SDN. Owns reliability, performance, scalability, and operational excellence for that one pillar. No cross-vertical work.Must have:Deep expertise in one domain (Storage / Compute / SDN).SRE as a fundamental minimum - has set up observability, done alert management, and worked with SLIs/SLOs within that domain.Ability to code the infrastructure: extend out-of-the-box product parameters to build observability layers, and tie SLAs/SLOs to business KPIs.Domain specifics:o Storage: block, blob, and file storage; distributed storage; performance diagnostics and data-path optimization.Critical screening filter - operator vs. SRE:Crusoe explicitly does not want another \"operator\" (storage admin doing patching, installs, upgrades). Screen hard for SRE substance (observability built, alerting, SLI/SLO ownership), not just domain tenure.Role 3 - Platform Engineer (L2, depth-first)Specialized domain expert embedded in a single foundation team: Storage, Compute, or SDN. Owns reliability, performance, scalability, and operational excellence for that one pillar. No cross-vertical work.Must have:Deep expertise in one domain (Storage / Compute / SDN).SRE as a fundamental minimum - has set up observability, done alert management, and worked with SLIs/SLOs within that domain.Ability to code the infrastructure: extend out-of-the-box product parameters to build observability layers, and tie SLAs/SLOs to business KPIs.Domain specifics:o Compute: Linux systems, KVM/QEMU, Cloud Hypervisor, kernel tuning, CPU/memory/VM optimization.Critical screening filter - operator vs. SRE:Crusoe explicitly does not want another \"operator\" (storage admin doing patching, installs, upgrades). Screen hard for SRE substance (observability built, alerting, SLI/SLO ownership), not just domain tenure.Role 4 - Platform Engineer (L2, depth-first)Specialized domain expert embedded in a single foundation team: Storage, Compute, or SDN. Owns reliability, performance, scalability, and operational excellence for that one pillar. No cross-vertical work.Must have:Deep expertise in one domain (Storage / Compute / SDN).SRE as a fundamental minimum - has set up observability, done alert management, and worked with SLIs/SLOs within that domain.Ability to code the infrastructure: extend out-of-the-box product parameters to build observability layers, and tie SLAs/SLOs to business KPIs.Domain specifics:SDN: OVS/OVN, network virtualization, high-performance networking, NIC tuning.Critical screening filter - operator vs. SRE:Crusoe explicitly does not want another \"operator\" (storage admin doing patching, installs, upgrades). Screen hard for SRE substance (observability built, alerting, SLI/SLO ownership), not just domain tenure.","company":"Gcb Services","rawCompany":"gcb services","city":"Sunnyvale","state":"CA","isRemote":false,"isActive":true,"createdAt":"2026-10-02T13:26:41.116Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Platform Engineer","description":"Role 1 - Core Platform Engineer (L1, breadth-first)The engineering first line of defense: incident response, triage, reliability, and automation across the full infrastructure stack.Day-to-day:Runs incident response drills, post-mortems, and root cause analysis; learns from past incidents to prevent recurrence.Starts the day reviewing overnight alerts and system performance metrics, triaging anomalies.Participates in team stand-ups on projects, incidents, and daily priorities.Automates routine processes, analyzes system logs, and builds tools to strengthen monitoring.Works alongside software engineers advising on resilient-code best practices and reviewing changes pre-deployment.Maintains high SLIs/SLOs; documents work and shares insights with a customer-centric mindset.Must have:Architecture, design patterns, reliability, and scaling of new and existing systems.Incident command experience - driving RCA, coordinating cross-functional teams, ensuring corrective-action follow-through.Observability built from the ground up - defining SLOs/SLIs, closing monitoring gaps, alerting strategies that catch failures before customers do.Linux kernel internals - scheduler, memory allocation, driver subsystems.High-quality code in at least one language (Python, Go, or similar).System-level debugging - kdump, kernel panic analysis.IaC (Ansible, Terraform, Kubernetes) and CI/CD (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure.TCP/IP and network programming.Distributed storage systems - object, block, and/or file storage paradigms.Strong communication skills.Nice to have:Hardware and GPU troubleshooting.OVN/OVS-based networking stack exposure.Sourcing note: This is a deep SRE profile, not a pure generalist. The kernel-internals and system-level debugging bar is real and higher than a typical \"L1\" label implies - screen for genuine engineering depth, not helpdesk/NOC-tier breadth.Role 2 - Platform Engineer (L2, depth-first)Specialized domain expert embedded in a single foundation team: Storage, Compute, or SDN. Owns reliability, performance, scalability, and operational excellence for that one pillar. No cross-vertical work.Must have:Deep expertise in one domain (Storage / Compute / SDN).SRE as a fundamental minimum - has set up observability, done alert management, and worked with SLIs/SLOs within that domain.Ability to code the infrastructure: extend out-of-the-box product parameters to build observability layers, and tie SLAs/SLOs to business KPIs.Domain specifics:o Storage: block, blob, and file storage; distributed storage; performance diagnostics and data-path optimization.Critical screening filter - operator vs. SRE:Crusoe explicitly does not want another \"operator\" (storage admin doing patching, installs, upgrades). Screen hard for SRE substance (observability built, alerting, SLI/SLO ownership), not just domain tenure.Role 3 - Platform Engineer (L2, depth-first)Specialized domain expert embedded in a single foundation team: Storage, Compute, or SDN. Owns reliability, performance, scalability, and operational excellence for that one pillar. No cross-vertical work.Must have:Deep expertise in one domain (Storage / Compute / SDN).SRE as a fundamental minimum - has set up observability, done alert management, and worked with SLIs/SLOs within that domain.Ability to code the infrastructure: extend out-of-the-box product parameters to build observability layers, and tie SLAs/SLOs to business KPIs.Domain specifics:o Compute: Linux systems, KVM/QEMU, Cloud Hypervisor, kernel tuning, CPU/memory/VM optimization.Critical screening filter - operator vs. SRE:Crusoe explicitly does not want another \"operator\" (storage admin doing patching, installs, upgrades). Screen hard for SRE substance (observability built, alerting, SLI/SLO ownership), not just domain tenure.Role 4 - Platform Engineer (L2, depth-first)Specialized domain expert embedded in a single foundation team: Storage, Compute, or SDN. Owns reliability, performance, scalability, and operational excellence for that one pillar. No cross-vertical work.Must have:Deep expertise in one domain (Storage / Compute / SDN).SRE as a fundamental minimum - has set up observability, done alert management, and worked with SLIs/SLOs within that domain.Ability to code the infrastructure: extend out-of-the-box product parameters to build observability layers, and tie SLAs/SLOs to business KPIs.Domain specifics:SDN: OVS/OVN, network virtualization, high-performance networking, NIC tuning.Critical screening filter - operator vs. SRE:Crusoe explicitly does not want another \"operator\" (storage admin doing patching, installs, upgrades). Screen hard for SRE substance (observability built, alerting, SLI/SLO ownership), not just domain tenure.","datePosted":"2026-10-02T13:26:41.116Z","dateModified":"2026-10-02T13:26:41.116Z","hiringOrganization":{"@type":"Organization","name":"Gcb Services","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Sunnyvale","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"c028f313f1f21b480dcdcd3f"},"url":"https://jobsearcher.com/jobs/c028f313f1f21b480dcdcd3f"}}