{"schemaVersion":"jobsearcher.job.v1","id":"bedb739d0dff1cc3ee8776f5","url":"https://jobsearcher.com/jobs/bedb739d0dff1cc3ee8776f5","canonicalUrl":"https://jobsearcher.com/jobs/bedb739d0dff1cc3ee8776f5","title":"Platform Engineer","description":"Job Title : Platform Engineer (L1 Support)Location : Sunnyvale, CA & San Jose CA (Final round F2F) -Need local profiles only.Type : Full Time/ContractInterview: Final interview- face-to-faceRequired Qualifications5 to 10+ years of experience in Site Reliability Engineering, Platform Engineering,Infrastructure Operations, or Systems Engineering.Strong infrastructure troubleshooting experience.Deep expertise in at least ertise in at least one of the following:GPU infrastructureKVM/virtualizationSDN (OVN/OVS)Storage (Lightbits/Pure Storage)Preferred QualificationsExperience with incident.io or similar incident management platforms.Kubernetes production operations experience.Cloud-native infrastructure experience.Experience supporting large-scale AI or GPU environments.Strong communication and stakeholder management skills.ResponsibilitiesAct as first responder during infrastructure incidents.Lead incident bridges and coordinate cross-functional response efforts.Perform incident triage and identify impacted infrastructure domains.Gather evidence and telemetry to route incidents to the correct SME team.Drive incident communications and stakeholder updates.Improve reliability processes across platform engineering teams.Define and promote SRE best practices and operational standards. (SLO,SLI)Identify observability gaps and implement improvements.Build automation for incident response workflows.Manage and optimize incident management tooling (e.g., incident.io).Support change management and operational readiness processes.Assist foundation engineering teams in identifying reliability risks and trends.Participate in on-call activities and operational reviews.","company":"Centraprise","rawCompany":"centraprise","city":"Sunnyvale","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-05T09:44:10.390Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Platform Engineer","description":"Job Title : Platform Engineer (L1 Support)Location : Sunnyvale, CA & San Jose CA (Final round F2F) -Need local profiles only.Type : Full Time/ContractInterview: Final interview- face-to-faceRequired Qualifications5 to 10+ years of experience in Site Reliability Engineering, Platform Engineering,Infrastructure Operations, or Systems Engineering.Strong infrastructure troubleshooting experience.Deep expertise in at least ertise in at least one of the following:GPU infrastructureKVM/virtualizationSDN (OVN/OVS)Storage (Lightbits/Pure Storage)Preferred QualificationsExperience with incident.io or similar incident management platforms.Kubernetes production operations experience.Cloud-native infrastructure experience.Experience supporting large-scale AI or GPU environments.Strong communication and stakeholder management skills.ResponsibilitiesAct as first responder during infrastructure incidents.Lead incident bridges and coordinate cross-functional response efforts.Perform incident triage and identify impacted infrastructure domains.Gather evidence and telemetry to route incidents to the correct SME team.Drive incident communications and stakeholder updates.Improve reliability processes across platform engineering teams.Define and promote SRE best practices and operational standards. (SLO,SLI)Identify observability gaps and implement improvements.Build automation for incident response workflows.Manage and optimize incident management tooling (e.g., incident.io).Support change management and operational readiness processes.Assist foundation engineering teams in identifying reliability risks and trends.Participate in on-call activities and operational reviews.","datePosted":"2026-08-05T09:44:10.390Z","dateModified":"2026-08-05T09:44:10.390Z","hiringOrganization":{"@type":"Organization","name":"Centraprise","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Sunnyvale","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"bedb739d0dff1cc3ee8776f5"},"url":"https://jobsearcher.com/jobs/bedb739d0dff1cc3ee8776f5"}}