{"schemaVersion":"jobsearcher.job.v1","id":"dbfa29fcc161a1eace452622","url":"https://jobsearcher.com/jobs/dbfa29fcc161a1eace452622","canonicalUrl":"https://jobsearcher.com/jobs/dbfa29fcc161a1eace452622","title":"Senior Platform Engineer – AI/ML Infrastructure & Reliability","description":"Senior Platform Engineer – AI/ML Infrastructure & ReliabilityLocation: San Francisco, CA — SOMAWork Arrangement: Onsite five days per weekEmployment Type: Full-timeCompensation: $210,000–$260,000 base salary plus equity plus benefitsNote: No C2C arrangements will be considered.Any attempt to use personal or household contact information for solicitation, candidate submission, or vendor outreach is strictly prohibited and will be reported to LinkedIn.Why This Role Is DifferentThis is a Staff-caliber platform engineering role with deep Site Reliability Engineering ownership—not a ticket-taking, infrastructure-maintenance, or pipeline-only DevOps position.You will help define how mission-critical machine learning and real-time analytics systems are deployed, observed, scaled, and operated in production. Your work will influence reliability strategy, deployment standards, infrastructure architecture, incident response, and platform performance across engineering.We are looking for a hands-on builder who has developed from strong Linux, systems, networking, DevOps, or infrastructure foundations into a senior platform and reliability engineer capable of diagnosing systemic failures and improving how entire engineering teams operate.This role is intentionally onsite five days per week in San Francisco. Infrastructure, software, data, and machine-learning engineers work side by side on system design, production debugging, performance tuning, and incident reviews. Technical decisions happen quickly, feedback loops are tight, and this engineer will have meaningful access to engineering leadership.What You’ll OwnProduction reliability for machine-learning, real-time analytics, and data-intensive workloadsCI/CD architecture, deployment automation, release controls, and rollback designObservability frameworks, including SLOs, monitoring, alerting, dashboards, and incident-response processesKubernetes, containerized workloads, and infrastructure-as-code environmentsCapacity planning, scalability, platform performance, and operational readinessSystemic troubleshooting across Linux systems, networking, infrastructure, deployments, services, and application dependenciesPost-incident reviews and root-cause analysis that drive measurable, long-term improvementsReliability and operational standards across engineering—not just within a single serviceInternal tools and automation that reduce manual work and improve engineering velocityTechnical documentation, operational runbooks, and processes that make complex systems easier to operateYou will partner directly with software, infrastructure, data, machine-learning, and security teams to ensure production systems are reliable by design.What We’re Looking ForEight or more years of experience in Site Reliability Engineering, platform engineering, DevOps, infrastructure engineering, systems engineering, or production operationsDemonstrated impact as a Senior or Staff-level SRE, Platform Engineer, DevOps Engineer, Infrastructure Engineer, or Production EngineerDeep hands-on experience operating and troubleshooting Linux systems in productionStrong networking fundamentals and the ability to diagnose service, infrastructure, and connectivity issuesExperience supporting distributed, data-intensive, analytics, or machine-learning systems in productionStrong hands-on experience with Docker, Kubernetes, and container orchestrationInfrastructure-as-code experience using Terraform, Ansible, or similar technologiesExperience building, owning, or substantially improving CI/CD pipelines and deployment automationExperience designing and operating modern observability systems using tools such as Prometheus, Grafana, Datadog, ELK, OpenTelemetry, or similar platformsProduction incident-response experience, including on-call support, root-cause analysis, post-incident reviews, and long-term corrective actionStrong scripting and automation ability using Bash, Python, or similar languagesAbility to debug complex failures across infrastructure, deployments, networking, services, and workloadsStrong communication skills and the ability to work effectively across infrastructure, engineering, data, security, and business teamsA willingness and desire to author clear documentation, operational runbooks, and internal technical processesComfort working onsite five days per week in San FranciscoAbility to operate effectively in a fast-moving startup environment where priorities can shift and problems may not be fully definedThese Skills Are a PlusExperience operating machine-learning platforms at scale, including training or inference workloadsExperience supporting Spark, Airflow, Kafka, Elasticsearch, or other data-platform technologiesExperience with Azure, AWS, or cloud-managed infrastructure servicesExperience supporting GPU infrastructure, high-performance compute, or high-volume data-processing environmentsExperience with feature flags, progressive delivery, service mesh, release controls, or deployment-safety practicesExperience in SOC 2, regulated, security-sensitive, or high-compliance environmentsFamiliarity with cloud-native security, identity and access management, secrets management, and secure infrastructure designExperience supporting external customer requirements, service commitments, or client-facing production systemsExperience in lean startup environments where engineers take broad ownership and solve undefined problemsYou’ll Thrive Here IfYou enjoy solving problems that do not have predefined answersYou would rather build automation than repeat the same manual taskYou take ownership of systems from design through productionYou are comfortable moving between infrastructure, reliability, networking, deployments, and application-level problemsYou can identify what needs to be done without waiting for daily task directionYou are curious, low ego, and willing to go deep technicallyYou are energized by changing priorities and working across multiple technical areasYou care about improving the system, not simply resolving the immediate alertYou enjoy working directly with engineers across infrastructure, software, data, AI/ML, and productOn-Call ExpectationsThis role participates in a rotating on-call schedule, currently structured as one week every three weeks, covering a 7:00 a.m.–7:00 p.m. operational window.The engineer in this role will help improve alerting, automation, runbooks, system design, and incident processes so that on-call support becomes more effective and sustainable over time.Why JoinOwn reliability for mission-critical AI, machine-learning, and real-time analytics systemsInfluence platform architecture and reliability standards across engineeringWork in a highly visible role with direct access to engineering leadershipBuild foundational systems rather than simply maintain inherited infrastructureJoin a small, highly collaborative engineering environment where technical depth and thoughtful execution matterWork alongside experienced engineers who value curiosity, pragmatism, ownership, and continuous learningReceive competitive compensation and a meaningful equity opportunityWe are looking for builders—engineers who enjoy implementing solutions, automating repetitive work, improving systems, and stepping into whatever technical problem needs to be solved.Success in this role comes from curiosity, ownership, sound technical judgment, and a willingness to operate across traditional infrastructure boundaries.Work Authorization: Candidates must be authorized to work in the United States without current or future employer sponsorship.","company":"Stratitech","rawCompany":"stratitech","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-18T13:36:00.753Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior Platform Engineer – AI/ML Infrastructure & Reliability","description":"Senior Platform Engineer – AI/ML Infrastructure & ReliabilityLocation: San Francisco, CA — SOMAWork Arrangement: Onsite five days per weekEmployment Type: Full-timeCompensation: $210,000–$260,000 base salary plus equity plus benefitsNote: No C2C arrangements will be considered.Any attempt to use personal or household contact information for solicitation, candidate submission, or vendor outreach is strictly prohibited and will be reported to LinkedIn.Why This Role Is DifferentThis is a Staff-caliber platform engineering role with deep Site Reliability Engineering ownership—not a ticket-taking, infrastructure-maintenance, or pipeline-only DevOps position.You will help define how mission-critical machine learning and real-time analytics systems are deployed, observed, scaled, and operated in production. Your work will influence reliability strategy, deployment standards, infrastructure architecture, incident response, and platform performance across engineering.We are looking for a hands-on builder who has developed from strong Linux, systems, networking, DevOps, or infrastructure foundations into a senior platform and reliability engineer capable of diagnosing systemic failures and improving how entire engineering teams operate.This role is intentionally onsite five days per week in San Francisco. Infrastructure, software, data, and machine-learning engineers work side by side on system design, production debugging, performance tuning, and incident reviews. Technical decisions happen quickly, feedback loops are tight, and this engineer will have meaningful access to engineering leadership.What You’ll OwnProduction reliability for machine-learning, real-time analytics, and data-intensive workloadsCI/CD architecture, deployment automation, release controls, and rollback designObservability frameworks, including SLOs, monitoring, alerting, dashboards, and incident-response processesKubernetes, containerized workloads, and infrastructure-as-code environmentsCapacity planning, scalability, platform performance, and operational readinessSystemic troubleshooting across Linux systems, networking, infrastructure, deployments, services, and application dependenciesPost-incident reviews and root-cause analysis that drive measurable, long-term improvementsReliability and operational standards across engineering—not just within a single serviceInternal tools and automation that reduce manual work and improve engineering velocityTechnical documentation, operational runbooks, and processes that make complex systems easier to operateYou will partner directly with software, infrastructure, data, machine-learning, and security teams to ensure production systems are reliable by design.What We’re Looking ForEight or more years of experience in Site Reliability Engineering, platform engineering, DevOps, infrastructure engineering, systems engineering, or production operationsDemonstrated impact as a Senior or Staff-level SRE, Platform Engineer, DevOps Engineer, Infrastructure Engineer, or Production EngineerDeep hands-on experience operating and troubleshooting Linux systems in productionStrong networking fundamentals and the ability to diagnose service, infrastructure, and connectivity issuesExperience supporting distributed, data-intensive, analytics, or machine-learning systems in productionStrong hands-on experience with Docker, Kubernetes, and container orchestrationInfrastructure-as-code experience using Terraform, Ansible, or similar technologiesExperience building, owning, or substantially improving CI/CD pipelines and deployment automationExperience designing and operating modern observability systems using tools such as Prometheus, Grafana, Datadog, ELK, OpenTelemetry, or similar platformsProduction incident-response experience, including on-call support, root-cause analysis, post-incident reviews, and long-term corrective actionStrong scripting and automation ability using Bash, Python, or similar languagesAbility to debug complex failures across infrastructure, deployments, networking, services, and workloadsStrong communication skills and the ability to work effectively across infrastructure, engineering, data, security, and business teamsA willingness and desire to author clear documentation, operational runbooks, and internal technical processesComfort working onsite five days per week in San FranciscoAbility to operate effectively in a fast-moving startup environment where priorities can shift and problems may not be fully definedThese Skills Are a PlusExperience operating machine-learning platforms at scale, including training or inference workloadsExperience supporting Spark, Airflow, Kafka, Elasticsearch, or other data-platform technologiesExperience with Azure, AWS, or cloud-managed infrastructure servicesExperience supporting GPU infrastructure, high-performance compute, or high-volume data-processing environmentsExperience with feature flags, progressive delivery, service mesh, release controls, or deployment-safety practicesExperience in SOC 2, regulated, security-sensitive, or high-compliance environmentsFamiliarity with cloud-native security, identity and access management, secrets management, and secure infrastructure designExperience supporting external customer requirements, service commitments, or client-facing production systemsExperience in lean startup environments where engineers take broad ownership and solve undefined problemsYou’ll Thrive Here IfYou enjoy solving problems that do not have predefined answersYou would rather build automation than repeat the same manual taskYou take ownership of systems from design through productionYou are comfortable moving between infrastructure, reliability, networking, deployments, and application-level problemsYou can identify what needs to be done without waiting for daily task directionYou are curious, low ego, and willing to go deep technicallyYou are energized by changing priorities and working across multiple technical areasYou care about improving the system, not simply resolving the immediate alertYou enjoy working directly with engineers across infrastructure, software, data, AI/ML, and productOn-Call ExpectationsThis role participates in a rotating on-call schedule, currently structured as one week every three weeks, covering a 7:00 a.m.–7:00 p.m. operational window.The engineer in this role will help improve alerting, automation, runbooks, system design, and incident processes so that on-call support becomes more effective and sustainable over time.Why JoinOwn reliability for mission-critical AI, machine-learning, and real-time analytics systemsInfluence platform architecture and reliability standards across engineeringWork in a highly visible role with direct access to engineering leadershipBuild foundational systems rather than simply maintain inherited infrastructureJoin a small, highly collaborative engineering environment where technical depth and thoughtful execution matterWork alongside experienced engineers who value curiosity, pragmatism, ownership, and continuous learningReceive competitive compensation and a meaningful equity opportunityWe are looking for builders—engineers who enjoy implementing solutions, automating repetitive work, improving systems, and stepping into whatever technical problem needs to be solved.Success in this role comes from curiosity, ownership, sound technical judgment, and a willingness to operate across traditional infrastructure boundaries.Work Authorization: Candidates must be authorized to work in the United States without current or future employer sponsorship.","datePosted":"2026-07-18T13:36:00.753Z","dateModified":"2026-07-18T13:36:00.753Z","hiringOrganization":{"@type":"Organization","name":"Stratitech","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"dbfa29fcc161a1eace452622"},"url":"https://jobsearcher.com/jobs/dbfa29fcc161a1eace452622"}}