{"schemaVersion":"jobsearcher.job.v1","id":"6034754b00a029aa950db295","url":"https://jobsearcher.com/jobs/6034754b00a029aa950db295","canonicalUrl":"https://jobsearcher.com/jobs/6034754b00a029aa950db295","title":"AI/ML Infrastructure Engineer","description":"Job DescriptionAI/ML Infrastructure EngineerMid-Level to Senior | Engineering & Platform OpsUS-based — CA preferred, open to West Coast + remote. Hybrid-friendly.Address: 2211 Park Blvd, Palo Alto, CA 94306━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ABOUT PREDII━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━Predii builds the intelligence layer that runs the automotive service and parts industry. Our platform, Predii 360, turns messy repair-order, DMS, and parts data into real-time intelligence — powering parts lookup, diagnostics, and repair search for dealership and aftermarket customers at scale, processing billions of repair orders and serving live search at sub-second latency.We're small, fast, and allergic to red tape. No 12-layer approval chains, no work that disappears into a backlog forever. If you build something here, it ships — and real customers use it. Learn more at www.predii.com.PREDII RESEARCH━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━We do real research, not just integration. We continue to submit state-of-the-art research on topics including: engineering-diagram and technical-document understanding, domain-calibrated evaluation frameworks for technical content, multi-agent architectures that optimize for correctness and honesty, detecting \"confident-but-wrong\" failures that standard monitoring misses, moving from reactive detection to causal, explainable prognosis, and multilingual evaluation of technical and repair content.We've found that multi-agent systems that just concatenate outputs get less trustworthy as they get more capable, so we design ours to contest and qualify each other's findings instead. And we run open-weight models in production at enterprise scale, because repair-grade accuracy shouldn't cost frontier-model money. All of it is deliberately vertical: deep automotive domain expertise applied to automotive problems, not a general-purpose model with an automotive skin.THE VIBE━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━We need a AI/ML Infrastructure Engineer who wants more than tickets — someone ready to actually own infrastructure across multiple clouds and help shape how we build. This is real ownership, not busywork. You'll touch:Multi-cloud infra (Azure, GCP, AWS)Kubernetes, CI/CD, automation-everythingSecurity, compliance, access — keeping the house lockedMonitoring & reliability — catching problems before customers doIncident response & disaster recoveryCloud cost optimization (yes, we care about the bill too)Senior folks: expect to shape architecture and mentor the team, not just execute someone else's roadmap.WHAT YOU'LL ACTUALLY DO━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━Cloud & Platform — build + run deployments on Azure, GCP, AWS with Kubernetes/AKS, Terraform, Ansible, Helm.DevOps & CI/CD — ship pipelines that are reliable and repeatable, not held together with duct tape.Reliability & Observability — build monitoring, logging, alerting, tracing; hunt down root causes, not just symptoms.Security & Compliance — RBAC, auth, network security, vuln management, SOC 2 Type 2 controls.Resilience & Ops — own backups, DR, capacity planning, cloud costs, and production support.Keep Leveling Up — evaluate new tools across DevOps, infra, and DevSecOps; you're not stuck with 2019's stack.IT Support — jump in on Windows Server / macOS support when needed.YOUR TOOLKIT━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━Cloud & Containers: Azure, GCP, AWS, Docker, Kubernetes, AKSAutomation: Terraform, Ansible, Helm, Bash, PythonCI/CD: GitHub Actions, GitLab CI, Azure DevOps (or similar)Observability: Grafana, Prometheus, ELK (or equivalent)Networking: TCP/IP, DNS, load balancing, VPNs, firewalls/NSGs, segmentationIdentity: RBAC, cloud IAM, Okta, multi-tenant authSystems: Linux, Windows Server, macOSSecurity & Compliance: SOC 2 Type 2, access/change controls, business continuity + DRWHAT YOU BRING━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━3–5+ years hands-on AI/ML Infrastructure / DevOps / SRE / Platform Engineering, in production — not just labs.Strong cloud + Kubernetes chops — Azure/AKS preferred; GCP/AWS/Docker is a big plus.Solid networking and security fundamentals.Comfortable with Terraform, Ansible, Helm, Bash, Python (or similar).Real Git-based CI/CD experience — automated deploys, security scanning included.Battle-tested on observability & prod ops — monitoring, incident response, RCA, runbooks, backup, DR.Working knowledge of security/compliance across Linux, Windows Server, macOS.A self-starter mindset — comfortable working independently across a distributed US–India team.BONUS POINTS━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━Auth/IAM experience, especially multi-tenant setups.Background in automotive or data-heavy platforms.Been through a SOC 2 Type 2 audit before.Startup or small-team energy — you've worn more than one hat.Cloud, Kubernetes, or Terraform certs.Senior folks: mentoring or technical leadership experience.HOW WE ROLL━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━Own it — flag issues early, make the call, follow through.Be proactive — don't wait to be told; spot the risk, bring the fix.Share the load — security and reliability are everyone's job, not just yours.Make it count — your work ships to production and touches real customers, real fast.How To ApplyEmail: jobs@predii.com","company":"Predii","rawCompany":"predii","city":"Sonoma","state":"CA","isRemote":false,"isActive":true,"createdAt":"2026-09-01T12:25:57.845Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"AI/ML Infrastructure Engineer","description":"Job DescriptionAI/ML Infrastructure EngineerMid-Level to Senior | Engineering & Platform OpsUS-based — CA preferred, open to West Coast + remote. Hybrid-friendly.Address: 2211 Park Blvd, Palo Alto, CA 94306━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ABOUT PREDII━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━Predii builds the intelligence layer that runs the automotive service and parts industry. Our platform, Predii 360, turns messy repair-order, DMS, and parts data into real-time intelligence — powering parts lookup, diagnostics, and repair search for dealership and aftermarket customers at scale, processing billions of repair orders and serving live search at sub-second latency.We're small, fast, and allergic to red tape. No 12-layer approval chains, no work that disappears into a backlog forever. If you build something here, it ships — and real customers use it. Learn more at www.predii.com.PREDII RESEARCH━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━We do real research, not just integration. We continue to submit state-of-the-art research on topics including: engineering-diagram and technical-document understanding, domain-calibrated evaluation frameworks for technical content, multi-agent architectures that optimize for correctness and honesty, detecting \"confident-but-wrong\" failures that standard monitoring misses, moving from reactive detection to causal, explainable prognosis, and multilingual evaluation of technical and repair content.We've found that multi-agent systems that just concatenate outputs get less trustworthy as they get more capable, so we design ours to contest and qualify each other's findings instead. And we run open-weight models in production at enterprise scale, because repair-grade accuracy shouldn't cost frontier-model money. All of it is deliberately vertical: deep automotive domain expertise applied to automotive problems, not a general-purpose model with an automotive skin.THE VIBE━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━We need a AI/ML Infrastructure Engineer who wants more than tickets — someone ready to actually own infrastructure across multiple clouds and help shape how we build. This is real ownership, not busywork. You'll touch:Multi-cloud infra (Azure, GCP, AWS)Kubernetes, CI/CD, automation-everythingSecurity, compliance, access — keeping the house lockedMonitoring & reliability — catching problems before customers doIncident response & disaster recoveryCloud cost optimization (yes, we care about the bill too)Senior folks: expect to shape architecture and mentor the team, not just execute someone else's roadmap.WHAT YOU'LL ACTUALLY DO━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━Cloud & Platform — build + run deployments on Azure, GCP, AWS with Kubernetes/AKS, Terraform, Ansible, Helm.DevOps & CI/CD — ship pipelines that are reliable and repeatable, not held together with duct tape.Reliability & Observability — build monitoring, logging, alerting, tracing; hunt down root causes, not just symptoms.Security & Compliance — RBAC, auth, network security, vuln management, SOC 2 Type 2 controls.Resilience & Ops — own backups, DR, capacity planning, cloud costs, and production support.Keep Leveling Up — evaluate new tools across DevOps, infra, and DevSecOps; you're not stuck with 2019's stack.IT Support — jump in on Windows Server / macOS support when needed.YOUR TOOLKIT━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━Cloud & Containers: Azure, GCP, AWS, Docker, Kubernetes, AKSAutomation: Terraform, Ansible, Helm, Bash, PythonCI/CD: GitHub Actions, GitLab CI, Azure DevOps (or similar)Observability: Grafana, Prometheus, ELK (or equivalent)Networking: TCP/IP, DNS, load balancing, VPNs, firewalls/NSGs, segmentationIdentity: RBAC, cloud IAM, Okta, multi-tenant authSystems: Linux, Windows Server, macOSSecurity & Compliance: SOC 2 Type 2, access/change controls, business continuity + DRWHAT YOU BRING━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━3–5+ years hands-on AI/ML Infrastructure / DevOps / SRE / Platform Engineering, in production — not just labs.Strong cloud + Kubernetes chops — Azure/AKS preferred; GCP/AWS/Docker is a big plus.Solid networking and security fundamentals.Comfortable with Terraform, Ansible, Helm, Bash, Python (or similar).Real Git-based CI/CD experience — automated deploys, security scanning included.Battle-tested on observability & prod ops — monitoring, incident response, RCA, runbooks, backup, DR.Working knowledge of security/compliance across Linux, Windows Server, macOS.A self-starter mindset — comfortable working independently across a distributed US–India team.BONUS POINTS━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━Auth/IAM experience, especially multi-tenant setups.Background in automotive or data-heavy platforms.Been through a SOC 2 Type 2 audit before.Startup or small-team energy — you've worn more than one hat.Cloud, Kubernetes, or Terraform certs.Senior folks: mentoring or technical leadership experience.HOW WE ROLL━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━Own it — flag issues early, make the call, follow through.Be proactive — don't wait to be told; spot the risk, bring the fix.Share the load — security and reliability are everyone's job, not just yours.Make it count — your work ships to production and touches real customers, real fast.How To ApplyEmail: jobs@predii.com","datePosted":"2026-09-01T12:25:57.845Z","dateModified":"2026-09-01T12:25:57.845Z","hiringOrganization":{"@type":"Organization","name":"Predii","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Sonoma","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"6034754b00a029aa950db295"},"url":"https://jobsearcher.com/jobs/6034754b00a029aa950db295"}}