Platform Engineer
Department: EngineeringLocation: San FranciscoDescriptionYou'll own the infrastructure platform that our AI models run on. This isn't a CI/CD-focused DevOps role. You'll work across GPU orchestration, multi-cloud Kubernetes, real-time networking, and observability. You'll be the person who knows why a model pod took 4 minutes to schedule, why cross-region latency spiked, or why a media relay is dropping packets.We run production today across multiple Kubernetes clusters, regions, and GPU types, and we're actively expanding to additional cloud providers. You'll lead that expansion and keep everything running.What You'll DoProvision and manage multi-region Kubernetes clusters across AWS and GPU cloud providers using infrastructure-as-code.Own the GitOps deployment lifecycle (Helm charts, Kustomize overlays, image automation, and continuous delivery.)Manage GPU node infrastructure: scheduling, model weight caching, image prefetching for fast cold starts, and GPU observability.Operate and improve our networking layer: ingress and gateway management, load balancing, media relay infrastructure, and cross-region connectivity.Build and maintain our observability stack: metrics, logs, traces, and profiling across all services and GPU workloads.Maintain infrastructure security: IAM, secret management, certificate automation, and encryption at rest.Own CI/CD pipelines for monorepo builds spanning Go services, Python model containers, and Helm chart releases.Partner with ML engineers on model serving: container optimization, health checks and startup tuning, media pipeline performance, and multi-GPU configuration.What We're Looking For You've operated Kubernetes in production at scale, not just deployed to it, but debugged node-level scheduling issues, tuned autoscalers, and managed cluster upgrades.Strong infrastructure-as-code experience (Terraform, Pulumi, or similar) across multiple environments and regions.You've worked with GPU workloads on Kubernetes: device plugins, node taints/tolerations, GPU-aware scheduling. You understand why bin-packing mattersfor expensive hardware.Experience with GitOps tooling (FluxCD, ArgoCD, or similar) and Helm chart authoring.Comfortable with Redis or similar in-memory data stores (replication, persistence, pub/sub or streaming patterns)Familiarity with modern observability stacks (Prometheus, Grafana, OpenTelemetry, or equivalent) and knowing when to reach for metrics vs. logs vs. traces.Solid networking fundamentals: load balancers, TLS, DNS, NAT. Real-time or low-latency networking experience is a strong plus.You've worked in a startup where you owned infrastructure end-to-end, not just one slice of it.Nice to HaveExperience with GPU cloud providers beyond AWS (Crusoe, CoreWeave, Lambda Labs, Nebius)Real-time media or streaming infrastructureGo or Python proficiencyFamiliarity with ML model serving (container image optimization, weight loading, GPU driver and runtime management)FinOps and GPU cost optimizationWhat We're Not Looking ForPure CI/CD pipeline engineers who haven't operated Kubernetes clusters directlyCandidates whose infrastructure experience is limited to managed PaaS (Heroku, Vercel, Railway)People who need a fully defined scope, this role requires figuring out what to build next, not just executing ticketsBenefitsCompetitive San Francisco salary and meaningful equityWe sponsor visas and support relocation to the USGenerous health, dental, and vision coverage