JOBSEARCHER

Machine Learning Engineer 5 - Decisioning & Optimization

NetflixBrooklyn, NYL7 ManagerSeptember 15th, 2026
Overview In this role you will build and operate end-to-end real-time ML model serving infrastructure for ad decisioning at massive scale. You will enable dozens of concurrent models per ad request with sub-20ms latency and lead feature serving, model validation, and deployment workflows. You’ll drive production-grade monitoring, experimentation, and reliability to optimize the ad marketplace while collaborating with Data Science and Platform teams. This is a chance to shape the performance and efficiency of Netflix’s in-house ad tech ecosystem and impact global ad delivery. Compensation / BenefitsHealth Plans401(k) with employer matchStock Option ProgramHealth Savings Account or Flexible Spending AccountsDisability ProgramsFamily-forming benefits ResponsibilitiesBuild and operate end-to-end ML model serving infrastructure for real-time ad decisioning, including publishing, packaging, validation, and zero-downtime deploymentScale inference path to support dozens of concurrent models at 1M+ QPS with strict latency budgets, including batching, resource allocation, and model versioningDesign and optimize feature serving paths with sub-10ms P99 fetch latency and online/offline consistencyProductionize scoring and ranking models for multi-stage ad selection and integrate outputs into auctionsDevelop model performance monitoring in production (latency, drift, calibration, regression detection)Collaborate with Data Science & Platform teams to align on workflows and standardsBuild simulation infrastructure to replay production traffic for offline validation of marketplace changesDrive operational excellence for ML systems (reliability, observability, capacity planning, incident response)Support live-event scalability for 35M+ concurrent viewers Key requirements7+ years of software engineering experience3+ years focused on ML infrastructure, model serving, or ML platform work in ads or real-time decisioningBuilt and operated real-time model serving systems at high QPS with sub-20ms latencyProficiency in Java, Python, or Scala with knowledge of multi-threading and performance optimizationHands-on with ML serving frameworks (serialization, runtime optimization, deployment constraints)Experience with feature engineering pipelines for real-time systems (online/offline consistency, hydration, caching)Strong model monitoring skills in production (drift detection, distribution analysis, calibration, latency profiling)Able to translate model artifacts into production-ready services that meet SLAExperience navigating both big-tech scale and startup speedExperience with JVM ecosystemComfortable at the boundary between ML research and production engineeringcollaboration with cross-functional teamsownership and accountabilityproblem-solving under high scaleML model serving / inference at low latencyfeature stores and model registriesmodel canary and shadow rollout