Senior ML Engineer
Overview
You will own and advance the recommendation engine at the core of StormForge, CloudBolt’s Kubernetes resource optimization product. In a hands-on, production-focused ML role, you design time-series models, ensure data quality and safety, and collaborate cross-functionally to scale accurate, reliable recommendations. The work blends applied machine learning with production engineering to help customers optimize cloud spend without compromising workloads. You’ll shape the technical direction and raise the bar on performance and safety.
Compensation / BenefitsMedical/Dental/Vision coverage401k with Company MatchHealth & Dependent Care FSAUnlimited PTO11 Company HolidaysTuition Reimbursement`,`Paid Parental Leave
ResponsibilitiesOwn the end-to-end recommendation engine: model selection, algorithm design, preprocessing, and safe guardrails for live production workloadsDesign, evaluate, and productionize time-series forecasting and statistical models for right-sizing Kubernetes workloads across CPU, memory, GPU, and JVM heapBuild and maintain data-quality layer to detect anomalies and filter telemetry before modelingDefine and improve metrics and validation for recommendation quality (regression tests, behavioral validation, production metrics)Investigate and resolve customer-reported recommendation quality issues across data, preprocessing, and model behaviorServe as ML authority: guide technical direction, tradeoffs between models and heuristics, communicate to engineers and leadershipWrite production-grade Python for models and pipelines; share responsibility for surrounding service components (queues, caches, observability)Prototype and validate new optimization capabilities (new resource types, algorithms) from research to feature-flagged rolloutStay current on time-series forecasting and resource optimization techniques and evaluate practical applicability
Key requirementsMaster's degree or higher in a quantitative field5+ years software engineering, with 3+ years operating ML or statistical systems in productionExpert-level Python with strong typing and testing, fluency in numpyHands-on experience with time-series analysis and forecasting; traditional statistical methods and anomaly detectionRigorous ML system testing: regression testing against baselines, behavioral validation, numerical reproducibilityWorking knowledge of Kubernetes: resource requests/limits, autoscaling, understanding failures under-provisioningExperience owning production services (queues, caches, observability, debugging from logs/metrics)Clear written and verbal communication to explain model behavior and tradeoffsstrong communicationcuriosity and rigorproblem-solving with a collaborative mindsettime-series forecastingProphet or similar librariesPython (production-grade)