Machine Learning Engineer
About ShadeformShadeform is a GPU cloud marketplace. We aggregate compute across 30+ cloud providers and data centers through a single API and dashboard, so teams can find, deploy, and manage GPUs wherever capacity is available and affordable.We're YC-backed (S23), profitable, and growing fast — a team of around 15 building one of the most important layers of AI infrastructure.About the RoleWe're hiring an ML Engineer to build the intelligent layer that decides how and where ML workloads run across our fleet. That fleet spans dozens of providers and a wide range of hardware, and no two nodes perform the same. Your job is to turn that messy, heterogeneous supply into something that reliably runs workloads well: an optimization and orchestration layer that matches jobs to the right hardware, tunes how they run, and gets the most out of every node.What You'll DoBuild the optimization and orchestration logic that places and tunes ML workloads across a heterogeneous, multi-provider fleet.Benchmark GPUs, interconnects, and driver stacks to understand real per-node performance, and feed that back into how the platform makes decisions.Build the tooling, images, and reference stacks that take someone from a fresh instance to a running job without a day of setup.Dig into performance and reliability problems that cross the line between the workload and the underlying hardware.Turn what you learn into things that scale: defaults, playbooks, and platform improvements.Take part in a weekly on-call rotation (24/7 coverage shared across the team).What You'll BringSolid experience getting ML workloads running in production, not just notebooks or research code.A real understanding of GPUs and how ML workloads use them: memory, throughput, and where things actually bottleneck.Comfort down at the systems and Linux level. You're not lost when a problem turns out to be drivers or networking rather than the model.The ability to own an ambiguous problem end to end and build a process where none exists yet.Nice to HaveDistributed training or large-scale inference in production.Familiarity with NVIDIA driver software, interconnects (InfiniBand/RoCE), or cluster networking.What We OfferCompetitive base salary and meaningful equity.A remote-first team working across the US and EU.A high degree of ownership over work that's central to the business.Comprehensive health benefits and flexible time off.Working at ShadeformWe're a small team, which means people have a high degree of ownership and autonomy. The work moves quickly, but we care a lot about building durable systems and maintaining quality. Most problems here are ambiguous at first, and people who do well tend to enjoy figuring things out without a lot of process already in place.We care deeply about clear thinking, strong execution, and working with people who raise the bar for the team around them.Shadeform is an equal opportunity employer.