Senior Machine Learning Engineer
About the role:
We're looking for a Machine Learning Engineer who can own the full lifecycle of alarge language model in production — from training and fine-tuning throughdeployment at scale and rigorous evaluation of what actually ships. This isn't aresearch-only role and it isn't a pure infrastructure role: you'll need to begenuinely fluent in all three, because the hardest problems in this space live atthe seams between them — a training decision that breaks serving latency, adeployment optimization that silently degrades output quality, an eval result thatdoesn't predict real-world behavior.
What you'll do:
Training & fine-tuning
- Design and execute pretraining, continued pretraining, and fine-tuning runs (SFT, DPO/RLHF-style alignment, LoRA/QLoRA and full-parameter approaches) against clear, measurable objectives
- Own data pipeline decisions that materially affect model quality — curation, deduplication, mixture weighting, and contamination checks against eval sets- Run and interpret distributed training (multi-GPU, multi-node) using frameworks such as FSDP, DeepSpeed, or Megatron-style parallelism, and diagnose failures that only show up at scale (loss spikes, stragglers, checkpoint corruption)
- Make and defend real tradeoffs between model size, training cost, and downstream performance
Deployment
Take a trained model to production: quantization, batching strategy, KV-cache management, and serving framework selection (e.g. vLLM, TensorRT-LLM, TGI) with explicit latency/throughput/cost targets, not just "make it run"
Design for the failure modes specific to LLM serving — tail latency under load, graceful degradation, prompt injection surface area, and safe fallback behavior
Build the operational muscle around this: monitoring, alerting, and rollback paths for a model in production, treated with the same rigor as any other critical service, not as a one-off notebook export
Assessment & evaluation
Build and maintain evaluation harnesses that go beyond running published benchmarks — including task-specific eval sets that reflect your actual product's use cases, not just leaderboard performance
Design human evaluation protocols where automated metrics fall short, and know which is which
Own regression detection: catching quality drops introduced by a new checkpoint, a prompt template change, or a serving optimization before they reach users- Contribute to safety and robustness evaluation — hallucination rate, adversarial/ red-team testing, and behavior under distribution shift — as a first-class part of the release process, not an afterthought
What we're looking for
4+ years of applied ML engineering experience, with at least **2 years working directly on large language models in a production context
Real production deployment experience — you've shipped a model that served live traffic, and you can talk about the latency/cost/quality tradeoffs you made- Strong software engineering fundamentals
Fluency with the modern LLM tooling landscape (training frameworks, serving frameworks, eval tooling)
Comfort with ambiguity — you'll be asked to define what "good" means for a model behavior that doesn't have an established benchmark
Pay: $160,000.00 - $260,000.00 per year
Work Location: Remote