Imaging Data Scientist, Machine Learning Engineer, Computer Vision
*About Sonari*
Sonari is a well-funded startup and the world's first company to deploy AI-powered ultrasound imaging in professional sports, turning a fast, portable scan into objective data on an athlete's tissue health, performance, and endurance over time. Our ultimate goal is to predict and prevent soft-tissue and non-contact injuries.
Our technology is grounded in peer-reviewed research. We're backed by a global investment firm ($70B+ AUM), and we're already in partnership with a Major League Baseball team and a Premier League club. Our aim is to put this capability in the hands of every sports organization in the world, from championship rosters to the academies where the next generation is coming up.
We're mission-driven and looking for like-minded top engineers who want to build at the intersection of AI, medical imaging and sports.
*Your role: *
You own the *models and the measurements*: you train the AI models that run on the device and in the cloud, and you build the analysis services that convert ultrasound images and clips into validated quantitative metrics that track athlete performance and injury risk. You'll work in a fast-paced, hands-on environment, reporting directly to the CTO.
This role spans two parts of the system:
* *The Model Factory (offline):* the training pipeline. Move data to and from the labeling platform, preprocess, train, evaluate, and package new metrics models, then ship them back to the edge and the analysis services with full provenance.
* *The Analysis Services (online):* the inference/measurement services that take an exam's images or clips and produce segmentation masks, geometry, and the derived athlete metrics — both on the edge and in the cloud backend.
_Scope note:_ this is a wide surface for one person - model training and productionized inference. We're hiring for someone comfortable owning both early, with the expectation that a second ML/MLOps hire follows. If your strength is decisively one side, tell us - we'd rather place you well.
*What you'll work on: *
* *Image & video processing.* Build robust image- and video-processing pipelines over *DICOM and raw imagery* — still frames and clips — including physical calibration so measurements come out in real-world units.
* *Segmentation & measurement models.* Train and refine segmentation models of musculoskeletal anatomy in ultrasound, and derive the athlete metrics from their geometry — the numbers that end up on a team's dashboard.
* *Navigation models for the edge.* Train the on-device models that recognize views, score image quality, and produce initial segmentation to guide capture — then package them to run on-device (e.g. CoreML) in the navigation app.
* *The labeling loop, with the clinical team.* Own the round-trip with the labeling/annotation platform: push de-identified images to be labeled, pull labels back, and land them in the platform's single canonical *annotation schema* (in-app clinician annotations and platform labels share that schema — you consume both). Work closely with the clinical team to define and continually refine what gets labeled, how, and to what quality bar, so the training data stays consistent and clinically meaningful.
* *Reproducibility & provenance.* Stamp every model with a version + provenance record so any metric traces back to the model that produced it, and maintain the evaluation / regression harness that qualifies a model before it ships.
* *Data governance in practice.* Work only with de-identified data; honor consent gating on the training stream; keep athlete-identifying information out of everything you touch.
Required skills
* Strong *Python* and the scientific stack: *NumPy, SciPy, OpenCV, scikit-image, scikit-learn*.
* *Deep learning for computer vision*, hands-on: *TensorFlow / Keras* and/or *PyTorch*; semantic segmentation (U-Net-style) and/or object detection (YOLO-style) on images.
* Solid ML engineering fundamentals: reproducible training, evaluation metrics (e.g. Dice/IoU for segmentation), regression testing, and reasoning about model quality and failure modes.
* Data discipline: you treat de-identification, consent, and provenance as first-class requirements, not afterthoughts.
Nice to have
* Experience with *medical / scientific imaging* — e.g. DICOM, physical calibration, and turning pixels into real-world units.
* *Sports science, musculoskeletal, or physiology* domain experience.
* Model *deployment to constrained runtimes* — CoreML for iOS, or ONNX.
* Experience with *labeling / annotation platforms* and building label round-trip pipelines.
* Image-stitching / geometric processing.
* MLOps: training reproducibility, model registries, GPU scheduling, containerized inference.
* Experience with *transformer-based or foundation-model approaches* to segmentation (e.g. SAM and its medical variants).
* Awareness of clinical-software constraints (data handling, regulatory pathways) — depth here is a plus, not a requirement; we retain regulatory specialists separately.
Tech stack you'll touch
Python · TensorFlow / Keras · PyTorch · OpenCV · scikit-image · scikit-learn · pydicom · SimpleITK · a labeling / annotation platform · cloud object storage · on-device packaging (CoreML) · Docker · a Python web framework for the analysis services · a shared JSON-Schema-based annotation contract.
How this role fits the team
You'll be key member of our engineering team. The *Mobile iOS Engineer* ships the models you package into the capture app; the *Full-stack Engineer* runs the platform that stores your training data, dispatches your analysis jobs, and displays your metrics. You'll define the model↔platform seam together, working against shared, versioned contracts rather than ad-hoc handoffs.
Pay: $160,000.00 - $220,000.00 per year
Benefits:
* Dental insurance
* Health insurance
* Vision insurance
Work Location: Hybrid remote in Cambridge, MA 02142