{"schemaVersion":"jobsearcher.job.v1","id":"1e9d40b51dd0952e85b5ec93","url":"https://jobsearcher.com/jobs/1e9d40b51dd0952e85b5ec93","canonicalUrl":"https://jobsearcher.com/jobs/1e9d40b51dd0952e85b5ec93","title":"Founding Machine Learning / Data Engineer","description":"About the roleRunning CI on real hardware produces a kind of data most ML teams never get to touch: power traces, sensor streams, build artifacts, console logs, test results, all tied to specific physical machines doing specific physical work. There's serious signal in there — predicting flaky boards, classifying failure modes, scoring job risk, and eventually closed-loop optimization of how the fleet itself is used.\r\nWe're hiring a founding ML / data engineer to build the pipeline that turns that data into models, and the models into product features. You'll own this end-to-end — not \"hand off a notebook and hope.\" Data ingestion, labeling, training infrastructure, evaluation, deployment, monitoring. You'll set the technical direction for ML at Primitive and shape what good looks like — from the first models in production to ML as a core part of the product.\r\nThis is the first dedicated ML hire. It's a build-from-zero role on top of a rich, real-world dataset.\r\nWhat you'll doDesign and build the data platform end-to-end: extend instrumentation where signals are missing today, then ingest from Postgres, SeaweedFS / S3, and streaming telemetry into a clean, versioned analytical layer\r\nBuild the labeling workflow that lets us (and eventually customers) label hardware events without it becoming a permanent side project\r\nDesign and operate a reproducible training stack on AWS — distributed where it needs to be, with experiment tracking, dataset versioning, and a real eval harness\r\nShip inference for product features: low-latency serving where it matters, batch scoring where it fits\r\nOperate models in production: drift monitoring, regression gates, the dashboards that tell us when a model is silently rotting\r\nPartner with the full-stack and hardware teams to integrate predictions cleanly into the product surface\r\nSet the bar for ML rigor at Primitive: eval-first development, reproducibility, honest reporting of model quality\r\nMentor new engineers on data and ML patterns as the team grows; raise the bar on data contracts, eval design, and reviewability\r\nAbout you5+ years in ML / data engineering, with at least one production ML system you took from raw data to served predictions\r\nStrong Python; comfortable with modern ML tooling (PyTorch, Hugging Face, Ray, or equivalents — we're not religious)\r\nReal opinions about data versioning, feature stores, and experiment tracking — you've used DVC, LakeFS, MLflow, or Weights & Biases and know what each is good and bad at\r\nProduction data pipeline experience with a real data warehouse — schema design, contracts, ownership\r\nAWS chops: S3, EKS-hosted training (or SageMaker), IAM that doesn't terrify the security team\r\nBuilt or operated a human-in-the-loop labeling workflow\r\nComfortable setting architectural direction for ML/data at a small company, balancing vision with pragmatism\r\nBonusTime-series, sensor, or signal-processing ML — we have a lot of it\r\nLLM fine-tuning, retrieval, or agent eval experience (there's product surface here too)\r\nBackground in hardware, EE, or anything physical — helps a lot when the data is from real machines\r\nClickHouse, dbt, Airflow / Dagster / Prefect at production scale\r\nContributed to open-source ML or data tooling\r\nStackPython, PyTorch, AWS (S3, EKS, possibly SageMaker), PostgreSQL, SeaweedFS, Grafana / Mimir. Pipeline orchestration is open (Airflow / Dagster / Prefect).\r\nHow we workOffices inNew York, NYandSan Francisco, CA— flexible in-office attendance, no fixed days per week\r\nWe like working together in person: regular team meetups across both offices\r\nSmall team, high ownership — most engineers ship to production in their first week\r\nLight on-call for serving infrastructure once models are in production\r\nBenefitsHealth, dental, and vision for you and your dependents\r\nSubstantial equity, with early exercise and an extended post-termination exercise window\r\nUnlimited paid vacation, plus local holidays\r\nEquinox membership\r\nCompensationBase salary:$180,000 – $250,000 , adjusted based on location, level, and experience. Total compensation includes substantial equity and the benefits above.\r\nA note on applyingIf you don't tick every box on the list above, apply anyway. The bullets describe the engineer we'd be thrilled to hire; what we actually need is someone who can do the work and learn the rest. We especially encourage applications from people who don't see themselves represented in tech today.#J-18808-Ljbffr","company":"Primitive Instruments","rawCompany":"primitive instruments","city":"Brooklyn","state":"NY","isRemote":false,"isActive":false,"createdAt":"2026-09-26T01:26:45.547Z","occupations":[{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Founding Machine Learning / Data Engineer","description":"About the roleRunning CI on real hardware produces a kind of data most ML teams never get to touch: power traces, sensor streams, build artifacts, console logs, test results, all tied to specific physical machines doing specific physical work. There's serious signal in there — predicting flaky boards, classifying failure modes, scoring job risk, and eventually closed-loop optimization of how the fleet itself is used.\r\nWe're hiring a founding ML / data engineer to build the pipeline that turns that data into models, and the models into product features. You'll own this end-to-end — not \"hand off a notebook and hope.\" Data ingestion, labeling, training infrastructure, evaluation, deployment, monitoring. You'll set the technical direction for ML at Primitive and shape what good looks like — from the first models in production to ML as a core part of the product.\r\nThis is the first dedicated ML hire. It's a build-from-zero role on top of a rich, real-world dataset.\r\nWhat you'll doDesign and build the data platform end-to-end: extend instrumentation where signals are missing today, then ingest from Postgres, SeaweedFS / S3, and streaming telemetry into a clean, versioned analytical layer\r\nBuild the labeling workflow that lets us (and eventually customers) label hardware events without it becoming a permanent side project\r\nDesign and operate a reproducible training stack on AWS — distributed where it needs to be, with experiment tracking, dataset versioning, and a real eval harness\r\nShip inference for product features: low-latency serving where it matters, batch scoring where it fits\r\nOperate models in production: drift monitoring, regression gates, the dashboards that tell us when a model is silently rotting\r\nPartner with the full-stack and hardware teams to integrate predictions cleanly into the product surface\r\nSet the bar for ML rigor at Primitive: eval-first development, reproducibility, honest reporting of model quality\r\nMentor new engineers on data and ML patterns as the team grows; raise the bar on data contracts, eval design, and reviewability\r\nAbout you5+ years in ML / data engineering, with at least one production ML system you took from raw data to served predictions\r\nStrong Python; comfortable with modern ML tooling (PyTorch, Hugging Face, Ray, or equivalents — we're not religious)\r\nReal opinions about data versioning, feature stores, and experiment tracking — you've used DVC, LakeFS, MLflow, or Weights & Biases and know what each is good and bad at\r\nProduction data pipeline experience with a real data warehouse — schema design, contracts, ownership\r\nAWS chops: S3, EKS-hosted training (or SageMaker), IAM that doesn't terrify the security team\r\nBuilt or operated a human-in-the-loop labeling workflow\r\nComfortable setting architectural direction for ML/data at a small company, balancing vision with pragmatism\r\nBonusTime-series, sensor, or signal-processing ML — we have a lot of it\r\nLLM fine-tuning, retrieval, or agent eval experience (there's product surface here too)\r\nBackground in hardware, EE, or anything physical — helps a lot when the data is from real machines\r\nClickHouse, dbt, Airflow / Dagster / Prefect at production scale\r\nContributed to open-source ML or data tooling\r\nStackPython, PyTorch, AWS (S3, EKS, possibly SageMaker), PostgreSQL, SeaweedFS, Grafana / Mimir. Pipeline orchestration is open (Airflow / Dagster / Prefect).\r\nHow we workOffices inNew York, NYandSan Francisco, CA— flexible in-office attendance, no fixed days per week\r\nWe like working together in person: regular team meetups across both offices\r\nSmall team, high ownership — most engineers ship to production in their first week\r\nLight on-call for serving infrastructure once models are in production\r\nBenefitsHealth, dental, and vision for you and your dependents\r\nSubstantial equity, with early exercise and an extended post-termination exercise window\r\nUnlimited paid vacation, plus local holidays\r\nEquinox membership\r\nCompensationBase salary:$180,000 – $250,000 , adjusted based on location, level, and experience. Total compensation includes substantial equity and the benefits above.\r\nA note on applyingIf you don't tick every box on the list above, apply anyway. The bullets describe the engineer we'd be thrilled to hire; what we actually need is someone who can do the work and learn the rest. We especially encourage applications from people who don't see themselves represented in tech today.#J-18808-Ljbffr","datePosted":"2026-09-26T01:26:45.547Z","dateModified":"2026-09-26T01:26:45.547Z","hiringOrganization":{"@type":"Organization","name":"Primitive Instruments","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Brooklyn","addressRegion":"NY","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"1e9d40b51dd0952e85b5ec93"},"url":"https://jobsearcher.com/jobs/1e9d40b51dd0952e85b5ec93"}}