{"schemaVersion":"jobsearcher.job.v1","id":"c64c3de4d4f11ab5dd96d4e3","url":"https://jobsearcher.com/jobs/c64c3de4d4f11ab5dd96d4e3","canonicalUrl":"https://jobsearcher.com/jobs/c64c3de4d4f11ab5dd96d4e3","title":"Machine Learning Engineer","description":"Who We Are:\nWe build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models.\nWe're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.\nWe’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI.\nThe Role.\nWe are looking for a Machine Learning Engineer who is skilled and has significant experience in developing machine learning models. You will join a small, innovative team and lead efforts to advance our capabilities, drive model development, and support our vision for a future where Grass is transformative in the evolution of the internet.\nWho You Are.\nBachelor’s, Master’s, or Doctoral degree in Data Science, Computer Science, Statistics, or a related field.\nA minimum of 3 years of work or research experience dealing with large datasets.\nExperience working with large-scale text datasets, NLP pipelines, or data preparation for LLM training is highly preferred.\nStrong coding skills in Python or other object-oriented programming languages.\nExperience with text deduplication, dataset filtering, corpus curation, or data distillation is a strong plus.\nGraduate-level knowledge of statistics, including but not limited to hypothesis testing, regression analysis, and probability.\nExcellent work ethic and the ability to thrive in a fast-paced startup environment.\nStrong problem-solving skills and attention to detail.\nGood communication skills, with the ability to articulate complex data concepts to non-technical stakeholders.\nExperience working in a high output team.\nWhat You'll Be Doing.\nDeveloping data processing pipelines and machine learning solutions for large-scale NLP and LLM applications, including improving the quality, filtering, and preparation of training datasets.\nDesigning and implementing pipelines for processing and analyzing large datasets.\nAnalyzing and interpreting complex time series data to provide actionable insights and solutions.\nDesigning, implementing, and maintaining data-driven models and algorithms.\nDeveloping techniques for dataset curation to improve the quality and efficiency of AI training data.\nBuilding scalable pipelines for filtering, deduplicating, and improving large-scale text datasets used for LLM training.\nCollaborating with cross-functional teams to understand data needs and deliver timely solutions.\nEnsuring data quality and integrity throughout all processes.\nUtilizing Optical Character Recognition (OCR) technology to convert different types of documents into editable and searchable data.\nContinuously researching and implementing best practices in data science and machine learning.\nContributing to the development and improvement of internal data processing tools and infrastructure.\nWhy Work With Us:\nOpportunity. We are at the forefront of developing a web-scale crawler and knowledge graph that improves access to public web data and extends the value of AI to the people.\nCulture. We're a lean team with a high bar. We come to work not to be comfortable, but to find out what we're capable of and to do work that matters. We're not calling for people who keep things moving. We're calling for people who make everyone around them better.\nWe prioritize low ego and high output. This is a fully remote team.\nCompensation. You’ll receive a competitive salary, benefits and equity package.","company":"Wynd Labs","rawCompany":"wynd labs","city":"Denver","state":"CO","isRemote":false,"isActive":false,"createdAt":"2026-08-14T12:19:45.277Z","occupations":[{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Machine Learning Engineer","description":"Who We Are:\nWe build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models.\nWe're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.\nWe’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI.\nThe Role.\nWe are looking for a Machine Learning Engineer who is skilled and has significant experience in developing machine learning models. You will join a small, innovative team and lead efforts to advance our capabilities, drive model development, and support our vision for a future where Grass is transformative in the evolution of the internet.\nWho You Are.\nBachelor’s, Master’s, or Doctoral degree in Data Science, Computer Science, Statistics, or a related field.\nA minimum of 3 years of work or research experience dealing with large datasets.\nExperience working with large-scale text datasets, NLP pipelines, or data preparation for LLM training is highly preferred.\nStrong coding skills in Python or other object-oriented programming languages.\nExperience with text deduplication, dataset filtering, corpus curation, or data distillation is a strong plus.\nGraduate-level knowledge of statistics, including but not limited to hypothesis testing, regression analysis, and probability.\nExcellent work ethic and the ability to thrive in a fast-paced startup environment.\nStrong problem-solving skills and attention to detail.\nGood communication skills, with the ability to articulate complex data concepts to non-technical stakeholders.\nExperience working in a high output team.\nWhat You'll Be Doing.\nDeveloping data processing pipelines and machine learning solutions for large-scale NLP and LLM applications, including improving the quality, filtering, and preparation of training datasets.\nDesigning and implementing pipelines for processing and analyzing large datasets.\nAnalyzing and interpreting complex time series data to provide actionable insights and solutions.\nDesigning, implementing, and maintaining data-driven models and algorithms.\nDeveloping techniques for dataset curation to improve the quality and efficiency of AI training data.\nBuilding scalable pipelines for filtering, deduplicating, and improving large-scale text datasets used for LLM training.\nCollaborating with cross-functional teams to understand data needs and deliver timely solutions.\nEnsuring data quality and integrity throughout all processes.\nUtilizing Optical Character Recognition (OCR) technology to convert different types of documents into editable and searchable data.\nContinuously researching and implementing best practices in data science and machine learning.\nContributing to the development and improvement of internal data processing tools and infrastructure.\nWhy Work With Us:\nOpportunity. We are at the forefront of developing a web-scale crawler and knowledge graph that improves access to public web data and extends the value of AI to the people.\nCulture. We're a lean team with a high bar. We come to work not to be comfortable, but to find out what we're capable of and to do work that matters. We're not calling for people who keep things moving. We're calling for people who make everyone around them better.\nWe prioritize low ego and high output. This is a fully remote team.\nCompensation. You’ll receive a competitive salary, benefits and equity package.","datePosted":"2026-08-14T12:19:45.277Z","dateModified":"2026-08-14T12:19:45.277Z","hiringOrganization":{"@type":"Organization","name":"Wynd Labs","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Denver","addressRegion":"CO","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"c64c3de4d4f11ab5dd96d4e3"},"url":"https://jobsearcher.com/jobs/c64c3de4d4f11ab5dd96d4e3"}}