{"schemaVersion":"jobsearcher.job.v1","id":"3fa28d331b8f2aa4862a1116","url":"https://jobsearcher.com/jobs/3fa28d331b8f2aa4862a1116","canonicalUrl":"https://jobsearcher.com/jobs/3fa28d331b8f2aa4862a1116","title":"AI Data Strategy Engineer / Applied Scientist, LLM Data","description":"Fully Remote • Overland Park, KS • Engineering\nJob Type\nFull-time\nDescription\n\nPropio Language Services is a provider of the highest quality interpretation, translation, and localization services. Our people take pride in every resource we offer, and our users always have access to cutting-edge technology, exceptional support, and collaborative user experiences. We are driven by our passion for innovation, growth, and bridging communication gaps in a diverse world. If you’re passionate about delivering technology-driven solutions and building lasting client relationships while contributing to client growth, Propio could be the ideal place for you.\n\nWe are building AI-powered systems that enhance multilingual communication, improve interpreter workflows, and support next-generation AI applications across text, speech, and multimodal experiences.\n\nPropio is hiring an AI Data Strategy Engineer / Applied Scientist, LLM Data to own the data strategy, curation pipelines, annotation workflows, and evaluation datasets that power our multilingual AI systems.\n\nThis is a hands-on technical role for someone who understands how to manage the full AI data lifecycle, from acquisition, curation, annotation, and quality control to evaluation datasets and post-training data, to directly improve model performance.\n\nThe ideal candidate can build scalable data pipelines, design high-quality annotation and QA processes, identify model failure modes, and close performance gaps through targeted data acquisition, curation, and synthetic data generation.\n\nKey Responsibilities:\nDefine the end-to-end data roadmap for multilingual and multimodal AI systems, including text, speech, translation, interpretation, low-resource languages, and agentic AI workflows.\nDesign and build dataset curation pipelines for training, post-training, and evaluation, including cleaning, deduplication, filtering, PII redaction, quality scoring, sampling, balancing, and versioning.\nCreate annotation schemas, labeling guidelines, QA rubrics, golden datasets, and reviewer workflows for multilingual, speech, translation, and agentic AI data.\nBuild evaluation datasets and benchmarks, analyze model failure modes, and translate performance gaps into targeted data improvements.\nSupport post-training data workflows such as SFT, instruction tuning, preference data, RLHF/DPO-style data, reward model data, and synthetic data generation.\nUse modern annotation tools and AWS-based data infrastructure to scale secure, traceable, and compliant AI data workflows.\n\nRequirements\n\nBachelor’s degree in Computer Science, Machine Learning, Data Science, Computational Linguistics, Linguistics, Statistics, or a related field, or equivalent practical experience.\n4+ years of experience in AI data, ML data operations, NLP data engineering, applied ML, speech/translation data, or LLM data workflows.\nStrong hands-on experience with Python, SQL, and dataset curation pipelines.\nExperience with annotation workflows, QA rubrics, evaluation datasets, or human-in-the-loop data processes.\nFamiliarity with multilingual NLP, speech data, translation data, low-resource languages, conversational AI, or agentic AI datasets.\nWorking knowledge of AWS data and ML tools such as S3, Glue, SageMaker, Bedrock, Lambda, Step Functions, EKS/ECS, IAM, or KMS.\nStrong communication skills and ability to work with ML engineers, applied scientists, product teams, linguists, data teams, and vendors.\nPreferred Qualifications\nMaster’s or PhD in Computer Science, Machine Learning, NLP, Computational Linguistics, Data Science, Statistics, or a related field.\nExperience with LLM post-training workflows such as SFT, instruction tuning, preference data, RLHF, DPO, reward modeling, or evaluation data generation.\nExperience with synthetic data generation, active learning, weak supervision, LLM-as-judge workflows, or automated data quality scoring.\nExperience with modern annotation and data platforms such as Labelbox, Scale AI, Prodigy, Argilla, Snorkel, Humanloop, or custom internal tooling.","company":"Propio Ls","rawCompany":"propio ls","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-08-03T18:12:42.811Z","occupations":[{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"},{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"}],"industries":[{"code":"541990","title":"All Other Professional, Scientific, and Technical Services","slug":"all-other-professional-scientific-and-technical-services"},{"code":"541930","title":"Translation and Interpretation Services","slug":"translation-and-interpretation-services"},{"code":"541690","title":"Other Scientific and Technical Consulting Services","slug":"other-scientific-and-technical-consulting-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"AI Data Strategy Engineer / Applied Scientist, LLM Data","description":"Fully Remote • Overland Park, KS • Engineering\nJob Type\nFull-time\nDescription\n\nPropio Language Services is a provider of the highest quality interpretation, translation, and localization services. Our people take pride in every resource we offer, and our users always have access to cutting-edge technology, exceptional support, and collaborative user experiences. We are driven by our passion for innovation, growth, and bridging communication gaps in a diverse world. If you’re passionate about delivering technology-driven solutions and building lasting client relationships while contributing to client growth, Propio could be the ideal place for you.\n\nWe are building AI-powered systems that enhance multilingual communication, improve interpreter workflows, and support next-generation AI applications across text, speech, and multimodal experiences.\n\nPropio is hiring an AI Data Strategy Engineer / Applied Scientist, LLM Data to own the data strategy, curation pipelines, annotation workflows, and evaluation datasets that power our multilingual AI systems.\n\nThis is a hands-on technical role for someone who understands how to manage the full AI data lifecycle, from acquisition, curation, annotation, and quality control to evaluation datasets and post-training data, to directly improve model performance.\n\nThe ideal candidate can build scalable data pipelines, design high-quality annotation and QA processes, identify model failure modes, and close performance gaps through targeted data acquisition, curation, and synthetic data generation.\n\nKey Responsibilities:\nDefine the end-to-end data roadmap for multilingual and multimodal AI systems, including text, speech, translation, interpretation, low-resource languages, and agentic AI workflows.\nDesign and build dataset curation pipelines for training, post-training, and evaluation, including cleaning, deduplication, filtering, PII redaction, quality scoring, sampling, balancing, and versioning.\nCreate annotation schemas, labeling guidelines, QA rubrics, golden datasets, and reviewer workflows for multilingual, speech, translation, and agentic AI data.\nBuild evaluation datasets and benchmarks, analyze model failure modes, and translate performance gaps into targeted data improvements.\nSupport post-training data workflows such as SFT, instruction tuning, preference data, RLHF/DPO-style data, reward model data, and synthetic data generation.\nUse modern annotation tools and AWS-based data infrastructure to scale secure, traceable, and compliant AI data workflows.\n\nRequirements\n\nBachelor’s degree in Computer Science, Machine Learning, Data Science, Computational Linguistics, Linguistics, Statistics, or a related field, or equivalent practical experience.\n4+ years of experience in AI data, ML data operations, NLP data engineering, applied ML, speech/translation data, or LLM data workflows.\nStrong hands-on experience with Python, SQL, and dataset curation pipelines.\nExperience with annotation workflows, QA rubrics, evaluation datasets, or human-in-the-loop data processes.\nFamiliarity with multilingual NLP, speech data, translation data, low-resource languages, conversational AI, or agentic AI datasets.\nWorking knowledge of AWS data and ML tools such as S3, Glue, SageMaker, Bedrock, Lambda, Step Functions, EKS/ECS, IAM, or KMS.\nStrong communication skills and ability to work with ML engineers, applied scientists, product teams, linguists, data teams, and vendors.\nPreferred Qualifications\nMaster’s or PhD in Computer Science, Machine Learning, NLP, Computational Linguistics, Data Science, Statistics, or a related field.\nExperience with LLM post-training workflows such as SFT, instruction tuning, preference data, RLHF, DPO, reward modeling, or evaluation data generation.\nExperience with synthetic data generation, active learning, weak supervision, LLM-as-judge workflows, or automated data quality scoring.\nExperience with modern annotation and data platforms such as Labelbox, Scale AI, Prodigy, Argilla, Snorkel, Humanloop, or custom internal tooling.","datePosted":"2026-08-03T18:12:42.811Z","dateModified":"2026-08-03T18:12:42.811Z","hiringOrganization":{"@type":"Organization","name":"Propio Ls","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"3fa28d331b8f2aa4862a1116"},"url":"https://jobsearcher.com/jobs/3fa28d331b8f2aa4862a1116"}}