{"schemaVersion":"jobsearcher.job.v1","id":"0e962b2b16c60137e11efff3","url":"https://jobsearcher.com/jobs/0e962b2b16c60137e11efff3","canonicalUrl":"https://jobsearcher.com/jobs/0e962b2b16c60137e11efff3","title":"LLM Data Engineer","description":"Must-Have RequirementsStrong hands-on experience with AWSExperience working with data sets, data sources, and AWS data servicesStrong AI & LLM experienceExcellent communication skillsActive and complete LinkedIn profileHealthcare experience is a plus, but not requiredRole OverviewWe are looking for a Generalist Data Engineer to support a healthcare-focused AI benchmark and evaluation platform.The engineer will be responsible for the data lifecycle before a model receives it and after the model generates a response. This includes building ingestion, normalization, packaging, storage, and analysis layers to transform raw healthcare data into standardized benchmark inputs and actionable evaluation results.Key ResponsibilitiesBuild ingestion and normalization pipelines for:DICOM radiology studiesWhole-slide pathology imagesTabular EHR data, including labs, vitals, encounters, and medication recordsDesign a canonical benchmark record format that can support different task configurations and model adaptersDevelop strategies for packaging large medical imaging datasets within third-party API limitationsBuild tiling, region selection, downsampling, and compression workflows while maintaining diagnostic informationMaintain provenance metadata to ensure results are reproducible and defensibleDevelop cohort and label pipelines for clinical prediction tasks such as sepsis onset, survival horizons, and longitudinal lab trendsBuild results storage and analysis layers for per-run, per-model, and per-task outputsEnforce PHI handling requirements, including encryption, least-privilege access, audit logging, and de-identificationEnsure clear controls around data leaving the VPC when interacting with third-party APIsRequired Skills:Data EngineeringProduction-grade data pipeline developmentStrong testing disciplineExperience building deterministic, idempotent, and re-runnable jobsAWS Data StackDeep hands-on experience with S3AWS Glue and/or Spark on EMRAthenaStep FunctionsLambdaAWS BatchUnderstanding of storage layout, lifecycle policies, and storage economics at scaleSQL & Data ModelingComplex temporal joinsPoint-in-time correctnessStrong understanding of preventing label leakage in time-series dataData Quality & LineageData validation frameworksSchema enforcementVersioned datasetsData lineage and reproducibilityAWS Security & GovernanceIAM policy designKMSVPC endpoints and PrivateLinkExperience working within HIPAA-eligible AWS architecturesDesirable SkillsHealthcare data standards including DICOM, FHIR, HL7v2, OMOP CDMFamiliarity with clinical coding systems such as ICD, LOINC, RxNorm, and SNOMEDMedical imaging experience with pydicom and OpenSlideUnderstanding of WSI pyramid structures and tilingTerraform / Infrastructure as CodeDocker, ECR, and CI/CDFamiliarity with LLM APIs and multimodal payload construction","company":"Xaxis Solutions","rawCompany":"xaxis solutions","city":"Denver","state":"CO","isRemote":false,"isActive":false,"createdAt":"2026-08-14T14:59:36.076Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"621999","title":"All Other Miscellaneous Ambulatory Health Care Services","slug":"all-other-miscellaneous-ambulatory-health-care-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"LLM Data Engineer","description":"Must-Have RequirementsStrong hands-on experience with AWSExperience working with data sets, data sources, and AWS data servicesStrong AI & LLM experienceExcellent communication skillsActive and complete LinkedIn profileHealthcare experience is a plus, but not requiredRole OverviewWe are looking for a Generalist Data Engineer to support a healthcare-focused AI benchmark and evaluation platform.The engineer will be responsible for the data lifecycle before a model receives it and after the model generates a response. This includes building ingestion, normalization, packaging, storage, and analysis layers to transform raw healthcare data into standardized benchmark inputs and actionable evaluation results.Key ResponsibilitiesBuild ingestion and normalization pipelines for:DICOM radiology studiesWhole-slide pathology imagesTabular EHR data, including labs, vitals, encounters, and medication recordsDesign a canonical benchmark record format that can support different task configurations and model adaptersDevelop strategies for packaging large medical imaging datasets within third-party API limitationsBuild tiling, region selection, downsampling, and compression workflows while maintaining diagnostic informationMaintain provenance metadata to ensure results are reproducible and defensibleDevelop cohort and label pipelines for clinical prediction tasks such as sepsis onset, survival horizons, and longitudinal lab trendsBuild results storage and analysis layers for per-run, per-model, and per-task outputsEnforce PHI handling requirements, including encryption, least-privilege access, audit logging, and de-identificationEnsure clear controls around data leaving the VPC when interacting with third-party APIsRequired Skills:Data EngineeringProduction-grade data pipeline developmentStrong testing disciplineExperience building deterministic, idempotent, and re-runnable jobsAWS Data StackDeep hands-on experience with S3AWS Glue and/or Spark on EMRAthenaStep FunctionsLambdaAWS BatchUnderstanding of storage layout, lifecycle policies, and storage economics at scaleSQL & Data ModelingComplex temporal joinsPoint-in-time correctnessStrong understanding of preventing label leakage in time-series dataData Quality & LineageData validation frameworksSchema enforcementVersioned datasetsData lineage and reproducibilityAWS Security & GovernanceIAM policy designKMSVPC endpoints and PrivateLinkExperience working within HIPAA-eligible AWS architecturesDesirable SkillsHealthcare data standards including DICOM, FHIR, HL7v2, OMOP CDMFamiliarity with clinical coding systems such as ICD, LOINC, RxNorm, and SNOMEDMedical imaging experience with pydicom and OpenSlideUnderstanding of WSI pyramid structures and tilingTerraform / Infrastructure as CodeDocker, ECR, and CI/CDFamiliarity with LLM APIs and multimodal payload construction","datePosted":"2026-08-14T14:59:36.076Z","dateModified":"2026-08-14T14:59:36.076Z","hiringOrganization":{"@type":"Organization","name":"Xaxis Solutions","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Denver","addressRegion":"CO","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"0e962b2b16c60137e11efff3"},"url":"https://jobsearcher.com/jobs/0e962b2b16c60137e11efff3"}}