{"schemaVersion":"jobsearcher.job.v1","id":"0927fa7fe7c73abff29eaf8d","url":"https://jobsearcher.com/jobs/0927fa7fe7c73abff29eaf8d","canonicalUrl":"https://jobsearcher.com/jobs/0927fa7fe7c73abff29eaf8d","title":"Sr. AI Data Engineer","description":"iSoftStone, Inc. is seeking a Sr. AI Data Engineer (Image Generation Data) to join our Team!\r\nThis is a Contract ONSITE Opportunity in Menlo Park, CA\r\nThis is a one-year contract role, and candidates must have permanent authorization to work in the United States. Visa sponsorship is not available for this role and 3 rd party vendor candidates cannot be considered.\r\nSummary\r\nGenerative AI models are only as good as the data they consume. Unlike traditional data engineering, building data pipelines for generative AI requires orchestrating ML model invocations (content understanding classifiers, embedding models, LLM-based cleaners) alongside standard SQL-based transformations, all at billion-row scale.\r\nThis role sits at the intersection of Data Engineering and ML Systems. The Senior AI Data Engineer will own end-to-end data pipelines that don't just move and transform data, but enrich it through remote model inference, managing the systems complexity of async execution, capacity allocation, retry/fallback logic, and throughput optimization that comes with it. This is not a pure ETL-with-SQL role; it demands hands-on systems experience with distributed inference infrastructure.\r\nOur team develops comprehensive data curation and evaluation solutions for image generation models across quality dimensions including visual quality, prompt adherence, identity preservation, naturalness, and visual text generation.\r\nResponsibilities: Main Responsibilities: AI-Augmented Data Pipelines: Design and maintain AI-augmented, large-scale data pipelines (billions of images) integrating traditional transformations with ML models (classifiers, embeddings, LLMs) for cleaning and annotation.\r\nRemote Inference Orchestration: Own the systems for remote ML model inference orchestration within pipelines, managing batching, retries, async jobs, and ensuring graceful degradation.\r\nFeature Pipelines: Build and maintain scalable pipelines for generating, storing, and serving vector embeddings, including nearest-neighbor index management and quality validation.\r\nData Curation at Scale: Source, filter, and curate training datasets using a combination of SQL and model-derived signals (e.g., aesthetic scores, NSFW classifiers), owning the end-to-end data flow and maintaining governance, quality, and compliance.\r\nAdditional Responsibilities\r\nLLM-Assisted Annotation: Design and operate pipelines that use LLMs and vision models for automated annotation of training data, including auditing workflows to measure and improve annotation model performance.\r\nTooling & Frameworks: Contribute to shared tooling and frameworks that make it easier for the broader team to build AI-augmented data pipelines — e.g., reusable operators for model invocation, standard patterns for async job management.\r\nQualifications\r\nAdvanced SQL & data pipeline expertise. Complex queries, query optimization, pipeline orchestration frameworks (Airflow, Dataswarm, or equivalent).\r\nExperience integrating ML models into data pipelines. Calling inference endpoints, managing model versions, batching requests, handling inference failures at scale.\r\nProficiency with AI-assisted coding agents (e.g., Copilot, Cursor, Codex). Expected to leverage AI tools as a force multiplier for writing, debugging, and reviewing code, building pipelines faster, and accelerating day-to-day engineering workflows\r\nStrong verbal and written communication skills, problem-solving ability, and cross-functional collaboration. Preferred\r\nWorking knowledge of embeddings and vector representations like generating, storing, indexing, and querying embeddings (FAISS, Milvus, or equivalent).\r\nFamiliarity with content-understanding models like image classifiers, object detection, OCR, NSFW detection, aesthetic scoring.\r\nExperience with LLMs for data tasks like prompt engineering for annotation, data cleaning, or evaluation using LLM APIs.\r\nKnowledge of generative AI like diffusion models, image generation, evaluation metrics (FID, CLIP score, etc.).\r\nEducation / Experience\r\nBachelor's degree or higher in Computer Science, Data Engineering, Machine Learning, or a related STEM field.\r\n5+ years of industry experience in data engineering, ML engineering, or a hybrid role involving both data pipelines and model serving/inference.\r\nDemonstrated track record of building and operating production data pipelines that invoke ML models at scale.\r\nPrevious experience at Meta is preferred but not required.\r\nAdditional Requirements\r\nWork onsite in MPK 5 days per week, working closely with engineers and researchers.\r\nPrimary Location Pay Range: $105,000 - $110,000 per year\r\nBenefits\r\n1099/Contractors: No benefits\r\nTemp salaried employee benefits, if scheduled to work at least 30 hours per week: medical, dental, vision, 401k, holidays.\r\niSoftStone is a global IT service and consulting company that creates value and drives success through technology solutions, service excellence, and digital innovation. We specialize in web and application development, software testing and support, data and content management, digital experience, accessibility, and data for machine learning and AI. With 20 delivery centers and more than 90,000 employees worldwide, iSoftStone is proud to serve some of the world's most well-known businesses, including 90+ Fortune Global 500 companies.\r\nVisit us at https://www.isoftstoneinc.com .\r\niSoftStone is committed to the practice of equal opportunity for all its employees and applicants in employment, and does not discriminate on the basis of race or ethnicity, color, age, national origin, religion, creed, marital status, sex, pregnancy, gender, gender identity, sexual orientation, status as an honorably discharged veteran or disabled veteran or military status, political affiliation or belief, citizenship/status as a lawfully admitted immigrant authorized to work in the United States, or presence of any physical, sensory, or mental disability. In addition, reasonable accommodation will be made for known physical or mental limitations for all otherwise qualified persons with disabilities.\r\nJ-18808-Ljbffr","company":"Socket","rawCompany":"socket","city":"Menlo Park","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-08T01:49:43.052Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Sr. AI Data Engineer","description":"iSoftStone, Inc. is seeking a Sr. AI Data Engineer (Image Generation Data) to join our Team!\r\nThis is a Contract ONSITE Opportunity in Menlo Park, CA\r\nThis is a one-year contract role, and candidates must have permanent authorization to work in the United States. Visa sponsorship is not available for this role and 3 rd party vendor candidates cannot be considered.\r\nSummary\r\nGenerative AI models are only as good as the data they consume. Unlike traditional data engineering, building data pipelines for generative AI requires orchestrating ML model invocations (content understanding classifiers, embedding models, LLM-based cleaners) alongside standard SQL-based transformations, all at billion-row scale.\r\nThis role sits at the intersection of Data Engineering and ML Systems. The Senior AI Data Engineer will own end-to-end data pipelines that don't just move and transform data, but enrich it through remote model inference, managing the systems complexity of async execution, capacity allocation, retry/fallback logic, and throughput optimization that comes with it. This is not a pure ETL-with-SQL role; it demands hands-on systems experience with distributed inference infrastructure.\r\nOur team develops comprehensive data curation and evaluation solutions for image generation models across quality dimensions including visual quality, prompt adherence, identity preservation, naturalness, and visual text generation.\r\nResponsibilities: Main Responsibilities: AI-Augmented Data Pipelines: Design and maintain AI-augmented, large-scale data pipelines (billions of images) integrating traditional transformations with ML models (classifiers, embeddings, LLMs) for cleaning and annotation.\r\nRemote Inference Orchestration: Own the systems for remote ML model inference orchestration within pipelines, managing batching, retries, async jobs, and ensuring graceful degradation.\r\nFeature Pipelines: Build and maintain scalable pipelines for generating, storing, and serving vector embeddings, including nearest-neighbor index management and quality validation.\r\nData Curation at Scale: Source, filter, and curate training datasets using a combination of SQL and model-derived signals (e.g., aesthetic scores, NSFW classifiers), owning the end-to-end data flow and maintaining governance, quality, and compliance.\r\nAdditional Responsibilities\r\nLLM-Assisted Annotation: Design and operate pipelines that use LLMs and vision models for automated annotation of training data, including auditing workflows to measure and improve annotation model performance.\r\nTooling & Frameworks: Contribute to shared tooling and frameworks that make it easier for the broader team to build AI-augmented data pipelines — e.g., reusable operators for model invocation, standard patterns for async job management.\r\nQualifications\r\nAdvanced SQL & data pipeline expertise. Complex queries, query optimization, pipeline orchestration frameworks (Airflow, Dataswarm, or equivalent).\r\nExperience integrating ML models into data pipelines. Calling inference endpoints, managing model versions, batching requests, handling inference failures at scale.\r\nProficiency with AI-assisted coding agents (e.g., Copilot, Cursor, Codex). Expected to leverage AI tools as a force multiplier for writing, debugging, and reviewing code, building pipelines faster, and accelerating day-to-day engineering workflows\r\nStrong verbal and written communication skills, problem-solving ability, and cross-functional collaboration. Preferred\r\nWorking knowledge of embeddings and vector representations like generating, storing, indexing, and querying embeddings (FAISS, Milvus, or equivalent).\r\nFamiliarity with content-understanding models like image classifiers, object detection, OCR, NSFW detection, aesthetic scoring.\r\nExperience with LLMs for data tasks like prompt engineering for annotation, data cleaning, or evaluation using LLM APIs.\r\nKnowledge of generative AI like diffusion models, image generation, evaluation metrics (FID, CLIP score, etc.).\r\nEducation / Experience\r\nBachelor's degree or higher in Computer Science, Data Engineering, Machine Learning, or a related STEM field.\r\n5+ years of industry experience in data engineering, ML engineering, or a hybrid role involving both data pipelines and model serving/inference.\r\nDemonstrated track record of building and operating production data pipelines that invoke ML models at scale.\r\nPrevious experience at Meta is preferred but not required.\r\nAdditional Requirements\r\nWork onsite in MPK 5 days per week, working closely with engineers and researchers.\r\nPrimary Location Pay Range: $105,000 - $110,000 per year\r\nBenefits\r\n1099/Contractors: No benefits\r\nTemp salaried employee benefits, if scheduled to work at least 30 hours per week: medical, dental, vision, 401k, holidays.\r\niSoftStone is a global IT service and consulting company that creates value and drives success through technology solutions, service excellence, and digital innovation. We specialize in web and application development, software testing and support, data and content management, digital experience, accessibility, and data for machine learning and AI. With 20 delivery centers and more than 90,000 employees worldwide, iSoftStone is proud to serve some of the world's most well-known businesses, including 90+ Fortune Global 500 companies.\r\nVisit us at https://www.isoftstoneinc.com .\r\niSoftStone is committed to the practice of equal opportunity for all its employees and applicants in employment, and does not discriminate on the basis of race or ethnicity, color, age, national origin, religion, creed, marital status, sex, pregnancy, gender, gender identity, sexual orientation, status as an honorably discharged veteran or disabled veteran or military status, political affiliation or belief, citizenship/status as a lawfully admitted immigrant authorized to work in the United States, or presence of any physical, sensory, or mental disability. In addition, reasonable accommodation will be made for known physical or mental limitations for all otherwise qualified persons with disabilities.\r\nJ-18808-Ljbffr","datePosted":"2026-08-08T01:49:43.052Z","dateModified":"2026-08-08T01:49:43.052Z","hiringOrganization":{"@type":"Organization","name":"Socket","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Menlo Park","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"0927fa7fe7c73abff29eaf8d"},"url":"https://jobsearcher.com/jobs/0927fa7fe7c73abff29eaf8d"}}