{"schemaVersion":"jobsearcher.job.v1","id":"12c9fa63dcf4b3300764a014","url":"https://jobsearcher.com/jobs/12c9fa63dcf4b3300764a014","canonicalUrl":"https://jobsearcher.com/jobs/12c9fa63dcf4b3300764a014","title":"Data Engineer","description":"Data Engineer / Datamart DeveloperAbout the RoleWe are seeking a highly skilled and innovative Data Engineer / Datamart Developer to join our Global Pharmaceutical External Develop and Manufacturing Solutions (EDMS) organization. This role is central to building and maintaining the enterprise data foundation that powers advanced analytics, machine learning, and AI-driven insights across drug development, clinical operations, manufacturing, insourcing and outsourcing functions.The ideal candidate brings deep expertise in Snowflake, ETL pipeline development, datamart architecture, and data view engineering — with a strong focus on enabling AI data linkage across multiple heterogeneous pharmaceutical databases.This is a highly technical, hands-on role requiring the ability to design scalable, governed, and GxP-aware data solutions that serve both human analysts and AI/ML model pipelines.Key ResponsibilitiesDatamart Architecture & DevelopmentDesign, build, and maintain enterprise-grade datamarts and data warehouses on the Snowflake platformDevelop scalable dimensional models (star and snowflake schemas) tailored to pharmaceutical business domains including clinical, regulatory, safety, manufacturing, and commercialArchitect and implement data views, materialized views, and virtual layers optimized for AI/ML consumption and business intelligence reportingDefine and enforce data modeling standards, naming conventions, and schema governance across the data platformCollaborate with data architects to align datamart design with the broader enterprise data mesh or data lakehouse strategyOptimize Snowflake query performance through clustering keys, micro-partitioning strategies, and warehouse sizingETL / ELT Pipeline DevelopmentDesign, develop, and maintain robust ETL/ELT pipelines to ingest, transform, and load data from multiple source systems into SnowflakeBuild and manage pipelines sourcing data from diverse pharmaceutical systemsImplement incremental load strategies, change data capture (CDC), and slowly changing dimension (SCD) patternsDevelop and maintain pipeline orchestration using tools such as Dataiku or InformaticaMonitor pipeline health, implement alerting, and manage incident resolution for production data flowsAI Data Linkage & IntegrationDevelop and maintain data views and semantic layers purpose-built for AI and machine learning model training, validation, and inference pipelinesDesign entity resolution and data linkage frameworks to connect and harmonize data across multiple disparate pharmaceutical databases (clinical, genomic, safety, real-world evidence, commercial)Build feature stores and curated datasets that serve downstream AI/ML use cases including drug discovery, patient stratification, adverse event detection, and commercial forecastingPartner with data scientists and AI/ML engineers to understand model data requirements and translate them into governed, reproducible data assetsImplement master data management (MDM) principles to ensure consistent entity identification across linked data sourcesSupport integration of external data sources including real-world data (RWD), claims data, genomic databases, and third-party biomarker datasetsData Quality & GovernanceImplement and enforce data quality frameworks including profiling, validation rules, anomaly detection, and data quality scorecardsDefine and maintain data contracts between source systems and downstream consumersCollaboration & Stakeholder EngagementPartner with data scientists, AI/ML engineers, bioinformaticians, and business analysts to deliver data solutions that accelerate research and commercial outcomesCollaborate with IT infrastructure, cloud engineering, and security teams on Snowflake environment management, role-based access control (RBAC), and data encryptionTechnical Knowledge RequirementsSnowflakeExpert-level proficiency in Snowflake platform including database, schema, and table designDeep experience with Snowflake-specific features: Time Travel, Zero-Copy Cloning, Data Sharing, Streams, Tasks, and Dynamic TablesSnowflake performance optimization: clustering keys, search optimization, result caching, and virtual warehouse managementExperience with Snowflake data marketplace for ingesting third-party pharmaceutical and real-world datasetsFamiliarity with Snowpark for Python/Java-based data transformation within SnowflakeRole-based access control (RBAC) and column/row-level security implementation in SnowflakeETL / ELT & Pipeline OrchestrationStrong hands-on experience with one or more ETL/ELT toolsExperience with dbt for SQL-based transformation layer development, testing, and documentationProficiency in designing incremental, full-load, and CDC-based pipeline strategiesPipeline monitoring, alerting, and SLA management in production environmentsSQL & Data ModelingExpert-level SQL including complex joins, window functions, CTEs, and recursive queriesExperience designing semantic and consumption layers for BI tools and AI/ML pipelinesFamiliarity with SQL Server, PostgreSQL, or Oracle as source system databasesProgramming & ScriptingProficiency in Python for data pipeline development, data transformation, and automationExperience with PySpark or Spark SQL for large-scale data processing is a plusFamiliarity with REST API development and consumption for data ingestion from third-party platformsAI / ML Data EngineeringExperience building feature stores for ML model consumptionUnderstanding of vector databases (Pinecone, Weaviate, or pgvector) for LLM and AI search use casesExperience designing training, validation, and inference datasets with reproducibility and versioning in mindQualificationsEducationBachelor's degree in Computer Science, Data Science, Information Systems, Bioinformatics, Engineering, or related field requiredMaster's degree in Data Science, Computational Biology, or related discipline preferredExperience3+ years of hands-on data engineering experience with a strong focus on datamart development and ETL pipeline delivery2+ years of production-level Snowflake experience in an enterprise environmentDemonstrated experience building AI/ML-ready data pipelines and feature engineering frameworksProven track record delivering data solutions that span multiple heterogeneous source systemsCertifications (Preferred)Snowflake SnowPro Core CertificationSnowflake SnowPro Advanced: Data Engineerdbt Analytics Engineering CertificationMicrosoft Certified: Azure Data Engineer Associate (DP-203)Skills & CompetenciesCore CompetenciesDeep technical expertise balanced with strong communication skills to engage both engineering and business stakeholdersAbility to design elegant, scalable data solutions that serve diverse consumers from BI analysts to AI/ML engineersStrong problem-solving and debugging skills for complex, multi-system data integration challengesDetail-oriented approach to data quality, lineage, and governance in regulated pharmaceutical environmentsTools & PlatformsData platforms: Snowflake, Azure Synapse, DatabricksETL/ELT: dbt, Azure Data Factory, Airflow, Informatica, MatillionData cataloging: Collibra, Alation, Microsoft PurviewBI & visualization: Spotfire, Power BIDevelopment: Python, SQL,Nice to HaveExperience with Figma for collaborating on data product UI/UX wireframes and dashboard design handoffs with analytics and product teamsExposure to generative AI or LLM integration patterns using enterprise pharmaceutical dataCompensation:$60-85/hrExact compensation may vary based on several factors, including skills, experience, and education.Benefit packages for this role will start on the 31st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law.","company":"Insight Global","rawCompany":"insight global","city":"Groton","state":"CT","isRemote":false,"isActive":false,"createdAt":"2026-08-17T12:17:34.465Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1243.00","title":"Database Architects","slug":"database-architects"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Data Engineer","description":"Data Engineer / Datamart DeveloperAbout the RoleWe are seeking a highly skilled and innovative Data Engineer / Datamart Developer to join our Global Pharmaceutical External Develop and Manufacturing Solutions (EDMS) organization. This role is central to building and maintaining the enterprise data foundation that powers advanced analytics, machine learning, and AI-driven insights across drug development, clinical operations, manufacturing, insourcing and outsourcing functions.The ideal candidate brings deep expertise in Snowflake, ETL pipeline development, datamart architecture, and data view engineering — with a strong focus on enabling AI data linkage across multiple heterogeneous pharmaceutical databases.This is a highly technical, hands-on role requiring the ability to design scalable, governed, and GxP-aware data solutions that serve both human analysts and AI/ML model pipelines.Key ResponsibilitiesDatamart Architecture & DevelopmentDesign, build, and maintain enterprise-grade datamarts and data warehouses on the Snowflake platformDevelop scalable dimensional models (star and snowflake schemas) tailored to pharmaceutical business domains including clinical, regulatory, safety, manufacturing, and commercialArchitect and implement data views, materialized views, and virtual layers optimized for AI/ML consumption and business intelligence reportingDefine and enforce data modeling standards, naming conventions, and schema governance across the data platformCollaborate with data architects to align datamart design with the broader enterprise data mesh or data lakehouse strategyOptimize Snowflake query performance through clustering keys, micro-partitioning strategies, and warehouse sizingETL / ELT Pipeline DevelopmentDesign, develop, and maintain robust ETL/ELT pipelines to ingest, transform, and load data from multiple source systems into SnowflakeBuild and manage pipelines sourcing data from diverse pharmaceutical systemsImplement incremental load strategies, change data capture (CDC), and slowly changing dimension (SCD) patternsDevelop and maintain pipeline orchestration using tools such as Dataiku or InformaticaMonitor pipeline health, implement alerting, and manage incident resolution for production data flowsAI Data Linkage & IntegrationDevelop and maintain data views and semantic layers purpose-built for AI and machine learning model training, validation, and inference pipelinesDesign entity resolution and data linkage frameworks to connect and harmonize data across multiple disparate pharmaceutical databases (clinical, genomic, safety, real-world evidence, commercial)Build feature stores and curated datasets that serve downstream AI/ML use cases including drug discovery, patient stratification, adverse event detection, and commercial forecastingPartner with data scientists and AI/ML engineers to understand model data requirements and translate them into governed, reproducible data assetsImplement master data management (MDM) principles to ensure consistent entity identification across linked data sourcesSupport integration of external data sources including real-world data (RWD), claims data, genomic databases, and third-party biomarker datasetsData Quality & GovernanceImplement and enforce data quality frameworks including profiling, validation rules, anomaly detection, and data quality scorecardsDefine and maintain data contracts between source systems and downstream consumersCollaboration & Stakeholder EngagementPartner with data scientists, AI/ML engineers, bioinformaticians, and business analysts to deliver data solutions that accelerate research and commercial outcomesCollaborate with IT infrastructure, cloud engineering, and security teams on Snowflake environment management, role-based access control (RBAC), and data encryptionTechnical Knowledge RequirementsSnowflakeExpert-level proficiency in Snowflake platform including database, schema, and table designDeep experience with Snowflake-specific features: Time Travel, Zero-Copy Cloning, Data Sharing, Streams, Tasks, and Dynamic TablesSnowflake performance optimization: clustering keys, search optimization, result caching, and virtual warehouse managementExperience with Snowflake data marketplace for ingesting third-party pharmaceutical and real-world datasetsFamiliarity with Snowpark for Python/Java-based data transformation within SnowflakeRole-based access control (RBAC) and column/row-level security implementation in SnowflakeETL / ELT & Pipeline OrchestrationStrong hands-on experience with one or more ETL/ELT toolsExperience with dbt for SQL-based transformation layer development, testing, and documentationProficiency in designing incremental, full-load, and CDC-based pipeline strategiesPipeline monitoring, alerting, and SLA management in production environmentsSQL & Data ModelingExpert-level SQL including complex joins, window functions, CTEs, and recursive queriesExperience designing semantic and consumption layers for BI tools and AI/ML pipelinesFamiliarity with SQL Server, PostgreSQL, or Oracle as source system databasesProgramming & ScriptingProficiency in Python for data pipeline development, data transformation, and automationExperience with PySpark or Spark SQL for large-scale data processing is a plusFamiliarity with REST API development and consumption for data ingestion from third-party platformsAI / ML Data EngineeringExperience building feature stores for ML model consumptionUnderstanding of vector databases (Pinecone, Weaviate, or pgvector) for LLM and AI search use casesExperience designing training, validation, and inference datasets with reproducibility and versioning in mindQualificationsEducationBachelor's degree in Computer Science, Data Science, Information Systems, Bioinformatics, Engineering, or related field requiredMaster's degree in Data Science, Computational Biology, or related discipline preferredExperience3+ years of hands-on data engineering experience with a strong focus on datamart development and ETL pipeline delivery2+ years of production-level Snowflake experience in an enterprise environmentDemonstrated experience building AI/ML-ready data pipelines and feature engineering frameworksProven track record delivering data solutions that span multiple heterogeneous source systemsCertifications (Preferred)Snowflake SnowPro Core CertificationSnowflake SnowPro Advanced: Data Engineerdbt Analytics Engineering CertificationMicrosoft Certified: Azure Data Engineer Associate (DP-203)Skills & CompetenciesCore CompetenciesDeep technical expertise balanced with strong communication skills to engage both engineering and business stakeholdersAbility to design elegant, scalable data solutions that serve diverse consumers from BI analysts to AI/ML engineersStrong problem-solving and debugging skills for complex, multi-system data integration challengesDetail-oriented approach to data quality, lineage, and governance in regulated pharmaceutical environmentsTools & PlatformsData platforms: Snowflake, Azure Synapse, DatabricksETL/ELT: dbt, Azure Data Factory, Airflow, Informatica, MatillionData cataloging: Collibra, Alation, Microsoft PurviewBI & visualization: Spotfire, Power BIDevelopment: Python, SQL,Nice to HaveExperience with Figma for collaborating on data product UI/UX wireframes and dashboard design handoffs with analytics and product teamsExposure to generative AI or LLM integration patterns using enterprise pharmaceutical dataCompensation:$60-85/hrExact compensation may vary based on several factors, including skills, experience, and education.Benefit packages for this role will start on the 31st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law.","datePosted":"2026-08-17T12:17:34.465Z","dateModified":"2026-08-17T12:17:34.465Z","hiringOrganization":{"@type":"Organization","name":"Insight Global","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Groton","addressRegion":"CT","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"12c9fa63dcf4b3300764a014"},"url":"https://jobsearcher.com/jobs/12c9fa63dcf4b3300764a014"}}