{"schemaVersion":"jobsearcher.job.v1","id":"76b6057a938913beee8511a8","url":"https://jobsearcher.com/jobs/76b6057a938913beee8511a8","canonicalUrl":"https://jobsearcher.com/jobs/76b6057a938913beee8511a8","title":"Data Engineer","description":"Our Fortune 50 Healthcare client is seeking a Senior Data Engineer to support our mission of improving the health and well-being of our members. This role will focus on building scalable, secure, data centric solutions and compliant data platforms that power analytics, clinical insights, and business decision-making across the enterprise.\r\nThe ideal candidate will have strong experience with cloud-based data platforms, Databricks, PostgreSQL , and healthcare data, with a passion for delivering high-quality, trusted data solutions in a regulated environment.\r\nKey Responsibilities\r\nDesign, develop, and scalable data pipeline solutions using Databricks (Spark) and cloud-native services\r\nBuild and optimize ETL/ELT workflows for ingesting structured and unstructured healthcare data (claims, clinical, provider, and member data)\r\nDevelop and maintain data models in PostgreSQL and enterprise data warehouses\r\nSupport Lakehouse architecture leveraging Databricks, Delta Lake , and cloud storage\r\nImprove performance, reliability, and cost-efficiency of data platforms\r\nWork with healthcare datasets, including producer/agent, broker, commission, and distribution data, ensuring proper ingestion, normalization, and optimization for analytics and reporting\r\nEnsure compliance with HIPAA, HITECH , and enterprise data governance policies\r\nImplement data security, encryption, masking, and access controls\r\nMaintain data lineage, auditability, and regulatory reporting readiness\r\nAdvanced Data Processing\r\nBuild real-time and batch pipelines for analytics and operational use cases\r\nDevelop data transformations using PySpark and SQL within Databricks\r\nLeverage PostgreSQL for transactional and analytical workloads where applicable\r\nIntegrate data from APIs, third-party vendors, and internal systems\r\nCollaboration & Stakeholder Engagement\r\nPartner with business stakeholders to support data-driven initiatives and member acquisition strategies\r\nTranslate insurance distribution, agent/producer , and marketing requirements into scalable, high-quality data solutions\r\nSupport downstream consumers, including Power BI, marketing analytics teams, and operational reporting stakeholders, by delivering curated, analytics-ready datasets\r\nTechnical Leadership\r\nLead design and architecture discussions for enterprise data solutions\r\nEstablish and enforce best practices in data engineering, testing, and CI/CD\r\nContribute to enterprise data strategy and platform modernization\r\nAI & Advanced Analytics (Databricks Genie)\r\nLeverage Databricks Genie (AI/BI capabilities) to enable natural language querying and democratize data access for business stakeholders\r\nDesign and optimize semantic layers and governed datasets that power Genie-driven insights with trusted, high-quality data\r\nCollaborate with stakeholders to translate business questions into AI-assisted analytics workflows using Databricks\r\nEnsure AI outputs are accurate, explainable, and compliant with healthcare data governance and HIPAA requirements\r\nLeverage large language models (LLMs), including Anthropic Claude , to enhance data exploration, automate insight generation, and support conversational analytics use cases\r\nIntegrate Genie capabilities with Delta Lake and curated data models to support near real-time insights and decision-making\r\nPartner with data scientists and analytics teams to enhance AI-driven use cases, including producer performance insights, marketing attribution, and member engagement analysis\r\nRequired Qualifications\r\nBachelor's or Master's degree in Computer Science, Engineering, or related field\r\n5–8+ years of experience in data engineering\r\nStrong programming in Python (PySpark) and advanced SQL\r\nHands-on experience with:\r\nDatabricks (core requirement)\r\nDistributed data processing frameworks (Apache Spark)\r\nExperience with cloud platforms (Azure preferred; AWS acceptable)\r\nProficiency in building and maintaining ETL/ELT pipelines\r\nStrong understanding of data modeling and warehousing concepts\r\nPreferred Qualifications\r\nExperience in healthcare or insurance industry (payer experience strongly preferred)\r\nFamiliarity with healthcare standards (e.g., FHIR, HL7 )\r\nExperience with\r\nDelta Lake / Lakehouse architecture\r\nKnowledge of DevOps and CI/CD pipelines (Azure DevOps, GitHub Actions)\r\nExperience supporting machine learning pipelines\r\nDeep understanding of data pipelines at scale\r\nStrong experience with Databricks ecosystem and Spark optimization\r\nExpertise in PostgreSQL performance tuning and schema design\r\nStrong attention to data quality, governance, and compliance\r\nExcellent communication skills, especially with non-technical stakeholders\r\nAbility to work in a highly regulated healthcare environment\r\nTypical Technology Stack\r\nLanguages: Python, SQL\r\nVersion Control: Git\r\nKPIs / Success Metrics\r\nReliability and performance of Databricks pipelines\r\nData quality and compliance adherence (HIPAA standards)\r\nTime-to-delivery for new data products\r\nQuery performance improvements in PostgreSQL and data warehouse systems\r\nJ-18808-Ljbffr","company":"Brooksource","rawCompany":"brooksource","city":"Louisville","state":"KY","isRemote":false,"isActive":false,"createdAt":"2026-07-16T01:47:51.702Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1243.00","title":"Database Architects","slug":"database-architects"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Data Engineer","description":"Our Fortune 50 Healthcare client is seeking a Senior Data Engineer to support our mission of improving the health and well-being of our members. This role will focus on building scalable, secure, data centric solutions and compliant data platforms that power analytics, clinical insights, and business decision-making across the enterprise.\r\nThe ideal candidate will have strong experience with cloud-based data platforms, Databricks, PostgreSQL , and healthcare data, with a passion for delivering high-quality, trusted data solutions in a regulated environment.\r\nKey Responsibilities\r\nDesign, develop, and scalable data pipeline solutions using Databricks (Spark) and cloud-native services\r\nBuild and optimize ETL/ELT workflows for ingesting structured and unstructured healthcare data (claims, clinical, provider, and member data)\r\nDevelop and maintain data models in PostgreSQL and enterprise data warehouses\r\nSupport Lakehouse architecture leveraging Databricks, Delta Lake , and cloud storage\r\nImprove performance, reliability, and cost-efficiency of data platforms\r\nWork with healthcare datasets, including producer/agent, broker, commission, and distribution data, ensuring proper ingestion, normalization, and optimization for analytics and reporting\r\nEnsure compliance with HIPAA, HITECH , and enterprise data governance policies\r\nImplement data security, encryption, masking, and access controls\r\nMaintain data lineage, auditability, and regulatory reporting readiness\r\nAdvanced Data Processing\r\nBuild real-time and batch pipelines for analytics and operational use cases\r\nDevelop data transformations using PySpark and SQL within Databricks\r\nLeverage PostgreSQL for transactional and analytical workloads where applicable\r\nIntegrate data from APIs, third-party vendors, and internal systems\r\nCollaboration & Stakeholder Engagement\r\nPartner with business stakeholders to support data-driven initiatives and member acquisition strategies\r\nTranslate insurance distribution, agent/producer , and marketing requirements into scalable, high-quality data solutions\r\nSupport downstream consumers, including Power BI, marketing analytics teams, and operational reporting stakeholders, by delivering curated, analytics-ready datasets\r\nTechnical Leadership\r\nLead design and architecture discussions for enterprise data solutions\r\nEstablish and enforce best practices in data engineering, testing, and CI/CD\r\nContribute to enterprise data strategy and platform modernization\r\nAI & Advanced Analytics (Databricks Genie)\r\nLeverage Databricks Genie (AI/BI capabilities) to enable natural language querying and democratize data access for business stakeholders\r\nDesign and optimize semantic layers and governed datasets that power Genie-driven insights with trusted, high-quality data\r\nCollaborate with stakeholders to translate business questions into AI-assisted analytics workflows using Databricks\r\nEnsure AI outputs are accurate, explainable, and compliant with healthcare data governance and HIPAA requirements\r\nLeverage large language models (LLMs), including Anthropic Claude , to enhance data exploration, automate insight generation, and support conversational analytics use cases\r\nIntegrate Genie capabilities with Delta Lake and curated data models to support near real-time insights and decision-making\r\nPartner with data scientists and analytics teams to enhance AI-driven use cases, including producer performance insights, marketing attribution, and member engagement analysis\r\nRequired Qualifications\r\nBachelor's or Master's degree in Computer Science, Engineering, or related field\r\n5–8+ years of experience in data engineering\r\nStrong programming in Python (PySpark) and advanced SQL\r\nHands-on experience with:\r\nDatabricks (core requirement)\r\nDistributed data processing frameworks (Apache Spark)\r\nExperience with cloud platforms (Azure preferred; AWS acceptable)\r\nProficiency in building and maintaining ETL/ELT pipelines\r\nStrong understanding of data modeling and warehousing concepts\r\nPreferred Qualifications\r\nExperience in healthcare or insurance industry (payer experience strongly preferred)\r\nFamiliarity with healthcare standards (e.g., FHIR, HL7 )\r\nExperience with\r\nDelta Lake / Lakehouse architecture\r\nKnowledge of DevOps and CI/CD pipelines (Azure DevOps, GitHub Actions)\r\nExperience supporting machine learning pipelines\r\nDeep understanding of data pipelines at scale\r\nStrong experience with Databricks ecosystem and Spark optimization\r\nExpertise in PostgreSQL performance tuning and schema design\r\nStrong attention to data quality, governance, and compliance\r\nExcellent communication skills, especially with non-technical stakeholders\r\nAbility to work in a highly regulated healthcare environment\r\nTypical Technology Stack\r\nLanguages: Python, SQL\r\nVersion Control: Git\r\nKPIs / Success Metrics\r\nReliability and performance of Databricks pipelines\r\nData quality and compliance adherence (HIPAA standards)\r\nTime-to-delivery for new data products\r\nQuery performance improvements in PostgreSQL and data warehouse systems\r\nJ-18808-Ljbffr","datePosted":"2026-07-16T01:47:51.702Z","dateModified":"2026-07-16T01:47:51.702Z","hiringOrganization":{"@type":"Organization","name":"Brooksource","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Louisville","addressRegion":"KY","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"76b6057a938913beee8511a8"},"url":"https://jobsearcher.com/jobs/76b6057a938913beee8511a8"}}