JOBSEARCHER

Data Engineer, Databricks

We are a technology-led healthcare solutions provider. We are driven by our purpose to enable healthcare organizations to be future-ready. We offer accelerated, global growth opportunities for talent that is bold, industrious and nimble. With Indegene, you gain a unique career experience that celebrates entrepreneurship and is guided by passion, innovation, collaboration and empathy. To explore exciting opportunities at the convergence of healthcare and technology, check out www.careers.indegene.comWhat if we told you that you can move to an exciting role in an entrepreneurial organization without the usual risks associated with it?We understand that you are looking for growth and variety in your career at this point and we would love for you to join us in our journey and grow with us. At Indegene, our roles come with the excitement you require at this stage of your career with the reliability you seek. We hire the best and trust them from day 1 to deliver global impact, handle teams and be responsible for the outcomes while our leaders support and mentor you.We are a profitable rapidly growing global organization and are scouting for the best talent for this phase of growth. With us, you are at the intersection of two of the most exciting industries of healthcare and technology. We offer global opportunities with fast-track careers while you work with a team that is fueled by purpose. The combination of these will lead to a truly differentiated experience for you.If this excites you, then apply below.Job Description: Data Engineer – Databricks Genie, AWS, PySpark (Life Sciences Commercial Analytics) – Location – Raleigh, NC Position OverviewThe Data Engineer will be responsible for designing, developing, and optimizing large‑scale data pipelines and analytical solutions supporting Life Sciences Commercial Operations. This role requires deep expertise in Databricks (including Genie), AWS cloud services, and PySpark, with strong domain knowledge of pharmaceutical/biotech commercial datasets such as IQVIA, Symphony, DDD, NPA, Xponent, Claims, Specialty Pharmacy, and CRM/Sales data.The ideal candidate will partner closely with Commercial Analytics, Data Science, and Business Stakeholders to deliver high‑quality, scalable data products that enable sales insights, targeting, forecasting, incentive compensation, and field performance reporting.Key ResponsibilitiesDatabricks DevelopmentBuild, optimize, and maintain data pipelines using Databricks, Delta Lake, and Genie‑powered workflows.Implement LLM‑based automation and query generation using Databricks Genie for data exploration and business‑user self‑service.Develop notebooks, jobs, workflows, and ML‑ready datasets within the Databricks environment.AWS Cloud EngineeringDesign and manage data ingestion, storage, and processing using AWS services such as S3, Glue, Lambda, EMR, Redshift, and Athena.Implement secure, scalable architectures following AWS best practices (IAM, VPC, encryption, monitoring).PySpark & Big Data ProcessingBuild distributed data transformation pipelines using PySpark for high‑volume commercial datasets.Optimize PySpark jobs for performance, cost efficiency, and reliability.Implement unit testing, CI/CD, and code versioning using Git‑based workflows.Life Sciences Commercial Data ExpertiseIngest, harmonize, and model datasets including:IQVIA (Xponent, DDD, NPA, LAAD, NSP)Claims (medical, pharmacy)Specialty Pharmacy data feedsSales & CRM (Veeva, Salesforce)Roster, Territory, Alignment datasetsBuild commercial data marts supporting:Sales reporting and dashboardsTargeting and segmentationIncentive compensationForecasting and demand analyticsField performance insightsCross‑Functional CollaborationWork with Commercial Analytics, Data Science, IT, and Business teams to translate requirements into scalable data solutions.Ensure data quality, lineage, governance, and compliance with industry standards (HIPAA, GxP, SOC2).Required Qualifications5–10+ years of experience in Data Engineering or Big Data Analytics.Strong hands‑on experience with Databricks, Genie, Delta Lake, and MLflow.Advanced proficiency in PySpark, Python, SQL, and distributed computing.Deep experience with AWS (S3, Glue, Lambda, EMR, Redshift, IAM).Proven experience working with Life Sciences Commercial datasets.Strong understanding of data modeling, ETL/ELT, and cloud architecture.Experience with CI/CD, Git, DevOps, and automated workflow orchestration.Preferred QualificationsExperience with Databricks Unity Catalog and Lakehouse architecture.Familiarity with LLM‑based automation or AI‑assisted analytics.Experience supporting Commercial Operations, Sales Leadership, and Field Teams.Knowledge of incentive compensation methodologies and targeting algorithms.Experience with BI tools (Tableau, Power BI, Qlik).Role CompetenciesStrong analytical and problem‑solving skills.Excellent communication and documentation abilities.Ability to work in fast‑paced, cross‑functional environments.High attention to detail and commitment to data accuracy.EQUAL OPPORTUNITYIndegene is proud to be an Equal Employment Employer and is committed to the culture of Inclusion and Diversity. We do not discriminate on the basis of race, religion, sex, colour, age, national origin, pregnancy, sexual orientation, physical ability, or any other characteristics. All employment decisions, from hiring to separation, will be based on business requirements, candidate’s merit and qualification.We are an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, colour, religion, sex, national origin, gender identity, sexual orientation, disability status, protected veteran status, or any other characteristics.