{"schemaVersion":"jobsearcher.job.v1","id":"cf9c41b56b682c002616087f","url":"https://jobsearcher.com/jobs/cf9c41b56b682c002616087f","canonicalUrl":"https://jobsearcher.com/jobs/cf9c41b56b682c002616087f","title":"Sr. Data Expert, Data Engineer (Clinical & Multimodal Data Integration)","description":"The Position\nA healthier future. It’s what drives us to innovate. To continuously advance science and ensure everyone has access to the healthcare they need today and for generations to come. Creating a world where we all have more time with the people we love. That’s what makes us Roche.\nAdvances in AI, data and computational sciences are transforming drug discovery and development. Roche’s Research and Early Development organizations at Genentech (gRED) and Pharma (pRED) have demonstrated how these technologies accelerate R&D, leveraging data and novel computational models to drive impact. Seamless data sharing and access to models across gRED and pRED are essential to maximising these opportunities. The Computational Sciences Center of Excellence (CS CoE) is a strategic, unified group whose goal is to harness the transformative power of data and Artificial Intelligence (AI) to assist our scientists in both pRED and gRED to deliver more innovative and life-changing medicines for patients worldwide.\nThe Computational Sciences Center of Excellence (CS CoE) brings together data, AI, and computational expertise to accelerate innovation across gRED and pRED. Within CS CoE, the Data and Digital Catalyst (DDC) organization leads the modernization of our data ecosystem, enabling scalable, data-driven science.\nThe Data Capability organization within DDC is responsible for establishing foundational data capabilities, including data connectivity, data compliance, scientific content management and data ingestion, curation, integration, and delivery. The team ensures that high-quality, well-structured datasets are available to power analytics, AI/ML, and scientific discovery across Research and Early Development.\nThe Opportunity:\nWe are seeking a Sr. Data Expert, Data Engineer to lead the integration and delivery of clinically anchored, multimodal scientific datasets spanning clinical, sequencing, imaging, proteomics, and other emerging data modalities.\nIn this role, you will:\nLead the integration and harmonization of clinical and multimodal scientific datasets, applying industry data standards and metadata frameworks to improve interoperability and scientific usability.\nOwn end-to-end data delivery by designing, validating, documenting, and delivering high-quality, analysis-ready datasets that support research, AI/ML, and computational biology initiatives.\nDevelop scalable data workflows that automate data ingestion, quality control, transformation, and metadata management across diverse scientific data sources.\nPartner with computational scientists, bioinformaticians, and data engineers to understand scientific requirements and translate them into scalable, reusable data solutions.\nDrive data quality and continuous improvement by implementing validation frameworks, metadata standards, and AI-assisted data curation practices that improve data discoverability and reuse.\nSupport emerging AI and foundation model initiatives by preparing interoperable, metadata-rich datasets optimized for downstream analytics and machine learning applications.\n\nWho You Are:\nYou have a PhD with 2+ years, a Master's degree with 3–5 years, or a Bachelor's degree with 5+ years of experience in Bioinformatics, Data Science, Biomedical Engineering, Computer Science, Clinical Sciences, or a related discipline, with experience working with clinical, biomedical, or scientific datasets.\nYou have hands-on experience integrating clinical data with one or more scientific modalities, including sequencing, imaging, proteomics, or other omics datasets, and understand clinical data models and longitudinal patient data.\nYou are proficient in Python (Pandas), SQL, and scientific data processing, with experience working with scientific data formats such as FASTQ, BAM/CRAM, VCF, DICOM, AnnData, or Parquet, and familiarity with cloud data platforms (AWS or GCP).\nYou have experience developing or supporting automated data pipelines using workflow orchestration tools such as Airflow, Nextflow, Snakemake, or Prefect, and are comfortable using Git for collaborative software development.\nYou are a collaborative problem solver with a strong focus on data quality, metadata management, and scientific reproducibility, and enjoy partnering with multidisciplinary teams to deliver scalable data solutions.\nPreferred Qualifications:\nExperience with clinical data standards such as CDISC (SDTM/ADaM), OMOP, or FHIR.\nExperience integrating multimodal datasets (e.g., clinical + genomics, imaging + transcriptomics, or multi-omics).\nFamiliarity with biomedical ontologies, controlled vocabularies, FAIR data principles, and metadata standards.\nExperience preparing scientific datasets for AI/ML workflows or foundation model development.\nExperience supporting translational research, biomarker discovery, or drug discovery programs.\nOnsite presence, on our South San Francisco campus, is expected for at least 3 days a week.\nRelocation benefits are not available for this job posting.\nThe expected salary range for this position based on the primary location of California is $119,800 - $222,400. Actual pay will be determined based on experience, qualifications, geographic location, and other job-related factors permitted by law. A discretionary annual bonus may be available based on individual and Company performance. This position also qualifies for the benefits detailed at the link provided below.\nBenefits\n#LI-JD1\n#ComputationCoE\nGenentech is an equal opportunity employer. It is our policy and practice to employ, promote, and otherwise treat any and all employees and applicants on the basis of merit, qualifications, and competence. The company's policy prohibits unlawful discrimination, including but not limited to, discrimination on the basis of Protected Veteran status, individuals with disabilities status, and consistent with all federal, state, or local laws.\nIf you have a disability and need an accommodation in relation to the online application process, please contact us by completing this form Accommodations for Applicants.","company":"Genentech","rawCompany":"genentech","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-18T12:46:46.124Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"},{"code":"15-2051.02","title":"Clinical Data Managers","slug":"clinical-data-managers"}],"industries":[{"code":"541714","title":"Research and Development in Biotechnology (except Nanobiotechnology)","slug":"research-and-development-in-biotechnology-except-nanobiotechnology"},{"code":"541715","title":"Research and Development in the Physical, Engineering, and Life Sciences (except Nanotechnology and Biotechnology)","slug":"research-and-development-in-the-physical-engineering-and-life-sciences-except-nanotechnology-and-biotechnology"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Sr. Data Expert, Data Engineer (Clinical & Multimodal Data Integration)","description":"The Position\nA healthier future. It’s what drives us to innovate. To continuously advance science and ensure everyone has access to the healthcare they need today and for generations to come. Creating a world where we all have more time with the people we love. That’s what makes us Roche.\nAdvances in AI, data and computational sciences are transforming drug discovery and development. Roche’s Research and Early Development organizations at Genentech (gRED) and Pharma (pRED) have demonstrated how these technologies accelerate R&D, leveraging data and novel computational models to drive impact. Seamless data sharing and access to models across gRED and pRED are essential to maximising these opportunities. The Computational Sciences Center of Excellence (CS CoE) is a strategic, unified group whose goal is to harness the transformative power of data and Artificial Intelligence (AI) to assist our scientists in both pRED and gRED to deliver more innovative and life-changing medicines for patients worldwide.\nThe Computational Sciences Center of Excellence (CS CoE) brings together data, AI, and computational expertise to accelerate innovation across gRED and pRED. Within CS CoE, the Data and Digital Catalyst (DDC) organization leads the modernization of our data ecosystem, enabling scalable, data-driven science.\nThe Data Capability organization within DDC is responsible for establishing foundational data capabilities, including data connectivity, data compliance, scientific content management and data ingestion, curation, integration, and delivery. The team ensures that high-quality, well-structured datasets are available to power analytics, AI/ML, and scientific discovery across Research and Early Development.\nThe Opportunity:\nWe are seeking a Sr. Data Expert, Data Engineer to lead the integration and delivery of clinically anchored, multimodal scientific datasets spanning clinical, sequencing, imaging, proteomics, and other emerging data modalities.\nIn this role, you will:\nLead the integration and harmonization of clinical and multimodal scientific datasets, applying industry data standards and metadata frameworks to improve interoperability and scientific usability.\nOwn end-to-end data delivery by designing, validating, documenting, and delivering high-quality, analysis-ready datasets that support research, AI/ML, and computational biology initiatives.\nDevelop scalable data workflows that automate data ingestion, quality control, transformation, and metadata management across diverse scientific data sources.\nPartner with computational scientists, bioinformaticians, and data engineers to understand scientific requirements and translate them into scalable, reusable data solutions.\nDrive data quality and continuous improvement by implementing validation frameworks, metadata standards, and AI-assisted data curation practices that improve data discoverability and reuse.\nSupport emerging AI and foundation model initiatives by preparing interoperable, metadata-rich datasets optimized for downstream analytics and machine learning applications.\n\nWho You Are:\nYou have a PhD with 2+ years, a Master's degree with 3–5 years, or a Bachelor's degree with 5+ years of experience in Bioinformatics, Data Science, Biomedical Engineering, Computer Science, Clinical Sciences, or a related discipline, with experience working with clinical, biomedical, or scientific datasets.\nYou have hands-on experience integrating clinical data with one or more scientific modalities, including sequencing, imaging, proteomics, or other omics datasets, and understand clinical data models and longitudinal patient data.\nYou are proficient in Python (Pandas), SQL, and scientific data processing, with experience working with scientific data formats such as FASTQ, BAM/CRAM, VCF, DICOM, AnnData, or Parquet, and familiarity with cloud data platforms (AWS or GCP).\nYou have experience developing or supporting automated data pipelines using workflow orchestration tools such as Airflow, Nextflow, Snakemake, or Prefect, and are comfortable using Git for collaborative software development.\nYou are a collaborative problem solver with a strong focus on data quality, metadata management, and scientific reproducibility, and enjoy partnering with multidisciplinary teams to deliver scalable data solutions.\nPreferred Qualifications:\nExperience with clinical data standards such as CDISC (SDTM/ADaM), OMOP, or FHIR.\nExperience integrating multimodal datasets (e.g., clinical + genomics, imaging + transcriptomics, or multi-omics).\nFamiliarity with biomedical ontologies, controlled vocabularies, FAIR data principles, and metadata standards.\nExperience preparing scientific datasets for AI/ML workflows or foundation model development.\nExperience supporting translational research, biomarker discovery, or drug discovery programs.\nOnsite presence, on our South San Francisco campus, is expected for at least 3 days a week.\nRelocation benefits are not available for this job posting.\nThe expected salary range for this position based on the primary location of California is $119,800 - $222,400. Actual pay will be determined based on experience, qualifications, geographic location, and other job-related factors permitted by law. A discretionary annual bonus may be available based on individual and Company performance. This position also qualifies for the benefits detailed at the link provided below.\nBenefits\n#LI-JD1\n#ComputationCoE\nGenentech is an equal opportunity employer. It is our policy and practice to employ, promote, and otherwise treat any and all employees and applicants on the basis of merit, qualifications, and competence. The company's policy prohibits unlawful discrimination, including but not limited to, discrimination on the basis of Protected Veteran status, individuals with disabilities status, and consistent with all federal, state, or local laws.\nIf you have a disability and need an accommodation in relation to the online application process, please contact us by completing this form Accommodations for Applicants.","datePosted":"2026-07-18T12:46:46.124Z","dateModified":"2026-07-18T12:46:46.124Z","hiringOrganization":{"@type":"Organization","name":"Genentech","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"cf9c41b56b682c002616087f"},"url":"https://jobsearcher.com/jobs/cf9c41b56b682c002616087f"}}