{"schemaVersion":"jobsearcher.job.v1","id":"a87652a87b201ad4ee595e22","url":"https://jobsearcher.com/jobs/a87652a87b201ad4ee595e22","canonicalUrl":"https://jobsearcher.com/jobs/a87652a87b201ad4ee595e22","title":"Data Engineer","description":"Title : Data EngineerDuration : 12 - months contract (high possibility of extension)Location : Cambridge MAHours per week : 40 Hours Job RequirementWe are seeking a highly skilled Data Engineer with expertise in Databricks, Seqera Platform (Nextflow Tower), cloud data engineering, and scientific data workflows to support our Discovery R&D and data science initiatives. The ideal candidate will design, develop, and maintain scalable data platforms and automated bioinformatics/data science pipelines that enable researchers and scientists to efficiently process, analyze, and access large-scale scientific and business datasets. This role requires strong experience in cloud-native architectures, data engineering best practices, workflow orchestration, and collaboration with cross-functional teams including scientists, bioinformaticians, data scientists, and IT infrastructure teams.About the RoleThe role involves designing, developing, and maintaining scalable data platforms and automated bioinformatics/data science pipelines.ResponsibilitiesData Engineering & Platform DevelopmentDesign, develop, and maintain scalable data pipelines using Databricks, Apache Spark, and cloud-native technologies.Build and optimize ETL/ELT processes for structured, semi-structured, and unstructured data.Develop data ingestion frameworks for research, laboratory, clinical, and external scientific datasets.Implement data quality, validation, monitoring, and governance processes.Support enterprise data lakehouse architecture and data platform modernization initiatives.Databricks Administration & DevelopmentDevelop and maintain Databricks notebooks, workflows, Delta Live Tables, and Jobs.Create optimized Spark-based transformations and data processing solutions.Implement Medallion Architecture (Bronze, Silver, Gold) for data lifecycle management.Manage Delta Lake environments and optimize performance, scalability, and cost.Integrate Databricks with cloud-native services and enterprise applications.Seqera Platform & Scientific Workflow ManagementDeploy, configure, and support Seqera Platform (formerly Nextflow Tower).Develop and maintain Nextflow pipelines for bioinformatics, genomics, imaging, AI/ML, and scientific computing workloads.Integrate Seqera workflows with AWS cloud infrastructure and compute environments.Support containerized workflows using Docker and Kubernetes technologies.Enable reproducible, scalable, and compliant scientific data processing workflows.Cloud EngineeringDesign and implement cloud-based data solutions in AWS.Manage cloud storage solutions including S3 and data lifecycle policies.Develop Infrastructure-as-Code solutions using Terraform or CloudFormation.Implement security controls and access management following enterprise IT standards.Partner with data scientists, researchers, bioinformaticians, and business stakeholders to understand data requirements.Provide technical guidance on data engineering best practices and workflow automation.Troubleshoot pipeline failures, performance issues, and workflow bottlenecks.Contribute to platform roadmaps and continuous improvement initiatives.Maintain technical documentation, SOPs, and knowledge articles.QualificationsEducationBachelor's degree in Computer Science, Information Technology, Data Engineering, Bioinformatics, or a related technical field.Master's degree preferred.Experience5+ years of experience in data engineering, cloud engineering, or analytics platform development.3+ years of hands-on experience with Databricks and Apache Spark.2+ years of experience with Seqera Platform (Nextflow Tower) and Nextflow workflows.Experience supporting scientific research, life sciences, pharmaceutical, biotech, or healthcare environments preferred.Required SkillsDatabricks & Data Engineering, Databricks Lakehouse PlatformApache Spark (PySpark, Spark SQL)Delta Lake, Delta Live Tables (DLT)Databricks WorkflowsUnity CatalogSQL and PythonSeqera & Scientific Computing; Seqera Platform / Nextflow TowerNextflow pipeline developmentBioinformatics workflow automationDocker and container technologiesKubernetes orchestrationHigh-performance computing environmentsCloud Technologies, AWS (required)S3, IAM, EC2, VPC, LambdaTerraform or CloudFormationCloud monitoring and logging toolsData TechnologiesData Lake and Lakehouse architecturesETL/ELT frameworksData modelingData cataloging and governanceAPI integrationsData quality frameworksDevOps & AutomationGitHub/GitLabCI/CD pipelinesJenkins, GitHub Actions, or similar toolsInfrastructure as CodeAgile and DevOps methodologies","company":"Tech Observer","rawCompany":"tech observer","city":"Somerville","state":"MA","isRemote":false,"isActive":false,"createdAt":"2026-08-19T13:26:56.739Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Data Engineer","description":"Title : Data EngineerDuration : 12 - months contract (high possibility of extension)Location : Cambridge MAHours per week : 40 Hours Job RequirementWe are seeking a highly skilled Data Engineer with expertise in Databricks, Seqera Platform (Nextflow Tower), cloud data engineering, and scientific data workflows to support our Discovery R&D and data science initiatives. The ideal candidate will design, develop, and maintain scalable data platforms and automated bioinformatics/data science pipelines that enable researchers and scientists to efficiently process, analyze, and access large-scale scientific and business datasets. This role requires strong experience in cloud-native architectures, data engineering best practices, workflow orchestration, and collaboration with cross-functional teams including scientists, bioinformaticians, data scientists, and IT infrastructure teams.About the RoleThe role involves designing, developing, and maintaining scalable data platforms and automated bioinformatics/data science pipelines.ResponsibilitiesData Engineering & Platform DevelopmentDesign, develop, and maintain scalable data pipelines using Databricks, Apache Spark, and cloud-native technologies.Build and optimize ETL/ELT processes for structured, semi-structured, and unstructured data.Develop data ingestion frameworks for research, laboratory, clinical, and external scientific datasets.Implement data quality, validation, monitoring, and governance processes.Support enterprise data lakehouse architecture and data platform modernization initiatives.Databricks Administration & DevelopmentDevelop and maintain Databricks notebooks, workflows, Delta Live Tables, and Jobs.Create optimized Spark-based transformations and data processing solutions.Implement Medallion Architecture (Bronze, Silver, Gold) for data lifecycle management.Manage Delta Lake environments and optimize performance, scalability, and cost.Integrate Databricks with cloud-native services and enterprise applications.Seqera Platform & Scientific Workflow ManagementDeploy, configure, and support Seqera Platform (formerly Nextflow Tower).Develop and maintain Nextflow pipelines for bioinformatics, genomics, imaging, AI/ML, and scientific computing workloads.Integrate Seqera workflows with AWS cloud infrastructure and compute environments.Support containerized workflows using Docker and Kubernetes technologies.Enable reproducible, scalable, and compliant scientific data processing workflows.Cloud EngineeringDesign and implement cloud-based data solutions in AWS.Manage cloud storage solutions including S3 and data lifecycle policies.Develop Infrastructure-as-Code solutions using Terraform or CloudFormation.Implement security controls and access management following enterprise IT standards.Partner with data scientists, researchers, bioinformaticians, and business stakeholders to understand data requirements.Provide technical guidance on data engineering best practices and workflow automation.Troubleshoot pipeline failures, performance issues, and workflow bottlenecks.Contribute to platform roadmaps and continuous improvement initiatives.Maintain technical documentation, SOPs, and knowledge articles.QualificationsEducationBachelor's degree in Computer Science, Information Technology, Data Engineering, Bioinformatics, or a related technical field.Master's degree preferred.Experience5+ years of experience in data engineering, cloud engineering, or analytics platform development.3+ years of hands-on experience with Databricks and Apache Spark.2+ years of experience with Seqera Platform (Nextflow Tower) and Nextflow workflows.Experience supporting scientific research, life sciences, pharmaceutical, biotech, or healthcare environments preferred.Required SkillsDatabricks & Data Engineering, Databricks Lakehouse PlatformApache Spark (PySpark, Spark SQL)Delta Lake, Delta Live Tables (DLT)Databricks WorkflowsUnity CatalogSQL and PythonSeqera & Scientific Computing; Seqera Platform / Nextflow TowerNextflow pipeline developmentBioinformatics workflow automationDocker and container technologiesKubernetes orchestrationHigh-performance computing environmentsCloud Technologies, AWS (required)S3, IAM, EC2, VPC, LambdaTerraform or CloudFormationCloud monitoring and logging toolsData TechnologiesData Lake and Lakehouse architecturesETL/ELT frameworksData modelingData cataloging and governanceAPI integrationsData quality frameworksDevOps & AutomationGitHub/GitLabCI/CD pipelinesJenkins, GitHub Actions, or similar toolsInfrastructure as CodeAgile and DevOps methodologies","datePosted":"2026-08-19T13:26:56.739Z","dateModified":"2026-08-19T13:26:56.739Z","hiringOrganization":{"@type":"Organization","name":"Tech Observer","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Somerville","addressRegion":"MA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"a87652a87b201ad4ee595e22"},"url":"https://jobsearcher.com/jobs/a87652a87b201ad4ee595e22"}}