{"schemaVersion":"jobsearcher.job.v1","id":"f994007cd9a825615be58cb4","url":"https://jobsearcher.com/jobs/f994007cd9a825615be58cb4","canonicalUrl":"https://jobsearcher.com/jobs/f994007cd9a825615be58cb4","title":"Data Engineering & Pipelining Lead","description":"We are excited to be expanding our Life Science Data Analytics Platform group in our Boston and Cambridge offices. We are looking for a Data Science & Pipelining Lead to join our Life Science Data Analytics Platform Group.\n\nArrayo assists top tier clients in life sciences in implementing effective Data Analytics strategies. We make sure that data assets are available and accessible for advanced analytics, and so that the inherent value of data assets can be realized more readily.\n\nThe Data Science & Pipelining Lead will be responsible for driving the socialization and utilization of technologies, algorithms, models, and methods for science driven data analytics R&D projects. As an Arrayo Team member, you will work to understand users’ requirements, and drive the definition, design, implementation and validation of cutting-edge pipelines and models used to process and analyze diverse sources of data.\n\nResponsibilities:\nDevelop data pipelines to extract, transform, and load data from various data sources in various forms.\n\nWork in collaboration with key scientific personnel to build, test, adapt, support, and validate pipelines with integration into production systems.\n\nManage the definition, design, implementation, and validation of data pipelines and models to analyze data from diverse sources.\n\nWrite custom scripts to extract data from unstructured/semi-structured sources.\n\nMake great use of advanced pipeline technologies incl. Prefect, Nextflow, Airflow, Cromwell, KNIME, Databricks, Luigi, petl, AWS Data Pipeline.\n\nLeverage big-data technologies for data processing, including Apache Spark, Kubernetes, Apache Pulsar, AWS (Lambda, S3, Athena).\n\nDeliver solutions in an efficient agile manner.\n\nContribute to many different projects in a dynamic, fast-moving environment.\n\nCollaboratively translate scientific and business questions into data and analytics requirements.\n\nDrive rapid prototyping for further implementation of analytical products.\n\nPartner with SMEs to translate modeling outputs into business language.\n\nWork with IT resources to enable appropriate data flow/data models.\n\nRequirements:\nB.S. in information systems, computer science, computer engineering or related field with 4+ years of experience working within bioinformatics, genomics, genetics, or other science related environments. M.S. or PhD Is preferred.\n\nKnowledge of a subset of analytical approaches (ex. machine learning, statistical analysis, predictive modeling, visual analytics).\n\nProficiency building, running, and monitoring pipelines on cloud computing environments.\n\nExperience in commonly used command-line NGS tools is a plus (BWA, SAMTools, Bowtie2, Picard, PINDEL, GATK, etc.).\n\nAbility to understand and communicate statistical measures for interrogating the quality of data manipulation preferred.\n\nDemonstrated ability to communicate efficiently and work effectively with a team of scientists and other engineers.\n\nExperience with SQL and modeling relational databases. PostgreSQL experience preferred.\n\nExperience using / designing web services and REST APIs.\n\nKnowledge of software development best practices\n\nExperience working in Cloud Computing environments (ex AWS, Azure, etc) is preferred.","company":"Arrayo","rawCompany":"arrayo","city":"East Boston","state":"MA","isRemote":false,"isActive":false,"createdAt":"2026-08-05T10:57:16.526Z","occupations":[{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"},{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"19-1029.01","title":"Bioinformatics Scientists","slug":"bioinformatics-scientists"}],"industries":[{"code":"541714","title":"Research and Development in Biotechnology (except Nanobiotechnology)","slug":"research-and-development-in-biotechnology-except-nanobiotechnology"},{"code":"541690","title":"Other Scientific and Technical Consulting Services","slug":"other-scientific-and-technical-consulting-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Data Engineering & Pipelining Lead","description":"We are excited to be expanding our Life Science Data Analytics Platform group in our Boston and Cambridge offices. We are looking for a Data Science & Pipelining Lead to join our Life Science Data Analytics Platform Group.\n\nArrayo assists top tier clients in life sciences in implementing effective Data Analytics strategies. We make sure that data assets are available and accessible for advanced analytics, and so that the inherent value of data assets can be realized more readily.\n\nThe Data Science & Pipelining Lead will be responsible for driving the socialization and utilization of technologies, algorithms, models, and methods for science driven data analytics R&D projects. As an Arrayo Team member, you will work to understand users’ requirements, and drive the definition, design, implementation and validation of cutting-edge pipelines and models used to process and analyze diverse sources of data.\n\nResponsibilities:\nDevelop data pipelines to extract, transform, and load data from various data sources in various forms.\n\nWork in collaboration with key scientific personnel to build, test, adapt, support, and validate pipelines with integration into production systems.\n\nManage the definition, design, implementation, and validation of data pipelines and models to analyze data from diverse sources.\n\nWrite custom scripts to extract data from unstructured/semi-structured sources.\n\nMake great use of advanced pipeline technologies incl. Prefect, Nextflow, Airflow, Cromwell, KNIME, Databricks, Luigi, petl, AWS Data Pipeline.\n\nLeverage big-data technologies for data processing, including Apache Spark, Kubernetes, Apache Pulsar, AWS (Lambda, S3, Athena).\n\nDeliver solutions in an efficient agile manner.\n\nContribute to many different projects in a dynamic, fast-moving environment.\n\nCollaboratively translate scientific and business questions into data and analytics requirements.\n\nDrive rapid prototyping for further implementation of analytical products.\n\nPartner with SMEs to translate modeling outputs into business language.\n\nWork with IT resources to enable appropriate data flow/data models.\n\nRequirements:\nB.S. in information systems, computer science, computer engineering or related field with 4+ years of experience working within bioinformatics, genomics, genetics, or other science related environments. M.S. or PhD Is preferred.\n\nKnowledge of a subset of analytical approaches (ex. machine learning, statistical analysis, predictive modeling, visual analytics).\n\nProficiency building, running, and monitoring pipelines on cloud computing environments.\n\nExperience in commonly used command-line NGS tools is a plus (BWA, SAMTools, Bowtie2, Picard, PINDEL, GATK, etc.).\n\nAbility to understand and communicate statistical measures for interrogating the quality of data manipulation preferred.\n\nDemonstrated ability to communicate efficiently and work effectively with a team of scientists and other engineers.\n\nExperience with SQL and modeling relational databases. PostgreSQL experience preferred.\n\nExperience using / designing web services and REST APIs.\n\nKnowledge of software development best practices\n\nExperience working in Cloud Computing environments (ex AWS, Azure, etc) is preferred.","datePosted":"2026-08-05T10:57:16.526Z","dateModified":"2026-08-05T10:57:16.526Z","hiringOrganization":{"@type":"Organization","name":"Arrayo","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"East Boston","addressRegion":"MA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"f994007cd9a825615be58cb4"},"url":"https://jobsearcher.com/jobs/f994007cd9a825615be58cb4"}}