{"schemaVersion":"jobsearcher.job.v1","id":"2a0605d45fca64d3c11d6502","url":"https://jobsearcher.com/jobs/2a0605d45fca64d3c11d6502","canonicalUrl":"https://jobsearcher.com/jobs/2a0605d45fca64d3c11d6502","title":"Founding Data Engineer","description":"A team I'm working with is building AI systems to improve decision-making in drug development, specifically around predicting safety risks earlier in the process. One of the biggest challenges in the space is that teams still rely on a mix of fragmented experimental data, prior research, and manual judgment to decide whether something is safe to advance. The focus here is on bringing that into a more unified, data-driven approach. They've put together large, complex datasets across a range of scientific domains and are using those to power models that help researchers better understand risk and underlying mechanisms. They're looking to add a Data Engineer to help build out the core data layer. The role is pretty end-to-end—designing and scaling pipelines, building internal APIs and tooling, and turning messy raw data into clean, structured datasets that can be used for modeling and analysis. You'd also work on automating data ingestion from different sources, improving data quality and reliability, and supporting model-related workflows. There's some exposure to LLM-driven workflows as well, particularly around extracting and structuring information. It's an early, high-impact role where the data systems you build will directly influence how the broader team operates and scales.\r\nWhat you'll do\r\nOwn and scale core data infrastructure across research, ML, and product systems\r\nBuild and maintain data pipelines for ingesting, processing, validating, and serving large, complex datasets from multiple sources\r\nDevelop internal platforms that connect experimental workflows, data capture, processing, and downstream usage\r\nTransform raw, messy data into clean, versioned, and ML-ready datasets\r\nDesign and build APIs and data tools that make it easy for different teams to access and use data\r\nWork on systems that automate data ingestion, cleaning, normalization, and structuring across a variety of inputs\r\nSupport and scale model-related workflows, including batch processing and inference pipelines\r\nImplement data quality systems (validation, testing, monitoring, lineage, observability) to ensure reliability\r\nPartner closely with domain experts to understand workflows and translate them into scalable infrastructure\r\nHelp support internal and external data delivery, including datasets, outputs, and derived insights\r\nBuild systems that improve speed and efficiency across the organization\r\nWhat they're looking for\r\nExperience building and maintaining large-scale data platforms used by multiple teams\r\nComfortable working with messy, unstructured, or heterogeneous datasets\r\nAbility to operate across backend engineering, data systems, and infrastructure\r\nStrong focus on data quality, correctness, and reproducibility\r\nExperience working with cross-functional teams and translating real-world needs into technical solutions\r\nFamiliarity with AI/ML or LLM-related data workflows is a plus\r\nComfortable operating in ambiguous, fast-moving environments\r\nInterested in owning critical systems and having broad impact\r\nCuriosity across technical and applied problem spaces\r\nTechnical background (nice to have)\r\nStrong Python and SQL fundamentals\r\nExperience with distributed systems or large-scale data processing frameworks\r\nCloud infrastructure (any major provider)\r\nInfrastructure as code and modern deployment practices\r\nData platforms (warehouses, lakes, or similar storage systems)\r\nWorkflow orchestration tools\r\nExperience building APIs or internal data tooling\r\nExposure to ML or LLM-related infrastructure\r\nExperience handling large-scale datasets\r\nWhat tends to work well in this environment\r\nMoves quickly and takes ownership\r\nStrong engineering judgment and attention to detail\r\nProactively identifies problems and builds solutions\r\nBalances speed with reliability\r\nThink(s) in systems rather than one-off fixes\r\nComfortable navigating complexity and changing requirements\r\nEnjoys working with technical and research-oriented teams\r\nFocuses on building scalable, long-term solutions\r\nMotivated by high-impact work in an early-stage environment\r\nJ-18808-Ljbffr","company":"Glocomms","rawCompany":"glocomms","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-16T01:56:17.955Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Founding Data Engineer","description":"A team I'm working with is building AI systems to improve decision-making in drug development, specifically around predicting safety risks earlier in the process. One of the biggest challenges in the space is that teams still rely on a mix of fragmented experimental data, prior research, and manual judgment to decide whether something is safe to advance. The focus here is on bringing that into a more unified, data-driven approach. They've put together large, complex datasets across a range of scientific domains and are using those to power models that help researchers better understand risk and underlying mechanisms. They're looking to add a Data Engineer to help build out the core data layer. The role is pretty end-to-end—designing and scaling pipelines, building internal APIs and tooling, and turning messy raw data into clean, structured datasets that can be used for modeling and analysis. You'd also work on automating data ingestion from different sources, improving data quality and reliability, and supporting model-related workflows. There's some exposure to LLM-driven workflows as well, particularly around extracting and structuring information. It's an early, high-impact role where the data systems you build will directly influence how the broader team operates and scales.\r\nWhat you'll do\r\nOwn and scale core data infrastructure across research, ML, and product systems\r\nBuild and maintain data pipelines for ingesting, processing, validating, and serving large, complex datasets from multiple sources\r\nDevelop internal platforms that connect experimental workflows, data capture, processing, and downstream usage\r\nTransform raw, messy data into clean, versioned, and ML-ready datasets\r\nDesign and build APIs and data tools that make it easy for different teams to access and use data\r\nWork on systems that automate data ingestion, cleaning, normalization, and structuring across a variety of inputs\r\nSupport and scale model-related workflows, including batch processing and inference pipelines\r\nImplement data quality systems (validation, testing, monitoring, lineage, observability) to ensure reliability\r\nPartner closely with domain experts to understand workflows and translate them into scalable infrastructure\r\nHelp support internal and external data delivery, including datasets, outputs, and derived insights\r\nBuild systems that improve speed and efficiency across the organization\r\nWhat they're looking for\r\nExperience building and maintaining large-scale data platforms used by multiple teams\r\nComfortable working with messy, unstructured, or heterogeneous datasets\r\nAbility to operate across backend engineering, data systems, and infrastructure\r\nStrong focus on data quality, correctness, and reproducibility\r\nExperience working with cross-functional teams and translating real-world needs into technical solutions\r\nFamiliarity with AI/ML or LLM-related data workflows is a plus\r\nComfortable operating in ambiguous, fast-moving environments\r\nInterested in owning critical systems and having broad impact\r\nCuriosity across technical and applied problem spaces\r\nTechnical background (nice to have)\r\nStrong Python and SQL fundamentals\r\nExperience with distributed systems or large-scale data processing frameworks\r\nCloud infrastructure (any major provider)\r\nInfrastructure as code and modern deployment practices\r\nData platforms (warehouses, lakes, or similar storage systems)\r\nWorkflow orchestration tools\r\nExperience building APIs or internal data tooling\r\nExposure to ML or LLM-related infrastructure\r\nExperience handling large-scale datasets\r\nWhat tends to work well in this environment\r\nMoves quickly and takes ownership\r\nStrong engineering judgment and attention to detail\r\nProactively identifies problems and builds solutions\r\nBalances speed with reliability\r\nThink(s) in systems rather than one-off fixes\r\nComfortable navigating complexity and changing requirements\r\nEnjoys working with technical and research-oriented teams\r\nFocuses on building scalable, long-term solutions\r\nMotivated by high-impact work in an early-stage environment\r\nJ-18808-Ljbffr","datePosted":"2026-07-16T01:56:17.955Z","dateModified":"2026-07-16T01:56:17.955Z","hiringOrganization":{"@type":"Organization","name":"Glocomms","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"2a0605d45fca64d3c11d6502"},"url":"https://jobsearcher.com/jobs/2a0605d45fca64d3c11d6502"}}