{"schemaVersion":"jobsearcher.job.v1","id":"df816f2f97757ee46bfe1d71","url":"https://jobsearcher.com/jobs/df816f2f97757ee46bfe1d71","canonicalUrl":"https://jobsearcher.com/jobs/df816f2f97757ee46bfe1d71","title":"Software Engineer - Data Infrastructure","description":"About The RoleAs a Data Infrastructure Engineer in Research at Luma, you will play a critical role in building and scaling the data infrastructure that supports our cutting-edge multimodal AI systems. Your work will focus on developing high-throughput, large-scale data processing pipelines tailored for machine learning research and internal ML platform needs. You will collaborate closely with ML researchers and product teams to create reliable, efficient, and easy-to-use data infrastructure that empowers innovation and accelerates development. This role requires a strong foundation in distributed systems and data engineering, with an emphasis on supporting complex machine learning workflows rather than traditional product data infrastructure.ResponsibilitiesBuild and maintain scalable data infrastructure for high-throughput machine learning workflowsCollaborate with ML researchers and product teams to ensure data systems meet evolving needsDevelop and optimize large-scale data pipelines and batch processing jobsContribute to the architecture and implementation of reliable, high-performance data platformsIntegrate open-source tools and continuously improve data infrastructure through monitoring and tuningParticipate in cross-functional projects to improve data reliability, scalability, and operational excellenceSupport the evaluation and adoption of new programming languages and frameworks relevant to data infrastructureEngage in continuous improvement of data infrastructure through monitoring, troubleshooting, and performance tuningCollaborate with research & engineering teams to help define and refine best practices for data infrastructure developmentQualificationsProficiency in Python (or similar languages with willingness to learn Python) and experience with large-scale, high-throughput data infrastructureFamiliarity with distributed computing frameworks (e.g., Ray, Spark, Beam)Ability to design and optimize data pipelines for ML research and internal teamsStrong problem-solving skills and understanding of data engineering at scaleCollaborative, product-focused mindset; comfortable in fast-paced environmentsExperience sourcing, integrating, and optimizing data from diverse and large datasetsComfortable working in a fast-paced, product-focused environment with a strong execution mindsetOpen to candidates across seniority levels, from mid-level individual contributors to senior engineers and managers.Nice to havePrior experience working with complex data infrastructure or AI/ML platforms highly desirableExperience with open source data infrastructure projects is a plusExperience working in the robotics industry preferredCompensation Range: $170K - $360K","company":"Luma","rawCompany":"luma","city":"Redwood City","state":"CA","isRemote":false,"isActive":true,"createdAt":"2026-08-12T10:52:41.640Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Software Engineer - Data Infrastructure","description":"About The RoleAs a Data Infrastructure Engineer in Research at Luma, you will play a critical role in building and scaling the data infrastructure that supports our cutting-edge multimodal AI systems. Your work will focus on developing high-throughput, large-scale data processing pipelines tailored for machine learning research and internal ML platform needs. You will collaborate closely with ML researchers and product teams to create reliable, efficient, and easy-to-use data infrastructure that empowers innovation and accelerates development. This role requires a strong foundation in distributed systems and data engineering, with an emphasis on supporting complex machine learning workflows rather than traditional product data infrastructure.ResponsibilitiesBuild and maintain scalable data infrastructure for high-throughput machine learning workflowsCollaborate with ML researchers and product teams to ensure data systems meet evolving needsDevelop and optimize large-scale data pipelines and batch processing jobsContribute to the architecture and implementation of reliable, high-performance data platformsIntegrate open-source tools and continuously improve data infrastructure through monitoring and tuningParticipate in cross-functional projects to improve data reliability, scalability, and operational excellenceSupport the evaluation and adoption of new programming languages and frameworks relevant to data infrastructureEngage in continuous improvement of data infrastructure through monitoring, troubleshooting, and performance tuningCollaborate with research & engineering teams to help define and refine best practices for data infrastructure developmentQualificationsProficiency in Python (or similar languages with willingness to learn Python) and experience with large-scale, high-throughput data infrastructureFamiliarity with distributed computing frameworks (e.g., Ray, Spark, Beam)Ability to design and optimize data pipelines for ML research and internal teamsStrong problem-solving skills and understanding of data engineering at scaleCollaborative, product-focused mindset; comfortable in fast-paced environmentsExperience sourcing, integrating, and optimizing data from diverse and large datasetsComfortable working in a fast-paced, product-focused environment with a strong execution mindsetOpen to candidates across seniority levels, from mid-level individual contributors to senior engineers and managers.Nice to havePrior experience working with complex data infrastructure or AI/ML platforms highly desirableExperience with open source data infrastructure projects is a plusExperience working in the robotics industry preferredCompensation Range: $170K - $360K","datePosted":"2026-08-12T10:52:41.640Z","dateModified":"2026-08-12T10:52:41.640Z","hiringOrganization":{"@type":"Organization","name":"Luma","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Redwood City","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"df816f2f97757ee46bfe1d71"},"url":"https://jobsearcher.com/jobs/df816f2f97757ee46bfe1d71"}}