Software Engineer, Data Infrastructure
Overview
In this role you will build and maintain the high-performance data layer powering AI training and evaluation at scale. You will work closely with researchers and engineers to tackle storage and data processing challenges. You’ll shape how petabyte-scale data is stored, accessed, and transformed to fuel cutting-edge models. This is a hands-on opportunity to impact the data backbone of world-class AI workloads and join a collaborative, fast-paced team.
Compensation / Benefitsopen and inclusive cultureweekly lunch stipend and snackshealth and dental benefits with mental health budget6 weeks vacationremote-flexible work locations with global officesco-working stipend
ResponsibilitiesBuild and maintain the high-performance data layer for training and evaluation jobsAddress petabyte-scale storage and networking performance challengesCollaborate daily with researchers and engineers to align data infrastructure with modeling needs
Key requirements4+ years of experience in data storage infrastructureStrong command of PythonKubernetes experience on the storage side (Persistent Volumes, CSI drivers)Ability to transform unstructured data into performant datasets across S3, GCS, and POSIXExperience with distributed data processing frameworks such as Apache Beam, Spark, or Flinkcollaboration with cross-functional teamscuriosity and willingness to work in the weedsgenuine excitement about AIPythonKubernetes storage (Persistent Volumes, CSI drivers)S3