Engineer, Supercomputing & Distributed Systems
About KreaAt Krea, we are building next-generation AI creative tools.We're dedicated to making AI intuitive and controllable for creatives - our mission is to build tools that empower human creativity, not replace it. We believe AI is a new medium that allows us to express ourselves through various formats - text, images, video, sound, and even 3D. We're building better, smarter, and more controllable tools to harness this medium. We recently took this a step forward with the launch of Krea 2, our first foundation model, built completely from scratch for aesthetic diversity and stylistic control.We've raised over $83M and are backed by world-class investors such as a16z, Bain Capital, and Abstract. We work full-time and in-person at our waterfront office in San Francisco. We care about creativity: our team includes musicians, designers, visual artists, and engineers.Supercomputing / AI Infra at KreaWe build and operate the infrastructure for Krea's research and inference. Distributed training, 1000+ K8s GPU clusters, petabyte scale data pipelines, etc. We build a lot of this from scratch — custom distributed datastores, job orchestration systems, and streaming pipelines that replace tools like Kafka and Ray for modern AI workloads at scale.Example projects:Distributed data systemsDesign multi-stage pipelines that turn petabytes of raw data into clean, annotated datasetsRun classification models on billions of imagesDeploy and combine LLMs to caption massive multimedia dataGPU infrastructureManage distributed training and inference on 1000+ GPU Kubernetes clustersSolve orchestration and scaling for large-scale GPU job processingScale workloads and research between clusters in multiple datacentersDistributed trainingProfile and optimize dataloaders streaming thousands of images per secondProfile and debug InfiniBand networking on huge training runsBuild fault tolerance systems for large-scale pretrainingCollaborate with researchers on evolving RL infrastructureApplied ML pipelinesFind clean scenes in millions of videos using distributed shot-boundary detectionCustomize and train models to filter billions of images for questions like "is this a screenshot?"Build the systems that bridge raw cluster capacity and research outputWho we're looking forSystems people. If you've read a blog post about InfiniBand debugging or building a custom distributed database and thought "I want to do that" — this is that team.You'll spend your time working heavily with Python, Kubernetes, Torch, and data tools like DuckDB, Arrow, etc. It's OK if you don't have K8s or ML experience — the main thing we hire for is an intuition for distributed systems, and a great mental model of how systems interact and function under different conditions.Strong candidates may have experience with…Python, PyArrow, DuckDB, SQL, massive relational databases, PyTorch, Pandas, NumPy…KubernetesDesigning and implementing large-scale ETL systemsFundamental knowledge of containerization, operating systems, file-systems, and networkingDistributed systems designDistributed training systems (NCCL, InfiniBand, RDMA)Streaming and event processing systems (Kafka, Pulsar, or similar)PyTorch internals, custom dataloaders, and training infrastructureWhat we offerTeam: Work alongside a world-class team building the future of AI creative toolingImpact: Significant scope and company-wide impactCompetitive compensation: generous salary & equity packagesHealth & wellness: 100% health & 99% dental/vision insurance premiums covered for employees, health FSA accounts, & long-term disability coverageTime off: Flexible PTO policyFinancial planning: 401k with a 4% company-sponsored matchMeals in the office: breakfast, lunch, dinner - you name it, we'll cover itTransit: Ubers covered to & from the officeSponsorship: We're open to sponsoring international visas where we can (e.g., STEM OPT, OPT, H-1B, O-1, E-3). And more!Please note the above benefits & perks are for full-time employees