{"schemaVersion":"jobsearcher.job.v1","id":"485e833e3aea81aa28999857","url":"https://jobsearcher.com/jobs/485e833e3aea81aa28999857","canonicalUrl":"https://jobsearcher.com/jobs/485e833e3aea81aa28999857","title":"Engineer, Supercomputing & Distributed Systems","description":"About Krea\nAt Krea, we are building next-generation AI creative tools. We are dedicated to making AI intuitive and controllable for creatives. Our mission is to build tools that empower human creativity, not replace it. We believe AI is a new medium that allows us to express ourselves through various formats—text, images, video, sound, and even 3D. We're building better, smarter, and more controllable tools to harness this medium.\n\nSupercomputing / AI Infra at Krea\nWe build and operate the infrastructure for Krea's research and inference. Distributed training, 1000+ K8s GPU clusters, petabyte scale data pipelines, etc. We build a lot of this from scratch — custom distributed datastores, job orchestration systems, and streaming pipelines that replace tools like Kafka and Ray for modern AI workloads at scale.\n\nExample Projects\nDistributed data systems\n\nDesign multi-stage pipelines that turn petabytes of raw data into clean, annotated datasets\n\nRun classification models on billions of images\n\nDeploy and combine LLMs to caption massive multimedia data\n\nGPU infrastructure\n\nManage distributed training and inference on 1000+ GPU Kubernetes clusters\n\nSolve orchestration and scaling for large-scale GPU job processing\n\nScale workloads and research between clusters in multiple datacenters\n\nDistributed training\n\nProfile and optimize dataloaders streaming thousands of images per second\n\nProfile and debug InfiniBand networking on huge training runs\n\nBuild fault tolerance systems for large-scale pretraining\n\nCollaborate with researchers on evolving RL infrastructure\n\nApplied ML pipelines\n\nFind clean scenes in millions of videos using distributed shot-boundary detection\n\nCustomize and train models to filter billions of images for questions like \"is this a screenshot?\"\n\nBuild the systems that bridge raw cluster capacity and research output\n\nWho We're Looking For\nSystems people. If you've read a blog post about InfiniBand debugging or building a custom distributed database and thought \"I want to do that\" — this is that team. You'll spend your time working heavily with Python, Kubernetes, Torch, and data tools like DuckDB, Arrow, etc. It's OK if you don't have K8s or ML experience — the main thing we hire for is an intuition for distributed systems, and a great mental model of how systems interact and function under different conditions.\n\nStrong Candidates May Have Experience With\n\nPython, PyArrow, DuckDB, SQL, massive relational databases, PyTorch, Pandas, NumPy\n\nKubernetes\n\nDesigning and implementing large-scale ETL systems\n\nFundamental knowledge of containerization, operating systems, file-systems, and networking\n\nDistributed systems design\n\nDistributed training systems (NCCL, InfiniBand, RDMA)\n\nStreaming and event processing systems (Kafka, Pulsar, or similar)\n\nPyTorch internals, custom dataloaders, and training infrastructure\n\n#J-18808-Ljbffr","company":"Krea","rawCompany":"krea","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-07-16T03:29:23.524Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Engineer, Supercomputing & Distributed Systems","description":"About Krea\nAt Krea, we are building next-generation AI creative tools. We are dedicated to making AI intuitive and controllable for creatives. Our mission is to build tools that empower human creativity, not replace it. We believe AI is a new medium that allows us to express ourselves through various formats—text, images, video, sound, and even 3D. We're building better, smarter, and more controllable tools to harness this medium.\n\nSupercomputing / AI Infra at Krea\nWe build and operate the infrastructure for Krea's research and inference. Distributed training, 1000+ K8s GPU clusters, petabyte scale data pipelines, etc. We build a lot of this from scratch — custom distributed datastores, job orchestration systems, and streaming pipelines that replace tools like Kafka and Ray for modern AI workloads at scale.\n\nExample Projects\nDistributed data systems\n\nDesign multi-stage pipelines that turn petabytes of raw data into clean, annotated datasets\n\nRun classification models on billions of images\n\nDeploy and combine LLMs to caption massive multimedia data\n\nGPU infrastructure\n\nManage distributed training and inference on 1000+ GPU Kubernetes clusters\n\nSolve orchestration and scaling for large-scale GPU job processing\n\nScale workloads and research between clusters in multiple datacenters\n\nDistributed training\n\nProfile and optimize dataloaders streaming thousands of images per second\n\nProfile and debug InfiniBand networking on huge training runs\n\nBuild fault tolerance systems for large-scale pretraining\n\nCollaborate with researchers on evolving RL infrastructure\n\nApplied ML pipelines\n\nFind clean scenes in millions of videos using distributed shot-boundary detection\n\nCustomize and train models to filter billions of images for questions like \"is this a screenshot?\"\n\nBuild the systems that bridge raw cluster capacity and research output\n\nWho We're Looking For\nSystems people. If you've read a blog post about InfiniBand debugging or building a custom distributed database and thought \"I want to do that\" — this is that team. You'll spend your time working heavily with Python, Kubernetes, Torch, and data tools like DuckDB, Arrow, etc. It's OK if you don't have K8s or ML experience — the main thing we hire for is an intuition for distributed systems, and a great mental model of how systems interact and function under different conditions.\n\nStrong Candidates May Have Experience With\n\nPython, PyArrow, DuckDB, SQL, massive relational databases, PyTorch, Pandas, NumPy\n\nKubernetes\n\nDesigning and implementing large-scale ETL systems\n\nFundamental knowledge of containerization, operating systems, file-systems, and networking\n\nDistributed systems design\n\nDistributed training systems (NCCL, InfiniBand, RDMA)\n\nStreaming and event processing systems (Kafka, Pulsar, or similar)\n\nPyTorch internals, custom dataloaders, and training infrastructure\n\n#J-18808-Ljbffr","datePosted":"2026-07-16T03:29:23.524Z","dateModified":"2026-07-16T03:29:23.524Z","hiringOrganization":{"@type":"Organization","name":"Krea","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"485e833e3aea81aa28999857"},"url":"https://jobsearcher.com/jobs/485e833e3aea81aa28999857"}}