{"schemaVersion":"jobsearcher.job.v1","id":"bdc10b162c5b36601bd6cd36","url":"https://jobsearcher.com/jobs/bdc10b162c5b36601bd6cd36","canonicalUrl":"https://jobsearcher.com/jobs/bdc10b162c5b36601bd6cd36","title":"Machine Learning Engineer","description":"## About the RolePluralis Research is seeking an ML Training Platform Engineer to architect, build, and scale the foundational infrastructure behind its decentralized machine learning training platform. The role focuses on infrastructure orchestration, distributed compute, and services integration to enable continuous experimentation and large-scale model training.The successful candidate will bring deep experience in infrastructure and platform engineering, distributed systems, ML infrastructure, and reliability. This is a highly technical role involving multi-cloud environments, GPU workloads, distributed training, heterogeneous compute nodes, real-world networking conditions, and production-grade reliability systems.## Key Responsibilities* Design resource management systems that provision and orchestrate compute across AWS, GCP, and Azure.* Use infrastructure-as-code tools such as Pulumi and Terraform to manage multi-cloud infrastructure.* Build systems capable of dynamic scaling, state synchronization, and concurrent operations across hundreds of heterogeneous nodes.* Architect fault-tolerant infrastructure for distributed machine learning training.* Manage GPU clusters and NVIDIA runtime environments.* Implement S3 checkpointing and large-scale dataset management and streaming.* Build health monitoring and resilient retry strategies for distributed training workloads.* Design systems that simulate and handle real-world network conditions, including bandwidth shaping, latency injection, and packet loss.* Manage dynamic node churn across consumer and non-co-located infrastructure.* Optimize data flow between workers with heterogeneous network connectivity.* Support distributed training workflows, checkpointing, data sharding, and model versioning.* Build and maintain long-running job orchestration systems.* Work with decentralized networking technologies and real-world bandwidth constraints.* Develop reliable infrastructure and platform services for large-scale ML experimentation.* Apply observability and SRE practices to improve system reliability and performance.* Monitor and profile system performance and participate in incident response.* Collaborate with ML researchers and engineering teams to build scalable training infrastructure.## Required Qualifications* 5+ years of work experience is ideally expected.* Deep experience with infrastructure and platform engineering.* Production experience with infrastructure-as-code, including Pulumi, Terraform, or CloudFormation.* Experience managing multi-cloud deployments and infrastructure lifecycle orchestration.* Experience with self-healing systems.* Experience with Docker and Kubernetes, including EKS.* Experience with GPU workloads and heterogeneous clusters at scale.* Deep understanding of distributed training workflows and ML infrastructure.* Experience with checkpointing, data sharding, model versioning, and long-running job orchestration.* Strong Python engineering skills.* Experience with Python asyncio, concurrency, retry logic, cloud SDKs, and CLI tooling.* Hands-on experience with observability and SRE practices.* Experience with monitoring tools such as Prometheus and Grafana.* Experience with performance profiling and incident response.* Deep understanding of multi-cloud infrastructure and distributed training systems.## Preferred Qualifications* Startup experience with an emphasis on microservices orchestration.* Big tech engineering background.* Experience with decentralized networking, including P2P and NAT traversal.* Experience with traffic shaping and systems operating under real-world bandwidth constraints.* Experience working with consumer nodes and non-co-located infrastructure.* Experience with heterogeneous network connectivity.* Experience in large-scale ML training environments.## Skills & Competencies* Machine Learning Infrastructure* ML Training Platforms* Distributed Systems* Distributed Machine Learning* Multi-Cloud Infrastructure* AWS* Google Cloud Platform (GCP)* Microsoft Azure* Infrastructure as Code (IaC)* Pulumi* Terraform* CloudFormation* Docker* Kubernetes* Amazon EKS* GPU Workloads* NVIDIA Runtime* Python* asyncio* Concurrency* Cloud SDKs* CLI Tooling* Distributed Training* Checkpointing* Data Sharding* Model Versioning* Job Orchestration* Microservices Orchestration* Decentralized Networking* P2P Networking* NAT Traversal* Traffic Shaping* Network Latency* Packet Loss* Bandwidth Management* Observability* SRE* Prometheus* Grafana* Performance Profiling* Incident Response* Reliability Engineering* Infrastructure Orchestration## Education & Experience### Education* No specific education requirement is stated in the original job description.### Experience* Ideally 5+ years of professional work experience.* Deep infrastructure and platform engineering experience.* Production experience with infrastructure-as-code and multi-cloud deployments.* Experience with distributed systems and ML training infrastructure.* Strong Python engineering experience, including asynchronous programming and concurrency.* Experience with cloud infrastructure, Kubernetes, GPU workloads, monitoring, observability, and reliability practices.## Work Arrangement & Schedule* Onsite* Location: San Francisco, CA, US* Employment type and specific weekly hours are not stated in the original job description.* Weekend requirements are not stated in the original job description.## Compensation & Benefits* Compensation is not specified in the original job description.* The company is described as being backed by Union Square Ventures and other tier-1 investors.## Compliance / Additional InformationPluralis Research is focused on Protocol Learning and decentralized machine learning infrastructure. The company describes its approach as enabling collaborative AI model training by pooling compute from multiple participants and preventing any single party from controlling a model's full weights.The original job description states that Pluralis is seeking candidates who are passionate about its approach and who are comfortable working in a highly technical, research-oriented environment.","company":"Torentify","rawCompany":"torentify","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-31T08:47:07.803Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Machine Learning Engineer","description":"## About the RolePluralis Research is seeking an ML Training Platform Engineer to architect, build, and scale the foundational infrastructure behind its decentralized machine learning training platform. The role focuses on infrastructure orchestration, distributed compute, and services integration to enable continuous experimentation and large-scale model training.The successful candidate will bring deep experience in infrastructure and platform engineering, distributed systems, ML infrastructure, and reliability. This is a highly technical role involving multi-cloud environments, GPU workloads, distributed training, heterogeneous compute nodes, real-world networking conditions, and production-grade reliability systems.## Key Responsibilities* Design resource management systems that provision and orchestrate compute across AWS, GCP, and Azure.* Use infrastructure-as-code tools such as Pulumi and Terraform to manage multi-cloud infrastructure.* Build systems capable of dynamic scaling, state synchronization, and concurrent operations across hundreds of heterogeneous nodes.* Architect fault-tolerant infrastructure for distributed machine learning training.* Manage GPU clusters and NVIDIA runtime environments.* Implement S3 checkpointing and large-scale dataset management and streaming.* Build health monitoring and resilient retry strategies for distributed training workloads.* Design systems that simulate and handle real-world network conditions, including bandwidth shaping, latency injection, and packet loss.* Manage dynamic node churn across consumer and non-co-located infrastructure.* Optimize data flow between workers with heterogeneous network connectivity.* Support distributed training workflows, checkpointing, data sharding, and model versioning.* Build and maintain long-running job orchestration systems.* Work with decentralized networking technologies and real-world bandwidth constraints.* Develop reliable infrastructure and platform services for large-scale ML experimentation.* Apply observability and SRE practices to improve system reliability and performance.* Monitor and profile system performance and participate in incident response.* Collaborate with ML researchers and engineering teams to build scalable training infrastructure.## Required Qualifications* 5+ years of work experience is ideally expected.* Deep experience with infrastructure and platform engineering.* Production experience with infrastructure-as-code, including Pulumi, Terraform, or CloudFormation.* Experience managing multi-cloud deployments and infrastructure lifecycle orchestration.* Experience with self-healing systems.* Experience with Docker and Kubernetes, including EKS.* Experience with GPU workloads and heterogeneous clusters at scale.* Deep understanding of distributed training workflows and ML infrastructure.* Experience with checkpointing, data sharding, model versioning, and long-running job orchestration.* Strong Python engineering skills.* Experience with Python asyncio, concurrency, retry logic, cloud SDKs, and CLI tooling.* Hands-on experience with observability and SRE practices.* Experience with monitoring tools such as Prometheus and Grafana.* Experience with performance profiling and incident response.* Deep understanding of multi-cloud infrastructure and distributed training systems.## Preferred Qualifications* Startup experience with an emphasis on microservices orchestration.* Big tech engineering background.* Experience with decentralized networking, including P2P and NAT traversal.* Experience with traffic shaping and systems operating under real-world bandwidth constraints.* Experience working with consumer nodes and non-co-located infrastructure.* Experience with heterogeneous network connectivity.* Experience in large-scale ML training environments.## Skills & Competencies* Machine Learning Infrastructure* ML Training Platforms* Distributed Systems* Distributed Machine Learning* Multi-Cloud Infrastructure* AWS* Google Cloud Platform (GCP)* Microsoft Azure* Infrastructure as Code (IaC)* Pulumi* Terraform* CloudFormation* Docker* Kubernetes* Amazon EKS* GPU Workloads* NVIDIA Runtime* Python* asyncio* Concurrency* Cloud SDKs* CLI Tooling* Distributed Training* Checkpointing* Data Sharding* Model Versioning* Job Orchestration* Microservices Orchestration* Decentralized Networking* P2P Networking* NAT Traversal* Traffic Shaping* Network Latency* Packet Loss* Bandwidth Management* Observability* SRE* Prometheus* Grafana* Performance Profiling* Incident Response* Reliability Engineering* Infrastructure Orchestration## Education & Experience### Education* No specific education requirement is stated in the original job description.### Experience* Ideally 5+ years of professional work experience.* Deep infrastructure and platform engineering experience.* Production experience with infrastructure-as-code and multi-cloud deployments.* Experience with distributed systems and ML training infrastructure.* Strong Python engineering experience, including asynchronous programming and concurrency.* Experience with cloud infrastructure, Kubernetes, GPU workloads, monitoring, observability, and reliability practices.## Work Arrangement & Schedule* Onsite* Location: San Francisco, CA, US* Employment type and specific weekly hours are not stated in the original job description.* Weekend requirements are not stated in the original job description.## Compensation & Benefits* Compensation is not specified in the original job description.* The company is described as being backed by Union Square Ventures and other tier-1 investors.## Compliance / Additional InformationPluralis Research is focused on Protocol Learning and decentralized machine learning infrastructure. The company describes its approach as enabling collaborative AI model training by pooling compute from multiple participants and preventing any single party from controlling a model's full weights.The original job description states that Pluralis is seeking candidates who are passionate about its approach and who are comfortable working in a highly technical, research-oriented environment.","datePosted":"2026-08-31T08:47:07.803Z","dateModified":"2026-08-31T08:47:07.803Z","hiringOrganization":{"@type":"Organization","name":"Torentify","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"bdc10b162c5b36601bd6cd36"},"url":"https://jobsearcher.com/jobs/bdc10b162c5b36601bd6cd36"}}