{"schemaVersion":"jobsearcher.job.v1","id":"86829f894e91b79c4cba7ab5","url":"https://jobsearcher.com/jobs/86829f894e91b79c4cba7ab5","canonicalUrl":"https://jobsearcher.com/jobs/86829f894e91b79c4cba7ab5","title":"Staff Replication Development Engineer","description":"DDN is seeking a Staff Replication Development Engineer to lead the design and development of the replication engine for the Infinia AI Data Platform. This role focuses on building enterprise-grade asynchronous replication capabilities that enable reliable and secure disaster recovery for large-scale data systems.\nYou will work on developing high-performance replication pipelines, efficient data synchronization mechanisms, and secure data transfer systems. This role requires deep expertise in distributed systems and strong technical leadership to deliver a scalable and resilient replication foundation.\n\nKey Responsibilities\nDesign and develop multi-threaded asynchronous replication systems with parallel streaming capabilities\nBuild object-level delta replication with checkpointing and resume functionality\nDevelop replication engines supporting bucket/share-level replication controls\nImplement secure data transfer mechanisms using TLS 1.3 with mutual authentication\nEnsure end-to-end data integrity through checksum validation and verification pipelines\nDesign and implement manual failover workflows for disaster recovery scenarios\nBuild and maintain REST APIs for replication configuration, control, and automation\nDevelop metadata tracking and change detection systems to enable efficient replication\nImplement RPO visibility, alerting, and operational insights for replication status\nContribute to monitoring dashboards focused on replication health and performance\nEnsure systems are designed for high availability, fault tolerance, and scalability\nPartner with QA teams to drive performance, resiliency, and scale validation\nCollaborate with backend, security, and platform teams to deliver end-to-end replication workflows\nParticipate in debugging, production issue resolution, and continuous improvement of replication reliability\nProvide technical leadership, architectural guidance, and mentorship to the engineering team\n\nRequired Qualifications\n8+ years of experience in distributed systems, storage systems, or backend software engineering\nStrong programming skills in one or more languages: C++, Go, Java, or Rust\nExperience designing and building data replication systems, data pipelines, or distributed data services\nDeep understanding of distributed systems concepts (consistency, availability, scalability, fault tolerance)\nStrong expertise in multi-threading, concurrency, and parallel processing\nKnowledge of networking protocols and secure communication (TCP/IP, HTTP/HTTPS, TLS)\nExperience implementing data integrity mechanisms (checksums, validation, consistency checks)\nExperience designing and building REST APIs and service-based architectures\nFamiliarity with checkpointing, failure recovery, and retry mechanisms in distributed systems\nBasic understanding of observability concepts (metrics, logging, alerting)\nStrong debugging, problem-solving, and system design skills\n\nPreferred Qualifications\nExperience with asynchronous replication, disaster recovery (DR), or backup systems\nFamiliarity with object storage or large-scale data storage systems\nKnowledge of delta encoding, change data capture, or incremental data synchronization techniques\nExperience building high-throughput, low-latency data movement systems\nExposure to security practices including mutual TLS, encryption, and authentication\nExperience working on enterprise-scale data platforms or storage products\nFamiliarity with performance optimization and large-scale system tuning","company":"Ddn","rawCompany":"ddn","city":"San Jose","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-05T16:13:01.393Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Staff Replication Development Engineer","description":"DDN is seeking a Staff Replication Development Engineer to lead the design and development of the replication engine for the Infinia AI Data Platform. This role focuses on building enterprise-grade asynchronous replication capabilities that enable reliable and secure disaster recovery for large-scale data systems.\nYou will work on developing high-performance replication pipelines, efficient data synchronization mechanisms, and secure data transfer systems. This role requires deep expertise in distributed systems and strong technical leadership to deliver a scalable and resilient replication foundation.\n\nKey Responsibilities\nDesign and develop multi-threaded asynchronous replication systems with parallel streaming capabilities\nBuild object-level delta replication with checkpointing and resume functionality\nDevelop replication engines supporting bucket/share-level replication controls\nImplement secure data transfer mechanisms using TLS 1.3 with mutual authentication\nEnsure end-to-end data integrity through checksum validation and verification pipelines\nDesign and implement manual failover workflows for disaster recovery scenarios\nBuild and maintain REST APIs for replication configuration, control, and automation\nDevelop metadata tracking and change detection systems to enable efficient replication\nImplement RPO visibility, alerting, and operational insights for replication status\nContribute to monitoring dashboards focused on replication health and performance\nEnsure systems are designed for high availability, fault tolerance, and scalability\nPartner with QA teams to drive performance, resiliency, and scale validation\nCollaborate with backend, security, and platform teams to deliver end-to-end replication workflows\nParticipate in debugging, production issue resolution, and continuous improvement of replication reliability\nProvide technical leadership, architectural guidance, and mentorship to the engineering team\n\nRequired Qualifications\n8+ years of experience in distributed systems, storage systems, or backend software engineering\nStrong programming skills in one or more languages: C++, Go, Java, or Rust\nExperience designing and building data replication systems, data pipelines, or distributed data services\nDeep understanding of distributed systems concepts (consistency, availability, scalability, fault tolerance)\nStrong expertise in multi-threading, concurrency, and parallel processing\nKnowledge of networking protocols and secure communication (TCP/IP, HTTP/HTTPS, TLS)\nExperience implementing data integrity mechanisms (checksums, validation, consistency checks)\nExperience designing and building REST APIs and service-based architectures\nFamiliarity with checkpointing, failure recovery, and retry mechanisms in distributed systems\nBasic understanding of observability concepts (metrics, logging, alerting)\nStrong debugging, problem-solving, and system design skills\n\nPreferred Qualifications\nExperience with asynchronous replication, disaster recovery (DR), or backup systems\nFamiliarity with object storage or large-scale data storage systems\nKnowledge of delta encoding, change data capture, or incremental data synchronization techniques\nExperience building high-throughput, low-latency data movement systems\nExposure to security practices including mutual TLS, encryption, and authentication\nExperience working on enterprise-scale data platforms or storage products\nFamiliarity with performance optimization and large-scale system tuning","datePosted":"2026-08-05T16:13:01.393Z","dateModified":"2026-08-05T16:13:01.393Z","hiringOrganization":{"@type":"Organization","name":"Ddn","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"San Jose","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"86829f894e91b79c4cba7ab5"},"url":"https://jobsearcher.com/jobs/86829f894e91b79c4cba7ab5"}}