Staff Engineer, Replication

Ddn

Raleigh (NC)

On-site

USD 150,000 - 190,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

DDN is seeking a Staff Replication Development Engineer to lead the design and development of the replication engine for the Infinia AI Data Platform. This role focuses on building enterprise-grade asynchronous replication capabilities for reliable disaster recovery in large-scale data systems.

You will develop high-performance replication pipelines, secure data transfer systems, and robust data integrity mechanisms, partnering with backend, security, and platform teams to deliver end-to-end

Qualifications

  • 8+ years of experience in distributed systems, storage systems, or backend software engineering.
  • Strong programming skills in C++, Go, Java, or Rust.
  • Experience designing and building data replication systems, data pipelines, or distributed data services.
  • Deep understanding of distributed systems concepts (consistency, availability, scalability, fault tolerance).
  • Strong expertise in multi-threading, concurrency, and parallel processing.
  • Knowledge of networking protocols and secure communication (TCP/IP, HTTP/HTTPS, TLS).
  • Experience implementing data integrity mechanisms (checksums, validation, consistency checks).
  • Experience designing and building REST APIs and service-based architectures.
  • Familiarity with checkpointing, failure recovery, and retry mechanisms in distributed systems.
  • Basic understanding of observability concepts (metrics, logging, alerting).
  • Strong debugging, problem-solving, and system design skills.

Responsibilities

  • Design and develop multi-threaded asynchronous replication systems with parallel streaming capabilities.
  • Build object-level delta replication with checkpointing and resume functionality.
  • Develop replication engines supporting bucket/share-level replication controls.
  • Implement secure data transfer mechanisms using TLS 1.3 with mutual authentication.
  • Ensure end-to-end data integrity through checksum validation and verification pipelines.
  • Design and implement manual failover workflows for disaster recovery scenarios.
  • Build and maintain REST APIs for replication configuration, control, and automation.
  • Develop metadata tracking and change detection systems to enable efficient replication.
  • Implement RPO visibility, alerting, and operational insights for replication status.
  • Contribute to monitoring dashboards focused on replication health and performance.
  • Ensure systems are designed for high availability, fault tolerance, and scalability.
  • Partner with QA teams to drive performance, resiliency, and scale validation.
  • Collaborate with backend, security, and platform teams to deliver end-to-end replication workflows.
  • Participate in debugging, production issue resolution, and continuous improvement of replication reliability.
  • Provide technical leadership, architectural guidance, and mentorship to the engineering team.

Skills

Distributed systems
C++
Go
Java
Rust
Multi-threading
Concurrency
Networking
Security
Observability

Job description

DDN is seeking a Staff Replication Development Engineer to lead the design and development of the replication engine for the Infinia AI Data Platform. This role focuses on building enterprise-grade asynchronous replication capabilities that enable reliable and secure disaster recovery for large-scale data systems.

You will work on developing high-performance replication pipelines, efficient data synchronization mechanisms, and secure data transfer systems. This role requires deep expertise in distributed systems and strong technical leadership to deliver a scalable and resilient replication foundation.

Key Responsibilities
  • Design and develop multi-threaded asynchronous replication systems with parallel streaming capabilities

  • Build object-level delta replication with checkpointing and resume functionality

  • Develop replication engines supporting bucket/share-level replication controls

  • Implement secure data transfer mechanisms using TLS 1.3 with mutual authentication

  • Ensure end-to-end data integrity through checksum validation and verification pipelines

  • Design and implement manual failover workflows for disaster recovery scenarios

  • Build and maintain REST APIs for replication configuration, control, and automation

  • Develop metadata tracking and change detection systems to enable efficient replication

  • Implement RPO visibility, alerting, and operational insights for replication status

  • Contribute to monitoring dashboards focused on replication health and performance

  • Ensure systems are designed for high availability, fault tolerance, and scalability

  • Partner with QA teams to drive performance, resiliency, and scale validation

  • Collaborate with backend, security, and platform teams to deliver end-to-end replication workflows

  • Participate in debugging, production issue resolution, and continuous improvement of replication reliability

  • Provide technical leadership, architectural guidance, and mentorship to the engineering team

Required Qualifications
  • 8+ years of experience in distributed systems, storage systems, or backend software engineering

  • Strong programming skills in one or more languages: C++, Go, Java, or Rust

  • Experience designing and building data replication systems, data pipelines, or distributed data services

  • Deep understanding of distributed systems concepts (consistency, availability, scalability, fault tolerance)

  • Strong expertise in multi-threading, concurrency, and parallel processing

  • Knowledge of networking protocols and secure communication (TCP/IP, HTTP/HTTPS, TLS)

  • Experience implementing data integrity mechanisms (checksums, validation, consistency checks)

  • Experience designing and building REST APIs and service-based architectures

  • Familiarity with checkpointing, failure recovery, and retry mechanisms in distributed systems

  • Basic understanding of observability concepts (metrics, logging, alerting)

  • Strong debugging, problem-solving, and system design skills

Preferred Qualifications
  • Experience with asynchronous replication, disaster recovery (DR), or backup systems

  • Familiarity with object storage or large-scale data storage systems

  • Knowledge of delta encoding, change data capture, or incremental data synchronization techniques

  • Experience building high-throughput, low-latency data movement systems

  • Exposure to security practices including mutual TLS, encryption, and authentication

  • Experience working on enterprise-scale data platforms or storage products

  • Familiarity with performance optimization and large-scale system tuning

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Replication Development Engineer
Staff Replication Development Engineer

DDN • San Francisco (CA)

On-site
USD 185,000 - 230,000
Staff Replication Engineer – High-Availability Data Platform
Staff Replication Engineer – High-Availability Data Platform

Ddn • Raleigh (NC)

On-site
USD 150,000 - 190,000
Lead Replication Systems Engineer
Lead Replication Systems Engineer

DDN • San Francisco (CA)

On-site
USD 185,000 - 230,000
Staff Engineer
Staff Engineer

Ddn • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Staff Software Engineer, AiDP
Staff Software Engineer, AiDP

Ddn • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Database Replication Engineer (On-site Indiana)
Database Replication Engineer (On-site Indiana)

DXC Technology Company • Indiana (PA)

On-site
USD 90,000 - 150,000
Principal Solutions Architect
Principal Solutions Architect

Software Placement Group • Long Beach (CA)

On-site
USD 100,000 - 130,000
Staff Software Engineer, Replication Foundations
Staff Software Engineer, Replication Foundations

United States Digital Space LLC • United States

On-site
USD 140,000 - 210,000
Database Replication Engineer (On-site Indiana)
Database Replication Engineer (On-site Indiana)

DXC Technology • Indiana (PA)

On-site
USD 110,000 - 150,000
Staff Engineer, Lakeflow Disaster Recovery & Replication
Staff Engineer, Lakeflow Disaster Recovery & Replication

Cacheflow • San Francisco (CA)

On-site
USD 192,000 - 260,000