Staff Software Engineer, Replication Foundations

United States Digital Space LLC

United States

On-site

USD 140,000 - 210,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

United States Digital Space LLC is hiring a Staff Software Engineer to join the Replication Foundations team within the CGS organization. You’ll help set the technical direction for the distributed replication stack, guide architecture decisions, and collaborate across teams to evolve reliable, scalable replication capabilities powering OSS and cloud products.

You’ll lead design, implementation, rollout, and operations, mentor engineers, and contribute to incident response and performance

Qualifications

  • Deep understanding of distributed systems fundamentals such as replication, partitioning, consistency, fault tolerance, durability, concurrency, and failure recovery.
  • Experience defining system architecture, protocol behavior, invariants, and trade-offs in ambiguous or evolving problem spaces.
  • Experience debugging complex production issues, including concurrency bugs, data inconsistencies, partial failures, and performance bottlenecks.

Responsibilities

  • Set the technical direction and evolve the architecture of the company’s OSS replication stack, from problem definition through rollout and operational support.
  • Lead the design and implementation of replication protocols that power High Availability namespaces, cross-cluster and cross-region replication, and migration between clusters.
  • Drive scalability and reliability initiatives such as multi-cell namespaces and load distribution improvements.
  • Define and communicate system-level guarantees, including consistency models, ordering, idempotency, failure recovery, performance, and operational behavior.
  • Identify architectural risks and opportunities, and shape the technical roadmap for replication capabilities that support current and future cloud products.
  • Partner with Cloud Enablement, CGS, Product, and other teams to align OSS replication foundations with customer needs.
  • Lead design reviews, mentor engineers, and provide technical guidance across the organization.
  • Lead or contribute to debugging production issues, incident response, and follow-up improvements related to replication and core system behavior.

Skills

Distributed systems
Replication
Consistency models
Fault tolerance
Concurrency
Failure recovery
System architecture
Mentoring
Communication

Tools

Go
Java
C++

Job description

Role Summary

We’re hiring a Staff Software Engineer to join the Replication Foundations team within the company’s Cloud Global Services (CGS) organization.

Replication Foundations owns and evolves the company’s core replication stack in the company OSS - the distributed systems backbone behind key the company Cloud capabilities such as High Availability namespaces, cross-cluster and cross-region failover, and migration products that enable customers to move workloads between self-hosted the company and the company Cloud. The team also builds foundational scalability and reliability mechanisms that support the company at scale.

In this role, you’ll help set the technical direction for the company’s distributed replication systems. You’ll lead complex, correctness-critical initiatives spanning architecture, design, implementation, rollout, and operations. You’ll work across teams to evolve reliable and scalable replication capabilities that support both the open source project and the company Cloud.

What You’ll Do
  • Set the technical direction and evolve the architecture of the company’s OSS replication stack, from problem definition through rollout and operational support.
  • Lead the design and implementation of replication protocols that power:
  • High Availability namespaces
  • Cross-cluster and cross-region replication
  • Migration between the company clusters, including cloud-to-self-hosted and cloud-to-cloud scenarios
  • Drive scalability and reliability initiatives such as:
  • Multi-cell namespaces
  • Enabling a namespace to span multiple clusters
  • Improving load distribution and handling hot spots
  • Define and communicate system-level guarantees, including consistency models, ordering, idempotency, failure recovery, performance, and operational behavior.
  • Identify architectural risks and opportunities, and shape the technical roadmap for replication capabilities that support current and future cloud products.
  • Partner with Cloud Enablement, CGS, Product, and other engineering teams to align OSS replication foundations with customer and product needs.
  • Lead design reviews, raise the quality of implementation and testing practices, mentor engineers, and provide technical guidance across the organization.
  • Lead or contribute to debugging complex production issues, incident response, and follow-up improvements related to replication and core system behavior.
What You’ll Bring
  • Deep understanding of distributed systems fundamentals such as replication, partitioning, consistency, fault tolerance, durability, concurrency, and failure recovery.
  • Experience defining system architecture, protocol behavior, invariants, and trade-offs in ambiguous or evolving problem spaces.
  • Experience debugging complex production issues, including concurrency bugs, data inconsistencies, partial failures, and performance bottlenecks.
  • Proficiency writing production-quality concurrent code in Go; experience with Java, C++, or similar systems languages is also welcome.
  • Strong written and verbal communication skills, including the ability to explain complex designs and trade-offs to both technical and cross-functional audiences.
  • Demonstrated ability to influence technical direction across teams, build alignment without direct authority, and mentor engineers.
  • A thoughtful and curious approach to understanding how systems behave under load, failure, and changing workload conditions.
Nice to Have
  • Experience designing or maintaining replication protocols or data-plane infrastructure.
  • Experience with multi-cluster or multi-region architectures, including active-active or active-passive systems.
  • Familiarity with database internals, log-based replication, or event-sourced systems.
  • Prior contributions to large open source projects or distributed systems infrastructure.

the company Technologies is an Equal Opportunity Employer. the company Technologies does not discriminate on the basis of race, religion, color, sex, gender identity, sexual orientation, age, non-disqualifying physical or mental disability, national origin, veteran status, or any other basis covered by appropriate law. All employment is decided on the basis of qualifications, merit, and business need. We embrace and celebrate differences and diversity.

the company is committed to providing access, equal opportunity, and reasonable accommodation for individuals with disabilities in employment, its services, programs, and activities. If you need to request a reasonable accommodation, please let your Recruiter know so we can assist.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Replication Foundations
Staff Software Engineer, Replication Foundations

Engg • United States

On-site
USD 210,000 - 260,000
Staff Software Engineer, Replication Foundations
Staff Software Engineer, Replication Foundations

Temporal • Seattle (WA)

On-site
USD 190,000 - 260,000
Staff Software Engineer, Replication Foundations
Staff Software Engineer, Replication Foundations

Engg • United States

Remote
USD 180,000 - 270,000
Staff Engineer: Distributed Replication Systems
Staff Engineer: Distributed Replication Systems

Temporal • Seattle (WA)

On-site
USD 190,000 - 260,000
Staff Replication Development Engineer
Staff Replication Development Engineer

DDN • San Francisco (CA)

On-site
USD 185,000 - 230,000
Software Engineer II, Open Source Server
Software Engineer II, Open Source Server

United States Digital Space LLC • United States

Remote
USD 120,000 - 180,000
Staff Engineer, Replication
Staff Engineer, Replication

Ddn • Raleigh (NC)

On-site
USD 150,000 - 190,000
Principal Solutions Architect
Principal Solutions Architect

Software Placement Group • Long Beach (CA)

On-site
USD 100,000 - 130,000
Senior Software Engineer, SecureBuild (remote) at Replicated
Senior Software Engineer, SecureBuild (remote) at Replicated

Feedinkoo • United States

Remote
USD 149,000 - 198,000
Health/Dental/Vision
Life/AD&D
LTD/STD
+8
Principal Data Systems Software Engineer - SRE
Principal Data Systems Software Engineer - SRE

Ll Oefentherapie • Seattle (WA), Herndon (VA)

On-site
USD 120,000 - 180,000
Top Secret clearance required