Staff Distributed Systems Engineer

Alexander Chapman

New York (NY)

On-site

USD 180,000 - 240,000

Full time

16 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Alexander Chapman is seeking a Staff Distributed Systems Engineer to own and scale critical production infrastructure. You will lead design and operation of high-throughput, multi-region services, ensuring reliability, performance, and data integrity across regions.

The role emphasizes building failover, disaster recovery, and observability capabilities while mentoring engineers and raising the team's technical bar for operating systems at scale.

Qualifications

  • 10+ years in backend, platform, or distributed systems engineering.
  • Deep experience with high-throughput, distributed production systems.
  • Strong PostgreSQL performance, replication, and failure mode knowledge.
  • Experience with AWS, Terraform, and reliability-focused architectures.

Responsibilities

  • Design and operate high-throughput, multi-region backend services.
  • Own reliability, scalability, and performance of shared infrastructure.
  • Improve datastore and cache performance, replication, and failure handling.
  • Implement backpressure, rate limiting, circuit breakers, and bounded retries.
  • Reduce cross-region latency and improve data locality.
  • Test failover procedures and disaster recovery strategies.
  • Define and validate RTO/RPO objectives for critical services.
  • Enhance observability and operational tooling across the platform.
  • Mentor engineers and raise the team's reliability bar.

Skills

Distributed systems
Backend engineering
Go / TypeScript
Observability
Failover / DR design
Performance tuning
Mentoring

Tools

Terraform
PostgreSQL
Redis
AWS (ECS/RDS/ElastiCache)
Datadog APM
NATS JetStream / Kafka

Job description

We are representing a rapidly growing financial technology company building a next-generation trading platform. Our client is developing a consumer-focused platform that combines real-time trading with social discovery, giving users a streamlined way to discover market activity, follow traders, receive real-time insights, and execute trades across a growing range of financial products. Behind the consumer product is a sophisticated, highly distributed backend platform supporting real-time market data, trading activity, social features, financial data, and user-facing services across multiple regions. With significant growth ahead, the engineering organization is expanding its core infrastructure capabilities and is looking for a Staff Distributed Systems Engineer to take ownership of reliability, scalability, and performance across the platform.

About the Role

This is a hands-on Staff-level position with direct ownership of critical production infrastructure. You will own shared systems including datastores, caches, messaging infrastructure, and regional application services. The focus will be on designing systems that remain predictable and resilient through traffic surges, dependency failures, infrastructure changes, and partial regional outages. You will also lead the development of new failover and disaster-recovery capabilities, including defining recovery objectives and implementing the systems and testing required to safely recover services and data.

Responsibilities
  • Design and operate high-throughput, multi-region backend services.
  • Own the reliability, scalability, and performance of critical shared infrastructure.
  • Improve datastore and cache performance, capacity, replication, and failure handling.
  • Implement backpressure, concurrency limits, load shedding, rate limiting, circuit breakers, and bounded retries.
  • Reduce cross-region latency and improve data locality.
  • Design and test service, datastore, and regional failover procedures.
  • Build disaster-recovery systems covering backup restoration, replication, and regional recovery.
  • Define and validate RTO/RPO objectives for critical services and data.
  • Identify production bottlenecks, capacity constraints, and failure modes before they become incidents.
  • Help architect new features so they can operate reliably at scale from day one.
  • Improve observability, debugging, and operational tooling across the platform.
  • Establish strong engineering practices around distributed systems and reliability.
  • Mentor engineers and raise the team's technical bar around operating systems at scale.
Qualifications
  • 10+ years of experience in backend, platform, infrastructure, or distributed systems engineering, or equivalent practical experience.
  • Deep experience designing, operating, and debugging distributed, high-throughput production systems.
  • Strong PostgreSQL experience, including query performance, indexing, connection pooling, replication, transaction contention, and database failure modes.
  • Strong experience with Redis or Redis-compatible systems such as Valkey, Dragonfly, or KeyDB, including sharding, replication, memory management, hot keys, and failure handling.
  • Experience operating production services on AWS, ideally using ECS, RDS, and ElastiCache.
  • Strong infrastructure-as-code experience, preferably Terraform.
  • Proficiency in Go, TypeScript/Node.js, or another systems-oriented programming language.
  • Hands-on experience designing and testing failover and disaster-recovery systems.
  • Experience with backup restoration, replication, regional failover, and validating RTO/RPO objectives.
  • Strong production engineering instincts and a track record of owning systems through real-world scale and failure scenarios.
Required Skills
  • Experience with NATS JetStream, Kafka, or another durable messaging system.
  • Familiarity with Datadog APM, AWS Performance Insights, or comparable observability tooling.
  • Experience performing live datastore or cache topology migrations.
  • Experience operating systems with highly variable or bursty traffic.
  • Background in financial technology, trading, cryptocurrency, gaming, or another high-throughput environment.
  • Experience working on systems where availability, latency, and data correctness are critical.
Why This Role

This is not a purely architectural Staff position. You will be deeply involved in production systems, making decisions around databases, caching, replication, regional architecture, failure recovery, traffic management, and system performance. The role is particularly well suited to an engineer who has spent years dealing with the realities of distributed systems: overloaded services, database contention, hot keys, failed deployments, network failures, dependency degradation, regional outages, and unpredictable traffic. It is an opportunity to have broad technical ownership over the distributed systems layer of a rapidly scaling financial platform, while helping shape the architecture and engineering practices as the platform grows.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Distributed Systems Engineer
Staff Distributed Systems Engineer

FOMO Labs Inc. • New York (NY), Northern (KY)

Hybrid
USD 190,000 - 270,000
Staff Distributed Systems Engineer
Staff Distributed Systems Engineer

fomo Labs • New York (NY)

On-site
USD 140,000 - 210,000
Staff Distributed Systems Engineer
Staff Distributed Systems Engineer

SOLANA FOUNDATION • New York (NY), Northern (KY)

Hybrid
USD 180,000 - 270,000
Staff Distributed Systems Engineer
Staff Distributed Systems Engineer

fomo • New York (NY)

On-site
USD 180,000 - 280,000
Staff Software Engineer (Data/Infrastructure)
Staff Software Engineer (Data/Infrastructure)

UMATR • New York (NY)

On-site
USD 200,000 - 250,000
Competitive salary up to $250k
Sizeable equity
High ownership in architecture
+2
Staff Software Engineer (Data/Infrastructure)
Staff Software Engineer (Data/Infrastructure)

UMATR • New York (NY)

On-site
USD 212,000 - 250,000
Competitive salary up to $250k plus sizeable equity
Opportunity to influence core architecture
Collaborative, low-bureaucracy culture
Senior Full-Stack Engineer: Infrastructure & Distributed Systems
Senior Full-Stack Engineer: Infrastructure & Distributed Systems

3M HEALTHCARE • San Francisco (CA)

On-site
USD 140,000 - 180,000
Senior Software Engineer
Senior Software Engineer

NJF Global Holdings Ltd • Chicago (IL)

Hybrid
USD 140,000 - 190,000
Staff Software Engineer (backend) Columbus HQ / Remote
Staff Software Engineer (backend) Columbus HQ / Remote

Motion LLC • Columbus (OH)

Hybrid
USD 150,000 - 190,000
Senior Staff Engineer
Senior Staff Engineer

Harnham • California (MO)

Hybrid
USD 180,000 - 210,000