Staff Distributed Systems Engineer

Recruiting from Scratch LLC

New York (NY)

On-site

USD 180,000 - 240,000

Full time

4 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive compensation
In-person office in New York

Job summary

Recruiting from Scratch LLC in New York seeks a highly autonomous Staff-level distributed systems engineer to own multi-region backend infrastructure, improve reliability, latency, and scaling for high-throughput trading workloads.

You will design and operate high-throughput services, manage datastores, caches, messaging, disaster recovery, and incident response, mentoring engineers and elevating the team's technical bar.

Qualifications

  • 8+ years of backend or distributed systems experience
  • Experience designing, operating, and debugging high-throughput production systems
  • Deep knowledge of production reliability and distributed-system failure modes
  • Strong knowledge of query performance, indexing, and replication
  • Experience with AWS and infrastructure-as-code practices ( Terraform )
  • Proficient programming in Go, TypeScript/Node.js, or equivalent
  • Experience building failover and disaster-recovery systems
  • Ability to mentor engineers and raise technical standards

Responsibilities

  • Own reliability, scalability, and performance of a multi-region backend platform
  • Design and operate high-throughput services supporting production trading workloads
  • Own critical shared infrastructure including datastores, caches, messaging systems, and regional services
  • Improve database and cache performance, capacity planning, replication, and failure handling
  • Build systems that remain predictable during traffic spikes, dependency failures, infrastructure changes, and outages
  • Implement resilience patterns including backpressure, concurrency limits, load shedding, rate limiting, circuit breakers, and bounded retries
  • Reduce cross-region latency and improve data locality
  • Design and implement service, datastore, and regional failover capabilities
  • Establish disaster-recovery procedures, recovery objectives, and validation processes
  • Build and test backup restoration and data-recovery workflows
  • Define and validate RTO and RPO targets
  • Help architect new product features for scale from the start
  • Improve observability across distributed services and data infrastructure
  • Remain hands-on with production systems and raise the distributed systems bar

Skills

Distributed systems
Go
AWS
Observability
Performance optimization
Database design
Caching strategies
Mentoring
Incident response
TypeScript/Node.js

Tools

Redis
NATS JetStream
Kafka
Datadog APM
Terraform

Job description

Recruiting from Scratch is a specialized talent firm dedicated to helping companies build exceptional teams. We partner closely with our clients to deeply understand their needs, then connect them with top-tier candidates who are not only highly skilled but also the right fit for the company’s culture and vision. Our mission is simple: place the best people in the right roles to drive long-term success for both clients and candidates.

https://www.recruitingfromscratch.com/

Location

New York

Company Stage of Funding

Early-Stage Consumer Trading / Crypto Company

Office Type

In Person

Company Description

We’re representing a consumer trading company building a simplified way for users to access on-chain assets without needing external wallets, bridges, or prior crypto expertise.

The product combines trading infrastructure with social discovery, allowing users to follow other traders, view portfolios and trades, discover emerging tokens, and access professional-grade execution and market data within a consumer-friendly experience.

As usage grows, the company is investing in the distributed systems infrastructure required to keep its backend reliable, performant, and predictable across regions and during periods of highly variable trading activity.

What You Will Do
  • Own the reliability, scalability, and performance of a multi-region backend platform.
  • Design and operate high-throughput services supporting production trading workloads.
  • Own critical shared infrastructure including datastores, caches, messaging systems, and regional application services.
  • Improve database and cache performance, capacity planning, replication, and failure handling.
  • Build systems that remain predictable during traffic spikes, dependency failures, infrastructure changes, and partial regional outages.
  • Implement resilience patterns including backpressure, concurrency limits, load shedding, rate limiting, circuit breakers, and bounded retries.
  • Reduce cross-region latency and improve data locality.
  • Design and implement service, datastore, and regional failover capabilities.
  • Establish disaster-recovery procedures, recovery objectives, and validation processes.
  • Build and test backup restoration and data-recovery workflows.
  • Define and validate RTO and RPO targets for critical services and data.
  • Help architect new product features so they can operate reliably at scale from the beginning.
  • Improve observability across distributed services and data infrastructure.
  • Remain directly hands-on with production systems rather than operating solely as an architect.
  • Raise the distributed systems bar across the broader engineering team and help other engineers reason more effectively about scale and failure modes.
Ideal Background
  • 8+ years of backend, platform, infrastructure, or distributed systems engineering experience, or equivalent practical experience.
  • Significant experience designing, operating, and debugging high-throughput distributed production systems.
  • Deep knowledge of production reliability and distributed-system failure modes.
  • Query performance and optimization
  • Indexing
  • Connection pooling
  • Replication
  • Transaction contention
  • Strong experience with Redis-compatible systems, such as Redis, Valkey, Dragonfly, or KeyDB.
  • Comfortable reasoning about cache sharding, replication, memory management, hot keys, and failure recovery.
  • Experience operating production services in AWS.
  • Familiarity with infrastructure-as-code practices, preferably Terraform.
  • Strong programming ability in Go, TypeScript/Node.js, or a comparable systems-oriented language.
  • Hands-on experience building and testing failover and disaster-recovery systems.
  • Experience with backup restoration, replication strategy, regional failover, and recovery validation.
  • Able to independently own critical production infrastructure from architecture through operation and incident response.
  • Strong systems judgment around performance, capacity, reliability, and operational complexity.
Preferred
  • Experience with NATS JetStream, Kafka, or another durable messaging platform.
  • Familiarity with Datadog APM.
  • Experience using AWS Performance Insights or similar database-performance tooling.
  • Hands-on experience performing live datastore or cache topology migrations.
  • Experience operating systems with highly bursty or unpredictable traffic patterns.
  • Background in financial systems, trading, cryptocurrency, gaming, or another high-throughput real-time domain.
  • Experience with AWS services such as ECS, RDS, and ElastiCache.
  • Experience improving data locality and latency across multi-region architectures.
  • Strong track record of mentoring engineers or raising the technical standard for distributed systems across a team.
Compensation and Benefits
  • Compensation: Competitive
  • Employment type and location details were not included in the source materials.
  • Highly hands-on Staff-level individual contributor role with direct ownership of critical production infrastructure.
  • Significant responsibility for the company’s multi-region architecture, datastore reliability, caching strategy, failure handling, and disaster recovery.
  • Opportunity to shape how the engineering organization approaches scalability and distributed systems as the trading platform grows.
  • Best suited for an engineer who enjoys working directly on databases, caches, messaging, regional infrastructure, and production failure modes, rather than a role focused primarily on application feature development.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Distributed Systems Engineer
Staff Distributed Systems Engineer

Alexander Chapman • New York (NY)

On-site
USD 180,000 - 240,000
Staff Distributed Systems Engineer - Multi-Region Backend
Staff Distributed Systems Engineer - Multi-Region Backend

Recruiting from Scratch LLC • New York (NY)

On-site
USD 180,000 - 240,000
Competitive compensation
In-person office in New York
Staff Distributed Systems Engineer
Staff Distributed Systems Engineer

fomo Labs • New York (NY)

On-site
USD 140,000 - 210,000
Staff Distributed Systems Engineer
Staff Distributed Systems Engineer

SOLANA FOUNDATION • New York (NY), Northern (KY)

On-site
USD 180,000 - 270,000
Staff Distributed Systems Engineer
Staff Distributed Systems Engineer

FOMO Labs Inc. • New York (NY), Northern (KY)

Hybrid
USD 190,000 - 270,000
Staff Systems Engineer
Staff Systems Engineer

globalplacementfirm • New York (NY)

On-site
USD 280,000 - 380,000
Competitive equity package
Relocation assistance
Software Engineer Backend Systems New York
Software Engineer Backend Systems New York

AIDA Recruitment • New York (NY)

On-site
USD 185,000 - 300,000
Salary range announced
Equity package
On-site in New York
+1
Software Engineer, Backend Systems
Software Engineer, Backend Systems

AIDA Recruitment • San Francisco (CA)

On-site
USD 185,000 - 300,000
Competitive equity package
On-site work in San Francisco
Staff Backend Engineer
Staff Backend Engineer

Serv Recruitment • United States

On-site
USD 160,000 - 190,000
Medical, Dental and Vision insurance
Health & Wellness stipend
Unlimited PTO
+1
Staff Software Engineer (Data/Infrastructure)
Staff Software Engineer (Data/Infrastructure)

UMATR • New York (NY)

On-site
USD 225,000 - 275,000
Competitive salary up to $250k plus sizeable equity
Opportunity to influence core architecture
Collaborative, low-bureaucracy culture