Recruiting from Scratch is a specialized talent firm dedicated to helping companies build exceptional teams. We partner closely with our clients to deeply understand their needs, then connect them with top-tier candidates who are not only highly skilled but also the right fit for the company’s culture and vision. Our mission is simple: place the best people in the right roles to drive long-term success for both clients and candidates.
https://www.recruitingfromscratch.com/
Location
New York
Company Stage of Funding
Early-Stage Consumer Trading / Crypto Company
Office Type
In Person
Company Description
We’re representing a consumer trading company building a simplified way for users to access on-chain assets without needing external wallets, bridges, or prior crypto expertise.
The product combines trading infrastructure with social discovery, allowing users to follow other traders, view portfolios and trades, discover emerging tokens, and access professional-grade execution and market data within a consumer-friendly experience.
As usage grows, the company is investing in the distributed systems infrastructure required to keep its backend reliable, performant, and predictable across regions and during periods of highly variable trading activity.
What You Will Do
- Own the reliability, scalability, and performance of a multi-region backend platform.
- Design and operate high-throughput services supporting production trading workloads.
- Own critical shared infrastructure including datastores, caches, messaging systems, and regional application services.
- Improve database and cache performance, capacity planning, replication, and failure handling.
- Build systems that remain predictable during traffic spikes, dependency failures, infrastructure changes, and partial regional outages.
- Implement resilience patterns including backpressure, concurrency limits, load shedding, rate limiting, circuit breakers, and bounded retries.
- Reduce cross-region latency and improve data locality.
- Design and implement service, datastore, and regional failover capabilities.
- Establish disaster-recovery procedures, recovery objectives, and validation processes.
- Build and test backup restoration and data-recovery workflows.
- Define and validate RTO and RPO targets for critical services and data.
- Help architect new product features so they can operate reliably at scale from the beginning.
- Improve observability across distributed services and data infrastructure.
- Remain directly hands-on with production systems rather than operating solely as an architect.
- Raise the distributed systems bar across the broader engineering team and help other engineers reason more effectively about scale and failure modes.
Ideal Background
- 8+ years of backend, platform, infrastructure, or distributed systems engineering experience, or equivalent practical experience.
- Significant experience designing, operating, and debugging high-throughput distributed production systems.
- Deep knowledge of production reliability and distributed-system failure modes.
- Query performance and optimization
- Indexing
- Connection pooling
- Replication
- Transaction contention
- Strong experience with Redis-compatible systems, such as Redis, Valkey, Dragonfly, or KeyDB.
- Comfortable reasoning about cache sharding, replication, memory management, hot keys, and failure recovery.
- Experience operating production services in AWS.
- Familiarity with infrastructure-as-code practices, preferably Terraform.
- Strong programming ability in Go, TypeScript/Node.js, or a comparable systems-oriented language.
- Hands-on experience building and testing failover and disaster-recovery systems.
- Experience with backup restoration, replication strategy, regional failover, and recovery validation.
- Able to independently own critical production infrastructure from architecture through operation and incident response.
- Strong systems judgment around performance, capacity, reliability, and operational complexity.
Preferred
- Experience with NATS JetStream, Kafka, or another durable messaging platform.
- Familiarity with Datadog APM.
- Experience using AWS Performance Insights or similar database-performance tooling.
- Hands-on experience performing live datastore or cache topology migrations.
- Experience operating systems with highly bursty or unpredictable traffic patterns.
- Background in financial systems, trading, cryptocurrency, gaming, or another high-throughput real-time domain.
- Experience with AWS services such as ECS, RDS, and ElastiCache.
- Experience improving data locality and latency across multi-region architectures.
- Strong track record of mentoring engineers or raising the technical standard for distributed systems across a team.
Compensation and Benefits
- Compensation: Competitive
- Employment type and location details were not included in the source materials.
- Highly hands-on Staff-level individual contributor role with direct ownership of critical production infrastructure.
- Significant responsibility for the company’s multi-region architecture, datastore reliability, caching strategy, failure handling, and disaster recovery.
- Opportunity to shape how the engineering organization approaches scalability and distributed systems as the trading platform grows.
- Best suited for an engineer who enjoys working directly on databases, caches, messaging, regional infrastructure, and production failure modes, rather than a role focused primarily on application feature development.