Designs, implements, and optimizes components in distributed systems with an emphasis on scalability, resiliency, and operability. Delivers features and load/performance tests; leverages data plane platforms and distributed state tools for high-volume retrieval, storage, and processing; and reviews peers’ implementations for scalability compliance. Builds fault-tolerant paths (redundancy, replication, automatic failover), applies recovery‑oriented principles, and implements retries, circuit breakers, and timeouts. Proactively detects and mitigates issues via tests, alarms, dashboards, and telemetry; authors runbooks and participates in incident response and RCAs. Implements standard replication and synchronization, develops automation/IaC for troubleshooting and maintenance, and applies advanced security controls (encryption, access, remediation) while ensuring change, compliance, and documentation standards are met.Qualifications:5+ years’ experience delivering and operating large scale, highly available distributed systems.Strong knowledge and interest in AI adoption including prompt engineering and agentic programming, with ChatGPT and Codex experience a plus.Strong knowledge of a base language such as Java, with a preference for functional programming language such as Scala.Strong knowledge of data structures, algorithms, operating systems, and distributed systems fundamentals.Experience with tools such as Terraform for Infrastructure as Code.Deep knowledge with networking protocols (TCP/IP, HTTP) and network architectures.Ability to design, troubleshoot and maintain networking infrastructure for high throughput use cases.Strong understanding of databases, storage, and distributed persistence technologies.Strong troubleshooting and performance tuning skills.Experience building multi-tenant, virtualized infrastructure a strong plus.