A complete application in a minute — tailored resume and cover letter, ready to send.
Ll Oefentherapie in Nashville, TN, onsite five days a week, seeks an engineer to design and optimize components in a distributed system. The role emphasizes scalability, resiliency, and operability, with features, load tests, and robust security controls.
You will build fault-tolerant paths, develop automation for maintenance, and participate in incident response and postmortems. Relocation is welcomed for this greenfield platform initiative.
Location: Nashville, TN — onsite 5 days/week. Relocation assistance is available; this is not a remote position.
The App Configuration Service team is building a new OCI platform service for application configuration and feature flighting. The team will enable OCI services to safely, consistently, and reliably manage dynamic configuration, targeted feature rollouts, experimentation, and emergency kill switches across distributed cloud environments. The team owns the end-to-end platform: service APIs, control plane, globally distributed configuration evaluation path, SDKs and self-service tooling, security and access controls, auditability, observability, and the operational model required for a Tier‑0 service. As a greenfield team, we are also building and operating in an AI‑first way—using AI‑assisted development, automation, and intelligent operational workflows to improve how we design, build, test, review, document, release, and run the service. This gives engineers the opportunity to shape both a high‑impact cloud platform and modern AI‑enabled engineering practices from the ground up.
Designs, implements, and optimizes components in distributed systems with an emphasis on scalability, resiliency, and operability. Delivers features and load/performance tests; leverages data plane platforms and distributed state tools for high‑volume retrieval, storage, and processing; and reviews peers’ implementations for scalability compliance. Builds fault‑tolerant paths (redundancy, replication, automatic failover), applies recovery‑oriented principles, and implements retries, circuit breakers, and timeouts. Proactively detects and mitigates issues via tests, alarms, dashboards, and telemetry; authors runbooks and participates in incident response and RCAs. Implements standard replication and synchronization, develops automation/IaC for troubleshooting and maintenance, and applies advanced security controls (encryption, access, remediation) while ensuring change, compliance, and documentation standards are met.