Job Description
NOTE: This role will sit onsite at our Nashville, TN location.
Are you interested in building large-scale distributed infrastructure for the cloud? Oracle’s Cloud Infrastructure team is building Infrastructure-as-a-Service technologies that operate at high scale in a broadly distributed multi-tenant cloud environment. Our customers run their businesses on our cloud, and our mission is to provide them with industry-leading compute, storage, networking, database, security, and an ever-expanding set of foundational cloud-based services.
As part of this effort, the Object Storage Service team is looking for hands‑on engineers with expertise and passion in solving difficult problems in distributed systems, large-scale storage, and highly available services. If this is you, you can be part of the team that drives the best-in-class Object Storage Service into the next phase of its development. These are exciting times for the service – we are growing fast, and delivering innovative, enterprise-class features to satisfy the most demanding big data and enterprise workloads for our customers. An engineer at any level can have significant technical and business impact.
Responsibilities
- Design, implement, and optimize components in distributed systems with an emphasis on scalability, resiliency, and operability.
- Deliver features and load/performance tests; leverage data‑plane platforms and distributed state tools for high‑volume retrieval, storage, and processing; and review peers’ implementations for scalability compliance.
- Build fault‑tolerant paths (redundancy, replication, automatic fail‑over), apply recovery‑oriented principles, and implement retries, circuit breakers, and timeouts.
- Proactively detect and mitigate issues via tests, alarms, dashboards, and telemetry; author runbooks and participate in incident response and root‑cause analyses.
- Implement standard replication and synchronization techniques, develop automation/IaC for troubleshooting and maintenance, and apply advanced security controls (encryption, access, remediation) while ensuring change, compliance, and documentation standards are met.
System Design & Architecture – System Scalability
- Implement and contribute to the development of components of distributed systems that support horizontal and vertical scaling including leveraging distributed state management tools.
- Optimize code and/or systems for large-scale data processing.
- Implement scalability requirements for assigned components and review implementation of team members.
- Leverage components of data‑plane platforms to handle large-scale data retrieval, storage, and processing.
- Implement performance and load testing.
System Design & Architecture – System Reliability Design
- Collaborate with the team to build fault‑tolerant components capable of withstanding in-service updates by implementing redundancy, replication, and automatic fail‑over mechanisms.
- Apply recovery-oriented computing principles to design components that effectively handle service disruptions.
- Implement retry mechanisms, circuit breakers, and timeouts to help handle network unreliability.
System Design & Architecture – System Reliability Performance
- Implement tests and alarm configurations to proactively detect and address issues or failures.
- Support efforts to recover from failures by drafting and executing runbooks and operational procedures.
- Build and customize dashboards, telemetry systems, and alerting mechanisms to monitor component health.
System Design & Architecture – Correctness / Availability
- Design and implement functional requirements and testing for assigned features within an existing system.
- Implement test scenarios (e.g., fault-injection, brown-out) to evaluate system correctness.
- Implement standard data replication and synchronization techniques to maintain data integrity and availability.
Operational Troubleshooting & Incident Management
- Diagnose, debug, and resolve issues in system components to support ongoing operations.
- Implement basic strategies to prevent interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.
- Design and implement automation scripts and tooling used to troubleshoot operational issues.
- Participate in operational support rotations, assisting in incident responses and root cause investigations.
Compliance & Security
- Apply advanced security measures to protect data and applications in multi-tenant environments, including encryption and access controls.
- Implement remediation plans to continuously improve security.
- Collaborate with the team to ensure cloud infrastructure complies with relevant industry standards and regulations and that documentation is up-to-date.
Automation & Change Management
- Maintain automation scripts and tools (e.g., Infrastructure as Code (IaC)) for managing cloud infrastructure.
- Adhere to change management plans for patching, updating, and rolling back applications.
Core Responsibilities
- Plan & execute tasks with minimal supervision, ensuring timely completion and alignment with project requirements. Prioritize and adjust work as resources or timelines change.
- Collaborate across teams to align on expectations and achieve shared objectives. Build and maintain a comprehensive understanding of business, stakeholder, and customer needs to support effective partnerships.
- Identify and address standard and non-standard issues independently, escalating more complex issues as appropriate. Analyze data and information from multiple sources to troubleshoot errors and contribute to knowledge sharing.
- Embrace continuous learning by actively seeking to build knowledge, acquire new skills and tools, and stay current with industry trends and best practices. Contribute to a culture of continuous learning and knowledge sharing.
- Develop ideas and recommend updates to increase the efficiency and effectiveness of processes, protocols, and workflows within a team. Seek input from team members on alternative approaches for improvement.
Qualifications
Disclaimer: Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.
Benefits
- Medical, dental, and vision insurance, including expert medical opinion
- Short-term and long-term disability
- Life insurance and AD&DD
- Supplemental life insurance (Employee/Spouse/Child)
- Health care and dependent care Flexible Spending Accounts
- Pre-tax commuter and parking benefits
- 401(k) Savings and Investment Plan with company match
- Paid time off according to company policy
- 11 paid holidays
- Paid sick leave: 72 hours upon hire, refreshed annually with a maximum cap of 112 hours
- Paid parental leave
- Adoption assistance
- Employee Stock Purchase Plan
- Financial planning and group legal services
- Voluntary benefits including auto, homeowner, and pet insurance
Career Level – IC3
About Us
Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. With AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.
We are committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494.
Equal Employment Opportunity
Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider qualified applicants with arrest and conviction records pursuant to applicable law.