Core Infrastructure Engineer

Oracle Corporation

Nashville (TN)

On-site

USD 120,000 - 180,000

Full time

10 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Oracle Corporation in Nashville, TN seeks an engineer to implement and optimize components within existing distributed systems, applying scalability, resiliency, and automation practices to handle network variability.

You will build telemetry, alerts, and runbook-driven procedures; deliver scoped features and fault-injection tests; and help implement basic data replication and synchronization under guidance, while following change, security, and compliance procedures.

Qualifications

  • Experience designing and maintaining distributed systems under guidance.
  • Familiarity with performance/load testing and resiliency patterns.
  • Ability to implement basic data replication and synchronization.
  • Experience with on-call rotations and incident response.
  • Knowledge of security and compliance best practices.

Responsibilities

  • Assist in diagnosing and debugging issues in system components to support ongoing operation.
  • Follow protocols to prevent interruptions and avoid customer maintenance windows.
  • Run basic automation scripts and tooling to troubleshoot operational issues.
  • Participate in incident responses and root cause investigations.
  • Assist in implementing basic security measures and encryption controls.
  • Maintain automation scripts and Infrastructure as Code (IaC) under guidance.

Skills

Distributed systems
On-call rotations
Performance testing
Automation
Cloud/IaC

Tools

Kubernetes
Terraform
Monitoring/Telemetry

Job description

Implements and optimizes components within existing distributed systems under guidance. Applies basic scalability requirements, conducts performance/load testing, and configures resiliency features (retries, circuit breakers, timeouts) to handle network variability. Builds telemetry, alerts, and runbook-driven procedures; delivers scoped features and fault-injection tests; and helps implement basic data replication and synchronization. Assists with on-call rotations, uses automation/IaC scripts to troubleshoot, and follows change, security, and compliance procedures (encryption, access controls, remediation plans) while escalating complex issues to senior engineers.

Internal Responsibilities

Key Responsibilities

System Design & Architecture - System Scalability:

  • -Assist in the implementation of components of distributed systems that support horizontal and vertical scaling under the guidance of senior engineers.
  • -Optimize code segments and/or systems for large-scale data processing with oversight from senior engineers.
  • -Implement scalability requirements for assigned components.
  • -Learn about the use of data plane platforms for large-scale data retrieval, storage, and processing.
  • -Execute performance and load testing, with guidance.

System Design & Architecture - System Reliability Design:

  • -Collaborate with the team to build fault-tolerant components capable of withstanding in-service updates by learning about redundancy, replication, and automatic failover mechanisms.
  • -Learn about recovery oriented computing principles and assist in applying them to component designs.
  • -Configure and test retry mechanisms, circuit breakers, and timeouts to help handle network unreliability, with guidance

System Design & Architecture - System Reliability Performance:

  • -Implement testing and alarming configurations to detect issues/failures.
  • -Support efforts to recover from failures by drafting and executing runbooks and operational procedures, under guidance.
  • -Help build dashboards, telemetry systems, and alerting mechanisms to monitor component health.

System Design & Architecture - Correctness / Availability:

  • -Implement functional requirements and testing for assigned features within an existing system.
  • -Implement test scenarios (e.g., fault-injection, brown-out) to evaluate system correctness, under guidance.
  • -Help implement basic data replication and synchronization techniques to maintain data integrity and availability.

Operational Troubleshooting & Incident Management:

  • -Assist in diagnosing and debugging issues in system components to support ongoing operation, under supervision.
  • -Follow protocols to prevent interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.
  • -Run basic automation scripts and tooling to troubleshoot operational issues.
  • -Participate in operational support rotations, assisting in incident responses and root cause investigations.

Compliance & Security:

  • -Assist in implementing basic security measures to protect data and applications in multi-tenant environments, including encryption and access controls.
  • -Assist in the execution of remediation plans to address identified security gaps, under supervision.
  • -Support the creation and updating of documentation to ensure cloud infrastructure is in compliance with relevant industry standards and regulations.

Automation & Change Management:

  • -Assist in maintaining basic automation scripts and tools (e.g., Infrastructure as Code (IaC)).
  • -Adhere to change management plans for patching, updating, and rolling back applications, under guidance

Core Responsibilities

Planning & Execution:

  • -Track timelines with minimal supervision, ensuring work is completed in a timely manner and is in alignment with project requirements.
  • -Prioritize and adjust work as resources or timelines change, with some guidance

Collaboration & Partnership:

  • -Collaborate within the team to better understand expectations and achieve shared objectives.
  • -Leverage a foundational understanding of business, stakeholder, and/or customer needs to build partnerships with limited guidance.
  • -Actively listen and ask questions to enhance collaboration.
  • -Build a basic understanding of business, stakeholder, and/or customer needs with guidance.

Problem Solving:

  • -Identify and address issues, escalating problems to senior staff as needed in accordance with standard procedures.
  • -Compile and review data and/or information from multiple sources to troubleshoot standard and non-standard errors.

Continuous Learning:

  • -Seek opportunities to gain knowledge and learn new skills and/or tools aligned with industry trends and best practices.
  • -Utilize feedback and training to improve skills.
  • -Participate in a culture of continuous learning and knowledge sharing.

Continuous Improvement:

  • -Implement updates to processes, protocols, and workflows to increase efficiency and effectiveness as directed, with some guidance.
  • -Contribute to ideation for future process improvements.
External Responsibilities

Key Responsibilities

System Design & Architecture - System Scalability:

  • -Assist in the implementation of components of distributed systems that support horizontal and vertical scaling under the guidance of senior engineers.
  • -Optimize code segments and/or systems for large-scale data processing with oversight from senior engineers.
  • -Implement scalability requirements for assigned components.
  • -Learn about the use of data plane platforms for large-scale data retrieval, storage, and processing.
  • -Execute performance and load testing, with guidance.

System Design & Architecture - System Reliability Design:

  • -Collaborate with the team to build fault-tolerant components capable of withstanding in-service updates by learning about redundancy, replication, and automatic failover mechanisms.
  • -Learn about recovery oriented computing principles and assist in applying them to component designs.
  • -Configure and test retry mechanisms, circuit breakers, and timeouts to help handle network unreliability, with guidance

System Design & Architecture - System Reliability Performance:

  • -Implement testing and alarming configurations to detect issues/failures.
  • -Support efforts to recover from failures by drafting and executing runbooks and operational procedures, under guidance.
  • -Help build dashboards, telemetry systems, and alerting mechanisms to monitor component health.

System Design & Architecture - Correctness / Availability:

  • -Implement functional requirements and testing for assigned features within an existing system.
  • -Implement test scenarios (e.g., fault-injection, brown-out) to evaluate system correctness, under guidance.
  • -Help implement basic data replication and synchronization techniques to maintain data integrity and availability.

Operational Troubleshooting & Incident Management:

  • -Assist in diagnosing and debugging issues in system components to support ongoing operation, under supervision.
  • -Follow protocols to prevent interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.
  • -Run basic automation scripts and tooling to troubleshoot operational issues.
  • -Participate in operational support rotations, assisting in incident responses and root cause investigations.

Compliance & Security:

  • -Assist in implementing basic security measures to protect data and applications in multi-tenant environments, including encryption and access controls.
  • -Assist in the execution of remediation plans to address identified security gaps, under supervision.
  • -Support the creation and updating of documentation to ensure cloud infrastructure is in compliance with relevant industry standards and regulations.

Automation & Change Management:

  • -Assist in maintaining basic automation scripts and tools (e.g., Infrastructure as Code (IaC)).
  • -Adhere to change management plans for patching, updating, and rolling back applications, under guidance

Core Responsibilities

Planning & Execution:

  • -Track timelines with minimal supervision, ensuring work is completed in a timely manner and is in alignment with project requirements.
  • -Prioritize and adjust work as resources or timelines change, with some guidance

Collaboration & Partnership:

  • -Collaborate within the team to better understand expectations and achieve shared objectives.
  • -Leverage a foundational understanding of business, stakeholder, and/or customer needs to build partnerships with limited guidance.
  • -Actively listen and ask questions to enhance collaboration.
  • -Build a basic understanding of business, stakeholder, and/or customer needs with guidance.

Problem Solving:

  • -Identify and address issues, escalating problems to senior staff as needed in accordance with standard procedures.
  • -Compile and review data and/or information from multiple sources to troubleshoot standard and non-standard errors.

Continuous Learning:

  • -Seek opportunities to gain knowledge and learn new skills and/or tools aligned with industry trends and best practices.
  • -Utilize feedback and training to improve skills.
  • -Participate in a culture of continuous learning and knowledge sharing.

Continuous Improvement:

  • -Implement updates to processes, protocols, and workflows to increase efficiency and effectiveness as directed, with some guidance.
  • -Contribute to ideation for future process improvements.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Software Engineer, Platform
Principal Software Engineer, Platform

Oracle Corporation • Seattle (WA)

On-site
USD 140,000 - 210,000
Health insurance
401(k) with company match
Lead Principal Platform Software Engineer
Lead Principal Platform Software Engineer

Oracle Corporation • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Lead Principal Data Center Facilities Development Manager
Lead Principal Data Center Facilities Development Manager

Oracle Corporation • United States

On-site
USD 120,000 - 180,000
Senior Manager, Data Center Operations
Senior Manager, Data Center Operations

Oracle Corporation • Ashburn (VA)

On-site
USD 180,000 - 240,000
Core Infrastructure Engineer
Core Infrastructure Engineer

Oracle • Nashville (TN)

On-site
USD 90,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Ll Oefentherapie • Reston (VA)

On-site
USD 50,000 - 70,000
Software Engineer Manager - Supply Chain RE (Remote)
Software Engineer Manager - Supply Chain RE (Remote)

The Home Depot • Atlanta (GA)

Remote
USD 180,000 - 240,000
Manager, Software Engineering
Manager, Software Engineering

InComm • Atlanta (GA)

On-site
USD 140,000 - 190,000
Senior Data Hall Designer
Senior Data Hall Designer

Ll Oefentherapie • Nashville (TN)

On-site
USD 90,000 - 150,000
Senior Data Center Support Services Technician
Senior Data Center Support Services Technician

Oracle • San Antonio (TX)

On-site
USD 58,000 - 118,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
+2