AI Systems Engineer (OCI/AI Infrastructure)

Oracle

Nashville (TN)

On-site

USD 97,000 - 306,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical/dental/vision insurance
401(k) with company match
Paid time off

Job summary

Oracle’s Hardware Platform Development Engineering team is seeking an AI Systems Engineer in Nashville to evaluate and characterize next-generation GPU and AI accelerators for OCI. This hands-on role brings up hardware platforms, enables AI training and inference stacks, runs representative workloads, and analyzes system performance under real operating conditions.

You will debug hardware/software issues, design experiments, and develop data-driven insights that explain system behavior beyond

Qualifications

  • Solid knowledge of AI/GPU/CPU platform architectures.
  • Experience with server platforms (x86/ARM) across multiple vendors.
  • Strong communication across engineering and executives.
  • Familiarity with high-speed interconnects and performance metrics.
  • Debug HW/SW interactions including vLLM on ROCm, CUDA memory leaks, and distributed runtimes.
  • Profiling training workloads and runtime graphs to optimize performance.

Responsibilities

  • Evaluate and characterize next-generation GPU/AI accelerator platforms for OCI.
  • Bring up hardware platforms, enable AI training/inference stacks, run workloads, and analyze performance.
  • Debug hardware/software integration issues and design experiments.
  • Provide data-driven insights and actionable platform recommendations for AI workloads.

Skills

AI GPU architecture
x86/ARM servers
Communication skills
Interconnects
Debugging & reliability
Profiling & performance

Education

Bachelor's degree or higher in CS/related field

Tools

CUDA
ROCm
NCCL
XLA
PyTorch
TensorFlow
JAX
Ray

Job description

Job Description

Oracle Hardware Platform Development Engineering is seeking a highly driven AI Systems Engineer to evaluate and characterize next-generation GPU and AI accelerator platforms for Oracle Cloud Infrastructure (OCI). This is a hands-on engineering role focused on bringing up new hardware platforms, enabling AI training and inference software stacks, running representative workloads, and analyzing system performance under real operating conditions. The engineer will identify whether workloads are HBM/memory-bandwidth, compute, scale-up, or scale-out bound, while characterizing power, thermals, memory behavior, utilization, scaling, and performance efficiency. Working directly in the lab, you will debug hardware/software integration issues, design and execute experiments, and develop data-driven insights that explain system behavior beyond benchmark results. A key part of the role is comparative architecture analysis across GPUs and emerging AI accelerators. You will evaluate architectural tradeoffs and translate performance findings into clear, actionable recommendations on which platforms are best suited for specific AI training and inference workloads. You will work closely with internal hardware and software teams as well as technology partners to help shape Oracle’s next generation of high-performance AI infrastructure. Position Overview: This position is ideal for someone who loves deep systems engineering, debugging complex hardware-software interactions, and optimizing performance at every layer of the ML stack. You will play a pivotal role in enabling the training and deployment of next-generation LLMs and generative AI models.


Responsibilities

Required Qualifications

  • Solid knowledge of AI / GPU or/and AI/CPU platform architecture and their capabilities.
  • Experience with the architecture, design, and implementation of modern server platforms consisting of multiple architectures and vendors, including x86 and ARM server architectures.
  • Strong communications skills and ability to clearly communicate complex technical issue across engineering disciplines as well as clearly and succinctly articulate issues for executives.
  • Experience and understanding of the latest high-speed busses and interconnect used in modern Compute and AI platforms. Familiarity with their startup connectivity and operational robustness as well as performance metrics.
  • Debugging & Reliability: Troubleshoot complex hardware-software interaction issues, including vLLM compilation failures on ROCm, CUDA memory leaks, distributed runtime failures, and kernel-level inconsistencies.
  • Profiling & Performance Analysis: Conduct detailed profiling of compilation graphs, training workloads, and runtime execution to optimize performance and eliminate bottlenecks.

Preferred Qualifications

  • Minimum of 8+ years of experience in developing software infrastructure for large scale AI systems.
  • Bachelor's degree or higher in Computer Science or a related technical field (or equivalent experience).
  • Strong debugging skills and experience in analyzing and triaging AI applications from the application level to the hardware level.
  • Hands-on experience maintaining or building ML training stacks involving CUDA, ROCm, NCCL, XLA, or similar technologies.
  • Experience in benchmarking AI workloads across different architectures.
  • Background in working with the large scale clusters
  • Good understanding on DL frameworks internal PyTorch, TensorFlow, JAX, and Ray

Qualifications

Disclaimer: Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.


Range and benefit information provided in this posting are specific to the stated locations only


US Salary Range

US: Hiring Range in USD from: $96,800 - $306,400 per year. May be eligible for bonus, equity, and compensation deferral. Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business. Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.


Benefits

  • Medical, dental, and vision insurance, including expert medical opinion
  • Short term disability and long term disability
  • Life insurance and AD&D
  • Supplemental life insurance (Employee/Spouse/Child)
  • Health care and dependent care Flexible Spending Accounts
  • Pre-tax commuter and parking benefits
  • 401(k) Savings and Investment Plan with company match
  • Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  • 11 paid holidays
  • Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
  • Paid parental leave
  • Adoption assistance
  • Employee Stock Purchase Plan
  • Financial planning and group legal
  • Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.


Career Level - IC5

About Us

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.


True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.


We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.


Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Systems Engineer (OCI/AI Infrastructure)
AI Systems Engineer (OCI/AI Infrastructure)

Oracle • United States

On-site
USD 97,000 - 306,000
Medical/dental/vision
401(k) plan
Paid time off
+5
Senior Software Developer - AI Infra Compute
Senior Software Developer - AI Infra Compute

Oracle • Nashville (TN)

On-site
USD 89,000 - 210,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
Architect, AI Infrastructure
Architect, AI Infrastructure

Ll Oefentherapie • San Juan (PR)

On-site
USD 170,000 - 355,000
Health insurance
Disability insurance
Life insurance
+6
Architect, AI Infrastructure
Architect, AI Infrastructure

Ll Oefentherapie • United States

On-site
USD 170,000 - 355,000
Medical, dental, and vision insurance
401(k) Savings with company match
Paid time off
Principal Core Infrastructure Software Engineer
Principal Core Infrastructure Software Engineer

Ll Oefentherapie • Nashville (TN)

On-site
USD 115,000 - 235,000
Medical, dental, and vision insurance
Paid time off
401(k) plan with company match
Principal AI Software Engineer
Principal AI Software Engineer

Oracle • Olympia (WA)

On-site
USD 126,000 - 264,000
Senior Software Developer - AI Infra Compute
Senior Software Developer - AI Infra Compute

Oracle • United States

On-site
USD 89,000 - 210,000
Medical insurance
Dental insurance
Vision insurance
+2
Principal Core Infrastructure Engineer
Principal Core Infrastructure Engineer

Ll Oefentherapie • Salem (OR)

On-site
USD 115,000 - 235,000
Program Manager 4
Program Manager 4

Oracle • Frankfort (KY)

On-site
USD 126,000 - 210,000
Medical insurance
Dental insurance
Vision insurance
+7
Business Development Consultant 3-Corp Plan
Business Development Consultant 3-Corp Plan

Ll Oefentherapie • San Francisco (CA)

On-site
USD 83,000 - 166,000