Remote SRE, Hardware Infra for AI Cloud

Nebius

United States

On-site

USD 130,000 - 180,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

100% company-paid medical, dental, and vision insurance
401(k) plan with company match
20 weeks paid parental leave for primary caregivers
Remote work reimbursement
Company-paid short-term, long-term, and life insurance

Job summary

Nebius is looking for a Site Reliability Engineer to join their Hardware Infrastructure team. This position offers the flexibility of primarily remote work with occasional travel to data centers.

Responsibilities include ensuring fault-tolerance, using advanced technology for infrastructure problems, and improving CI/CD processes. The ideal candidate should have strong Linux skills, proficiency in Python and Bash scripting, and a knack for troubleshooting complex systems. This role comes with excellent benefits including full health coverage and a competitive salary.

Qualifications

  • Proficiency in Linux systems, with expertise in Python and Bash scripting for automation.
  • Demonstrated ability to troubleshoot complex system issues, including hardware, software, and networking problems.
  • Strong analytical and problem-solving skills, with a focus on optimizing system performance.
  • Working proficiency in English.

Responsibilities

  • Ensure fault-tolerance, scale and uninterrupted operations for our services.
  • Use cutting-edge technology to solve a variety of infrastructure problems.
  • Implement and improve CI/CD processes.

Skills

Linux systems proficiency
Python scripting
Bash scripting
Troubleshooting complex system issues
Analytical skills
Problem-solving skills
Working proficiency in English

Job description

Nebius is looking for a Site Reliability Engineer to join their Hardware Infrastructure team. This position offers the flexibility of primarily remote work with occasional travel to data centers.

Responsibilities include ensuring fault-tolerance, using advanced technology for infrastructure problems, and improving CI/CD processes. The ideal candidate should have strong Linux skills, proficiency in Python and Bash scripting, and a knack for troubleshooting complex systems. This role comes with excellent benefits including full health coverage and a competitive salary.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote SRE — AI Cloud Hardware Infra
Remote SRE — AI Cloud Hardware Infra

Nebius • United States

On-site
USD 130,000 - 180,000
Health insurance
401(k) plan
Parental leave
+2
Network SRE: Reliability & Automation for Cloud Infra
Network SRE: Reliability & Automation for Cloud Infra

Nebius • United States

Remote
USD 140,000 - 210,000
Global Network SRE for AI Cloud Infrastructure
Global Network SRE for AI Cloud Infrastructure

Socket.dev • United States

On-site
USD 180,000 - 224,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Senior SRE - Compute Nodes (Linux & Virtualization)
Senior SRE - Compute Nodes (Linux & Virtualization)

Nebius • United States

Remote
USD 140,000 - 190,000
Senior SRE: Cloud Reliability, CI/CD & High-Load Ops
Senior SRE: Cloud Reliability, CI/CD & High-Load Ops

Nebius • United States

Remote
USD 120,000 - 170,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Remote SRE Engineer — Cloud Reliability & Automation
Remote SRE Engineer — Cloud Reliability & Automation

Noctua Technology • Virginia (MN), California (MO), Washington

Remote
USD 106,500 - 177,500
Senior SRE – AI-Driven Infra, Remote & Async
Senior SRE – AI-Driven Infra, Remote & Async

Remote • United States

Remote
USD 53,300 - 119,850
Budget for local social events
Flexible life-work balance
Remote Senior SRE: Build Reliable, Scalable AI Infra
Remote Senior SRE: Build Reliable, Scalable AI Infra

Runware • Town of Sweden (NY)

On-site
USD 140,000 - 190,000
Generous paid time off
Meaningful stock options
Remote-first setup
+3
Senior AI-Driven SRE for Cloud Reliability
Senior AI-Driven SRE for Cloud Reliability

Cerebras • Mountain View (CA)

Hybrid
USD 100,000 - 150,000
Competitive salary and benefits package
Opportunities for professional growth
Collaborative work environment
Infrastructure Site Reliability Engineer
Infrastructure Site Reliability Engineer

Socket.dev • United States

On-site
USD 180,000 - 224,000
Competitive compensation
Career growth
Flexibility and ownership
+3