Senior Global Server Ops & Reliability Engineer

Bytedance

San Jose (CA)

On-site

USD 122,000 - 272,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

ByteDance is seeking a server operations and maintenance expert to manage new product introduction and online quality operations for its global server fleet. Responsibilities include establishing operation baselines, hardware monitoring, and coordinating with R&D to resolve complex hardware issues.

The role requires strong Linux skills, server lifecycle knowledge, and the ability to work across teams. International travel up to four times per year is expected, with visa support as needed.

Qualifications

  • Bachelor’s degree or higher in Computer Engineering, Computer Science or related majors.
  • Experience in operation and maintenance of large-scale servers in the Internet industry.
  • Familiar with server hardware products lifecycle management and project documentation.

Responsibilities

  • Oversee operation management of server new product introduction and develop adaptation and maintenance baselines.
  • Design hardware monitoring plans and validate monitoring rules for new models and scenarios.
  • Analyze complex hardware issues using server-level logs and coordinate with R&D to find root causes.
  • Build and maintain a hardware failure knowledge base and promote monitoring automation and AIOps initiatives.
  • Establish unified online quality standards and proactive risk management across batches.
  • Travel up to four times per year for visa-supported assignments as required.

Skills

Linux proficiency
Server operation & maintenance
Hardware monitoring
BMC/IPMI/Redfish knowledge
Scripting (Python/Shell)

Education

Bachelor’s degree or higher in Computer Engineering, Computer Science or related majors

Tools

BMC
IPMI
Redfish

Job description

ByteDance is seeking a server operations and maintenance expert to manage new product introduction and online quality operations for its global server fleet. Responsibilities include establishing operation baselines, hardware monitoring, and coordinating with R&D to resolve complex hardware issues.

The role requires strong Linux skills, server lifecycle knowledge, and the ability to work across teams. International travel up to four times per year is expected, with visa support as needed.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Server Reliability & NPI Operations Engineer
Senior Server Reliability & NPI Operations Engineer

ByteDance • San Jose (CA)

On-site
USD 115,200 - 288,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Global SRE: Core Data Center & Service Reliability
Global SRE: Core Data Center & Service Reliability

ByteDance • San Jose (CA)

On-site
USD 162,000 - 388,000
Global Data Center Operations Supervisor
Global Data Center Operations Supervisor

Pangle • Ashburn (VA)

Hybrid
USD 86,000 - 167,000
Medical, dental & vision
401(k) with company match
Paid parental leave
+2
Graduate Backend Engineer — Global Infrastructure & Ops
Graduate Backend Engineer — Global Infrastructure & Ops

Bytedance • San Jose (CA)

On-site
USD 110,000 - 160,000
Senior GPU Infrastructure & Production Systems Engineer
Senior GPU Infrastructure & Production Systems Engineer

ByteDance • New York (NY)

On-site
USD 122,000 - 272,000
Senior SRE - Data Infrastructure & Reliability
Senior SRE - Data Infrastructure & Reliability

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Server Validation & Performance Engineer
Server Validation & Performance Engineer

Bytedance • San Jose (CA)

On-site
USD 162,000 - 388,000
Linux System Engineer – Graduate
Linux System Engineer – Graduate

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Cloud Hardware System Architect – Global Travel & AI-Ready
Cloud Hardware System Architect – Global Travel & AI-Ready

ByteDance • San Jose (CA)

On-site
USD 212,800 - 450,000
Server Fleet Engineer Intern — Linux, Automation & AI
Server Fleet Engineer Intern — Linux, Automation & AI

ByteDance • New York (NY)

On-site
USD 29,000 - 40,000
Health insurance
10 paid holidays per year
Paid sick time (56 hours / year)