AI Infrastructure Operations Engineer

Cerebras Systems

United States

On-site

USD 70,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Continuous learning and development
Diverse working environment

Job summary

Cerebras Systems is seeking an AI Infrastructure Operations Engineer to support deployment and monitoring of AI infrastructure in data centers. This entry-level position involves troubleshooting, monitoring hardware telemetry, and executing validation tests with guided support from senior engineers.

Ideal candidates should have a Bachelor's degree in a relevant engineering field and 0–3 years of experience in hardware operations or datacenter environments. Join a team that values innovation and inclusion, and play a role in shaping the future of AI.

Qualifications

  • 0–3 years experience in hardware operations, systems engineering, or datacenter environments.
  • Internship or early-career experience in datacenter or hardware lab environments.
  • Comfort working in data centers.

Responsibilities

  • Assist with deployment and bring-up of CS-X systems, cluster servers, and networking hardware.
  • Execute power-on sequencing, readiness checks, and validation tests.
  • Monitor hardware telemetry, alerts, and dashboards.
  • Perform first-line troubleshooting and structured escalation.
  • Collect logs, telemetry, and observations during incidents.

Skills

Basic familiarity with server hardware
Networking fundamentals
Linux systems

Education

Bachelor’s degree in a relevant engineering field

Tools

Monitoring and telemetry systems

Job description

About The Role

The AI Infrastructure Operations Engineer (SiteOps) is an entry‑level individual contributor role focused on the deployment, bring‑up, monitoring, and first‑line troubleshooting of Cerebras AI infrastructure in data center environments. The role supports CS systems, cluster server hardware, cluster networking hardware, and hardware telemetry and monitoring tools.

Support reliable operation and scale‑out of Cerebras AI clusters by executing defined hardware bring‑up and validation procedures, monitoring telemetry, performing first‑line troubleshooting, and escalating issues using established workflows.

Responsibilities
  • Assist with deployment and bring‑up of CS‑X systems, cluster servers, and networking hardware;
  • Execute power‑on sequencing, readiness checks, and validation tests.
  • Monitor hardware telemetry, alerts, and dashboards.
  • Perform first‑line troubleshooting and structured escalation.
  • Collect logs, telemetry, and observations during incidents.
Incident Support & Tooling
  • Participate in incident response under senior engineer guidance.
  • Use existing monitoring, telemetry, and incident tracking tools.
  • Provide feedback on tooling and process gaps.
Learning & Development
  • Build working knowledge of Cerebras system architecture.
  • Learn cluster hardware and networking fundamentals.
  • Shadow senior engineers during complex debugging.
  • Progress toward independent ownership of defined workflows.
Explicit Non‑Responsibilities
  • No people management.
  • No final escalation authority.
  • No ownership of cluster architecture, hardware design, or tooling architecture.
Required Qualifications

Bachelor’s degree in a relevant engineering field or equivalent experience; 0–3 years experience in hardware operations, systems engineering, or datacenter environments; basic familiarity with server hardware, networking fundamentals, and Linux systems.

Preferred Qualifications

Internship or early‑career experience in datacenter or hardware lab environments; exposure to monitoring or telemetry systems; comfort working in data centers.

What Success Looks Like

Consistent and correct execution of hardware bring‑up procedures, early identification and escalation of issues, improving documentation quality, and clear progression toward more independent operational responsibility.

Career Path

This role progresses naturally toward Senior and Principal IC roles within AI Infrastructure Operations (SiteOps), with an optional management track.

Why Join Cerebras

People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:

  1. Build a breakthrough AI platform beyond the constraints of the GPU.
  2. Publish and open source their cutting‑edge AI research.
  3. Work on one of the fastest AI supercomputers in the world.
  4. Enjoy job stability with startup vitality.
  5. Our simple, non‑corporate work culture that respects individual beliefs.

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

This website or its third‑party tools process personal data. For more details, please review our CCPA disclosure notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Operations Engineer
AI Infrastructure Operations Engineer

Cerebras • United States

On-site
USD 70,000 - 90,000
Inclusive work environment
Career growth opportunities
Equal opportunity employer
Software Engineer - Tools & Infrastructure / DevOps
Software Engineer - Tools & Infrastructure / DevOps

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 230,000
Site Reliability Engineer - Ops & Automation
Site Reliability Engineer - Ops & Automation

Cerebras • United States

On-site
USD 125,000 - 170,000
Site Reliability Engineer - Ops & Automation
Site Reliability Engineer - Ops & Automation

Cerebras Systems • United States

Hybrid
USD 120,000 - 150,000
Mentorship from seasoned engineers
Non-corporate work culture
Job stability with startup vitality
AI Inference Core - Senior SW Engineer for Platform & DevOps
AI Inference Core - Senior SW Engineer for Platform & DevOps

Cerebras • United States

On-site
USD 180,000 - 260,000
AI Inference Core - Senior SW Engineer for Platform & DevOps
AI Inference Core - Senior SW Engineer for Platform & DevOps

Cerebras Systems • United States

On-site
USD 180,000 - 270,000
Infrastructure Engineer (Data Center Operations)
Infrastructure Engineer (Data Center Operations)

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 120,000 - 170,000
Software Engineer - Tools & Infrastructure / DevOps
Software Engineer - Tools & Infrastructure / DevOps

Cerebras Systems • Sunnyvale (CA)

On-site
USD 140,000 - 190,000
Infrastructure Engineer (Data Center Operations)
Infrastructure Engineer (Data Center Operations)

Socket.dev • Sunnyvale (CA)

On-site
USD 120,000 - 190,000
Software Engineer, Cluster Deployment
Software Engineer, Cluster Deployment

Cerebras Systems • Sunnyvale (CA)

On-site
USD 120,000 - 180,000