Distributed Systems Engineer – AI Supercomputing Clusters

Cerebras

United States

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equal opportunity employer
Inclusive work environment
Opportunities for continuous learning
Non-corporate work culture

Job summary

Cerebras is looking for a skilled engineer to automate hardware and software workflows in large AI supercomputers. The candidate will work on ambient cluster configurations, resource allocation, and monitoring capabilities, contributing to cutting-edge AI technology.

The ideal individual has a solid background in software architecture and development using Kubernetes and distributed systems. Join a forward-thinking team committed to diversity and equal opportunity in the workplace.

Qualifications

  • Strong track record of software architecture, system design and development.
  • Strong track record of development in distributed cluster.
  • Strong understanding of Kubernetes (K8s) software ecosystem.
  • Development skills in GoLang, Python, and bash.
  • Strong debugging skills with distributed systems.

Responsibilities

  • Automate configuration of networking, OS, and applications on large clusters.
  • Implement workflows for upgrades, downgrades, and security patching.
  • Create orchestration and scheduler system for job management.
  • Support both on-premise and cloud deployment.
  • Monitor and handle resource failures in clusters.
  • Develop user-facing tools for job status and metrics.
  • Create administrator tools for managing large clusters.

Skills

Software architecture
System design
Distributed cluster development
Kubernetes ecosystem
Prometheus
Grafana
GoLang
Python
Bash
Debugging distributed systems

Job description

Cerebras is looking for a skilled engineer to automate hardware and software workflows in large AI supercomputers. The candidate will work on ambient cluster configurations, resource allocation, and monitoring capabilities, contributing to cutting-edge AI technology.

The ideal individual has a solid background in software architecture and development using Kubernetes and distributed systems. Join a forward-thinking team committed to diversity and equal opportunity in the workplace.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed Systems Engineer – AI Supercomputing Clusters
Distributed Systems Engineer – AI Supercomputing Clusters

Cerebras Systems • United States

On-site
USD 90,000 - 120,000
Job stability with startup vitality
Open-source AI research opportunities
Non-corporate work culture
Distributed Systems Engineer: AI Cluster Orchestration
Distributed Systems Engineer: AI Cluster Orchestration

Cerebras • Raleigh (NC)

On-site
USD 100,000 - 140,000
Non-corporate work culture
Job stability with startup vitality
Opportunities for continuous learning
Distributed Software Engineer
Distributed Software Engineer

Cerebras • Raleigh (NC)

On-site
USD 100,000 - 140,000
Non-corporate work culture
Job stability with startup vitality
Opportunities for continuous learning
Distributed Software Engineer
Distributed Software Engineer

Cerebras • United States

On-site
USD 120,000 - 160,000
Equal opportunity employer
Inclusive work environment
Opportunities for continuous learning
+1
Senior AI Systems Runtime Engineer
Senior AI Systems Runtime Engineer

Cerebras Systems • United States

On-site
USD 100,000 - 130,000
Equal opportunity employer
Continuous learning and growth opportunities
Distributed Software Engineer
Distributed Software Engineer

Cerebras Systems • United States

On-site
USD 90,000 - 120,000
Job stability with startup vitality
Open-source AI research opportunities
Non-corporate work culture
Senior Distributed AI Runtime Engineer
Senior Distributed AI Runtime Engineer

Cerebras • United States

On-site
USD 120,000 - 160,000
Job stability with startup vitality
Opportunity to work on advanced AI platforms
Open source AI research
Senior Runtime Engineer for Scalable AI Systems
Senior Runtime Engineer for Scalable AI Systems

Cerebras • Raleigh (NC)

On-site
USD 140,000 - 230,000
Compute Platform Architect — AI Cluster Scale
Compute Platform Architect — AI Cluster Scale

Cerebras • United States

On-site
USD 150,000 - 200,000
Equity options
Open-source contributions
Inclusive work culture
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras Systems • Sunnyvale (CA)

On-site
USD 180,000 - 240,000