ML Systems Performance Engineer

Cerebras Systems

Bengaluru

On-site

INR 2,500,000 - 4,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Breakthrough platform
Open source research
Fast AI supercomputer
Job stability
Flat culture

Job summary

Cerebras Systems, based in Bengaluru, is hiring an inference performance engineer at the intersection of hardware and software. You will drive end-to-end model inference speed on the Cerebras Wafer Scale Engine, tackling kernel performance, system-level analysis, and the development of performance tooling and diagnostics.

Join a team that builds a breakthrough AI platform, works with leading labs and enterprises, and values hands-on optimization, deep learning math exposure, and strong

Qualifications

  • Bachelor's degree or higher in Electrical Engineering or Computer Science.
  • Strong background in computer architecture and hardware-software interfaces.
  • Exposure to and understanding of low-level deep learning / LLM math.
  • Strong analytical and problem-solving mindset.

Responsibilities

  • Build performance models (kernel-level, end-to-end) to estimate the performance of state of the art and customer ML models.
  • Optimize and debug our kernel micro code and compiler algorithms to elevate ML model inference speed, throughput and compute utilization on the Cerebras WSE.
  • Debug and understand runtime performance on the system and cluster.
  • Develop tools and infrastructure to help visualize performance data collected from the Wafer Scale Engine and our compute cluster.

Skills

C++
Python
Computer Architecture
Kernel optimization
Performance profiling

Education

Bachelor's degree in Electrical Engineering or Computer Science

Tools

CPU/GPU simulators
Compiler algorithms

Job description

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.

Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.

About The Role

Engineers on the inference performance team operate at the intersection of hardware and software, driving end-to-end model inference speed and throughput. Their work spans low-level kernel performance debugging and optimization, system-level performance analysis, performance modeling and estimation, and the development of tooling for performanceprojection and diagnostics.

Responsibilities
  • Build performance models (kernel-level, end-to-end) to estimate the performance of state of the art and customer ML models.
  • Optimize and debug our kernel micro code and compiler algorithms to elevate ML model inference speed, throughput and compute utilization on the Cerebras WSE.
  • Debug and understand runtime performance on the system and cluster.
  • Develop tools and infrastructure to help visualize performance data collected from the Wafer Scale Engine and our compute cluster.
Requirements
  • Bachelors / Masters / PhD in Electrical Engineering or Computer Science.
  • Strong background in computer architecture.
  • Exposure to and understanding of low-level deep learning / LLM math.
  • Strong analytical and problem-solving mindset.
  • 3+ years of experience in a relevant domain (Computer Architecture, CPU/GPU Performance, Kernel Optimization, HPC).
  • Experience working on CPU/GPU simulators.
  • Exposure to performance profiling and debug on any system pipeline.
  • Comfort with C++ and Python.

Why Join Cerebras

People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we've reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:

  1. Build a breakthrough AI platform beyond the constraints of the GPU.
  2. Publish and open source their cutting-edge AI research.
  3. Work on one of the fastest AI supercomputers in the world.
  4. Enjoy job stability with startup vitality.
  5. Our simple, non-corporate work culture that respects individual beliefs.

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Performance Engineer
ML Systems Performance Engineer

Cerebras Systems, Inc. • India

On-site
INR 1,200,000 - 1,800,000
Full Stack LLM Engineer
Full Stack LLM Engineer

Cerebras Systems • Bengaluru

On-site
INR 2,000,000 - 3,000,000
Competitive salary and benefits package
Opportunities for professional growth
Dynamic work environment
+1
Lead Full Stack Machine Learning Engineer
Lead Full Stack Machine Learning Engineer

Cerebras Systems, Inc. • India

On-site
INR 2,000,000 - 3,000,000
Work with one of the fastest AI supercomputers
Enjoy job stability with startup vitality
Non-corporate work culture
Lead Full Stack Machine Learning Engineer
Lead Full Stack Machine Learning Engineer

Cerebras • India

On-site
INR 2,500,000 - 4,000,000
Non-corporate work culture
Job stability with startup vitality
Opportunity to publish and open source AI research
Manager kernel software
Manager kernel software

Cerebras • India

On-site
INR 3,500,000 - 5,500,000
Kernel Engineer
Kernel Engineer

Cerebras Systems, Inc. • India

On-site
INR 1,200,000 - 2,000,000
Open source cutting-edge AI research
Job stability with startup vitality
Inclusive work culture
Manager kernel software
Manager kernel software

Cerebras • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Job stability with startup vitality
Non-corporate work culture
Opportunity to publish cutting-edge AI research
ML Research Engineer (Inference)
ML Research Engineer (Inference)

Cerebras • India

On-site
INR 5,681,000 - 9,470,000
Publish and open source cutting-edge AI research
Work on one of the fastest AI supercomputers
Job stability with startup vitality
ML Research Engineer (Inference)
ML Research Engineer (Inference)

Cerebras • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Senior/Staff- Engineer: Post Silicon- Bring Up
Senior/Staff- Engineer: Post Silicon- Bring Up

Cerebras Systems, Inc. • Bengaluru

On-site
INR 16,556,000 - 26,018,000