Senior Software Development Engineer in Test (SDET) - AI Cluster Networking and Security

Cerebras

India

On-site

INR 3,000,000 - 6,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cerebras Systems in India seeks an experienced test engineer to innovate and execute tests on the world\'s largest AI chip infrastructure. You will validate thousands of nodes in large deployments and drive ultra-high reliability with 99.999% uptime.

The role emphasizes automation across cluster software (Kubernetes, Prometheus, Grafana) and hardware components, including ML wafer-scale accelerators and data transfer paths. Strong coding in Python/Go/C++ is essential.

Qualifications

  • Bachelor's or Master's in engineering related field.
  • 10+ years testing enterprise software, distributed systems, and datacenter hardware.
  • Experience with high-speed networking infra and vendor platforms (Juniper, Arista, Cisco).
  • Strong coding skills in Python, Go, or C/C++.
  • Familiarity with OS internals, memory management, and performance.

Responsibilities

  • Innovate and execute tests on AI infrastructure, validating thousands of nodes.
  • Define optimized test strategies to ensure cluster reliability (target 99.999%).
  • Develop understanding of large-scale distributed ML training and inference.
  • Adopt an automation-first approach across cluster features and security.
  • Champion cluster security and uptime for observability.

Skills

Python
Go
C/C++
Distributed systems
Networking
Kubernetes
Docker
GDB/strace
Memory management
Security basics

Education

Bachelor's or Master's degree in engineering

Tools

Ixia/Spirent
Juniper
Arista
Cisco
AWS
Kubernetes
Docker

Job description

Overview

Cerebras Systems builds the world\'s largest AI chip, 56 times larger than GPUs, enabling industry-leading training and inference speeds. This architecture delivers over 10x faster inference than GPU-based hyperscale cloud services and transforms the user experience of AI applications through real-time iteration and enhanced compute.

Cerebras collaborates with leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras to deploy scale and transform key workloads with ultra-high-speed inference.

Responsibilities
  • Innovate and execute tests on cutting-edge AI infrastructure. Define optimized test strategies and methodologies to validate thousands of nodes in large deployments and ensure cluster reliability (target 99.999%).

  • Contribute to testing and validation as Cerebras grows and the ML community evolves. Adapt to new technologies and bring a diverse skill set to a fast-moving team.

  • Develop a deep understanding of large-scale distributed ML training and inference. Break complex distributed challenges into testable components for unit testing.

  • Adopt an automation-first approach. Strive for high automation coverage across cluster features, including high availability, failure scenarios, performance, stress, and security.

  • Champion cluster security and reliability for uptime and observability.

  • Test all AI cluster components, including cluster software (Kubernetes, Prometheus, Grafana) and hardware components (ML wafer-scale accelerators, CPU runtime nodes, interconnects, and data transfer paths).

  • Evaluate cluster networking solutions (high-speed switches, routers, and optics from multiple vendors).

  • Evaluate cluster security features, OS security, network security, cloud compliance, user access, and security certifications.

Qualifications
  • Bachelor\'s or master\'s degree in engineering (computer science, electrical, AI, data science, or related field).

  • 10+ years of experience testing enterprise software, distributed systems, datacenter hardware and software.

  • Experience in large enterprise or cloud networking infrastructure (high-speed switches, routers, firewalls).

  • Experience qualifying networking vendor platforms (e.g., Juniper, Arista, Cisco) and network test equipment (Ixia/Spirent).

  • Experience in datacenter technologies (BGP, ECN, PFC).

  • Experience testing networking security, compliance, and firewalls.

  • Strong coding skills in Python, Go, or C/C++.

  • Strong debugging skills for large distributed systems, hardware and software; familiarity with tools like gdb, strace, and networking monitors.

  • Strong understanding of operating systems internals (memory management, file systems, security basics, performance).

  • Strong understanding of datacenter layout and device performance characteristics (PCIe, networking, storage).

  • Experience with cloud technologies (AWS, Kubernetes, Docker). Monitoring tools like Grafana and Prometheus are a plus.

  • Understanding and experience with ML model training and inference is a plus.

  • Understanding of ML hardware accelerators (GPUs, custom accelerator ASICs) is a plus.

Why Join Cerebras

People who are serious about software build their own hardware. Cerebras has a breakthrough architecture unlocking new opportunities for the AI industry. With ongoing model releases and rapid growth, we offer a fast-paced, opportunity-rich environment.

Find out more about what it\'s like to work at Cerebras.

Apply today and join the forefront of groundbreaking advancements in AI.

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate diverse backgrounds, perspectives, and skills, and strive to build a work environment that supports ongoing learning and growth for all team members.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Development Engineer in Test (SDET) - AI Cluster Networking and Security
Senior Software Development Engineer in Test (SDET) - AI Cluster Networking and Security

Cerebras • Bengaluru

Hybrid
INR 5,500,000 - 9,000,000
Cloud Quality Engineer
Cloud Quality Engineer

Cerebras • Bengaluru

On-site
INR 1,500,000 - 2,100,000
IT/DevOps Engineer
IT/DevOps Engineer

Cerebras Systems • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Work on cutting-edge AI research
Job stability with startup vitality
Non-corporate work culture
Lead Full Stack Machine Learning Engineer
Lead Full Stack Machine Learning Engineer

Cerebras Systems, Inc. • India

On-site
INR 2,000,000 - 3,000,000
Work with one of the fastest AI supercomputers
Enjoy job stability with startup vitality
Non-corporate work culture
ML Systems Performance Engineer
ML Systems Performance Engineer

Cerebras Systems • Bengaluru

On-site
INR 2,500,000 - 4,500,000
Breakthrough platform
Open source research
Fast AI supercomputer
+2
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras • Bengaluru

On-site
INR 900,000 - 1,500,000
Kernel Engineer
Kernel Engineer

Cerebras Systems, Inc. • India

On-site
INR 1,200,000 - 2,000,000
Open source cutting-edge AI research
Job stability with startup vitality
Inclusive work culture
Lead Full Stack Machine Learning Engineer
Lead Full Stack Machine Learning Engineer

Cerebras • India

On-site
INR 2,500,000 - 4,000,000
Non-corporate work culture
Job stability with startup vitality
Opportunity to publish and open source AI research
Manager kernel software
Manager kernel software

Cerebras • India

On-site
INR 3,500,000 - 5,500,000
Full Stack LLM Engineer
Full Stack LLM Engineer

Cerebras Systems • Bengaluru

On-site
INR 2,000,000 - 3,000,000
Competitive salary and benefits package
Opportunities for professional growth
Dynamic work environment
+1