Overview
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. Our wafer-scale architecture provides the AI compute power of dozens of GPUs on a single chip, with the programming simplicity of a single device. This approach enables industry-leading training and inference speeds and allows machine learning users to run large-scale ML applications without managing hundreds of GPUs or TPUs.
Cerebras' current customers include global corporations across multiple industries, national labs, and top-tier healthcare systems. In January, we announced a multi-year, multi-million-dollar partnership with Mayo Clinic, underscoring our commitment to transforming AI applications in various fields. In August, we launched Cerebras Inference, the fastest Generative AI inference solution in the world, over 10 times faster than GPU-based hyperscale cloud inference services.
As an IT/DevOps engineer, you will be part of a team responsible for our engineering on-premises and cloud computing environments. You will develop and maintain the IT/DevOps infrastructure required to run our day-to-day operations and help develop and implement future strategic initiatives. A proven ability to automate day-to-day tasks is critical to the role.
Responsibilities
- Design, deploy, and maintain infrastructure automation and configuration management.
- Implement and manage CI/CD pipelines and deployment processes.
- Monitor, troubleshoot, and optimize cloud and on-premises infrastructure.
- Develop and maintain Infrastructure as Code (IaC) using tools like Terraform and Ansible.
- Respond to general IT requests and infrastructure incidents.
- Ensure security best practices are implemented across all systems.
- Collaborate with engineering and development teams to evaluate and identify optimal cloud solutions.
- Modify and improve existing systems for scalability and performance.
- Develop and maintain cloud solutions in accordance with best practices.
- Monitor network infrastructure health and performance, with hands-on ability to configure and manage switches, VLANs, and firewalls as needed.
- Configure and manage Palo Alto firewalls and Panorama management platform.
- Design and implement comprehensive networking solutions spanning on-premises datacenters and cloud environments.
- Manage Kubernetes clusters and container orchestration with deep understanding of cloud-native networking.
- Troubleshoot complex network issues across hybrid cloud and datacenter environments.
Skills & Qualifications
- Minimum 5+ years of experience.
- Master's Degree in Computer Science.
- Programming & Scripting: Python (required); PowerShell (nice to have).
- Operating Systems: Deep Linux expertise (RHEL/CentOS/Rocky) including system administration, performance tuning, and troubleshooting; advanced Linux networking, storage, and security configurations; experience with Linux kernel optimization and system-level debugging; ability to support mixed Mac/PC/Linux users.
- Automation & Infrastructure: Automation tools such as Ansible, Terraform, Packer; Kubernetes (K8s) administration and deployment; Jenkins management experience.
- Cloud & On-Premises Networking: Strong networking foundation from switch configuration and VLAN management to cloud-native networking; experience with Palo Alto firewalls and Panorama; proficiency in configuring and managing Arista/Juniper switches, routing, and datacenter networking; deep understanding of cloud networking, hybrid connectivity, and multi-cloud architectures; AWS experience including VPC, Security Groups, EC2, EFS, ASG, Route 53, CloudFormation, Lambda, RDS; AWS Professional-level certification required; ability to monitor and troubleshoot complex network infrastructure.
- Monitoring & Security: Monitoring tools such as ELK/Grafana/Zabbix; familiarity with endpoint security tools deployment and administration.
- Identity & Access Management: Experience with Microsoft Azure AD/OKTA to enable SSO across SaaS applications; Office 365 administration (Azure AD, SharePoint, OneDrive); experience with Atlassian Jira ServiceDesk/Confluence/Slack; experience with MDM (Intune).
- Security & Compliance: Security best practices and compliance frameworks; security controls across cloud and on-premises environments; cloud cost optimization and budget management; proficiency in cloud cost monitoring tools and cost control strategies.
- ML/AI Infrastructure: Experience managing machine learning infrastructure and GPU-accelerated workloads; knowledge of latest GPU hardware deployment and optimization for AI/ML applications.
Why Join Cerebras
- Build a breakthrough AI platform beyond the constraints of the GPU.
- Publish and open source cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Enjoy job stability with startup vitality.
- Our simple, non-corporate work culture that respects individual beliefs.
Read our blog: Five Reasons to Join Cerebras in 2025.
Equal Opportunity Statement: Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We strive to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.
This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.
Accommodations
If you need accommodations during the interview process, please let us know. Applicants must be legally eligible to work in the country of the job location.