Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Cerebras Systems is seeking a Principal SRE to define and drive the architecture for scaling our AI inference fleet across datacenters and cloud-based solutions. You will build self-service platforms, observability, and automation, enabling product teams and customers to operate with strong guardrails.
In the first year, you will lead a transformation from ops-centric reliability to a shared engineering discipline, mentor senior engineers, and shape capacity management and rollout safety.
We are building a high-performance SRE function to support one of the world’s fastest-growing AI inference services, powered by the Wafer-Scale Engine (WSE). This team will help deliver world-class, ultra-reliable inference infrastructure for leading model builders such as OpenAI and other frontier labs.
As a Principal SRE, you will define and drive the technical architecture for scaling our inference fleet through self-service delivery, shared observability, capacity orchestration, rollout safety, and operational automation. This role starts with 2–3 weeks of hands-on operational immersion to build deep context on the current stack, production pain points, and high-stakes workflows.
From there, your mandate shifts to architecting the “tomorrow” layer: a unified capacity management and production control plane that enables reliable capacity planning, workload placement, rollout safety, validation, and operational decision-making across large-scale inference infrastructure.
Success in the first year means core engineering teams, product managers, external customers, and cluster stakeholders can execute critical operational workflows through self-service systems with strong guardrails, clear ownership, and minimal dependency on expert SRE operators.
You will collaborate with the tech leads and the leadership team across core, cluster, cloud, and product stakeholders. This work will shift reliability from an ops-only burden to a shared engineering discipline that underpins frontier AI inference at scale.
If you are a proven Principal engineer who enjoys turning complexity into elegant reliability at scale, this is your chance to lead this transformation from the front.
This role does not require 24/7 on-call rotations.
Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.
This website or its third-party tools process personal data. For more details, click to review our CCPA disclosure notice.