AI Inference Core - Senior Technical Program Manager

Cerebras Systems

United States

Hybrid

USD 160,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid schedule
Office presence 3 days/week

Job summary

Cerebras Systems is seeking an experienced Senior Technical Program Manager to own the end-to-end operating model for AI Inference Core, linking feature delivery, release testing, infrastructure readiness and deployment. You will ensure clear ownership, exit criteria, dependencies, and evidence across cross-functional teams.

This role requires deep technical context to challenge plans, identify hidden dependencies, and drive improvements in decision quality and execution.

Qualifications

  • Extensive experience leading complex technical programs across multiple engineering teams.
  • Strong ability to understand architecture, dependencies, quality evidence, and trade-offs.

Responsibilities

  • Define how features move from development to release readiness with clear owners and criteria.
  • Own the operating mechanism for pre- and post-merge stability including health and SLA tracking.
  • Translate priorities into goals, milestones, owners, risks, and review cadences.

Skills

Technical program management
Cross-team leadership
Technical fluency
Influencing without authority
Analytical skills
Executive communication
Startup environment experience

Tools

CI/CD tooling
Observability dashboards

Job description

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.

Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.

About the Role

We are looking for an experienced Senior Technical Program Manager to create one accountable operating layer across feature delivery, release integration testing, Core Infrastructure, branch stability, and release readiness within AI Inference Core.

The operating mechanism for Inference Core — making complex technical execution explicit, measurable, and consistently followed.

You will ensure that cross-team initiatives have clear ownership, documented entry and exit criteria, visible health and SLA tracking, timely escalation, and dependable follow-through. You will turn priorities into an integrated execution model without taking technical ownership away from engineering.

This is not a project-tracking or meeting-coordination role. You will need enough technical depth to understand complex AI systems, challenge unclear plans, identify hidden dependencies, improve decision quality, and create operating mechanisms that engineering teams trust and use.

What Makes This Role Distinct
  • End-to-end operating ownership: Connect feature development, qualification, release testing, infrastructure readiness, branch health, release qualification, and deployment.

  • Mechanism—not technical outcome: Own process design, planning cadence, dashboards, SLA reporting, dependency tracking, decision logs, and escalation follow-through; engineering retains architecture and quality decisions.

  • Feature-flow clarity: Make owners, entry and exit criteria, evidence, dependencies, and exception paths explicit as work moves toward release.

  • Stability accountability: Create visibility and follow-through for pre-merge and post-merge E2E health and SLA breaches.

  • Cross-team leverage: Coordinate roadmaps, staffing, risks, and ownership boundaries across multiple engineering functions and partner teams.

  • Leadership leverage: Reduce coordination load on engineering leads so they can focus on architecture, technical strategy, infrastructure, and integration quality.

What You Will Do
  • Define how features move from development and feature qualification into release integration testing and release qualification, including owners, entry and exit criteria, required evidence, dependencies, and exception paths.

  • Own the operating mechanism for pre-merge and post-merge E2E stability: publish health, track SLA breaches, drive triage and escalation, document decisions, and close recurring-failure loops.

  • Convert Inference Core priorities into clear goals, milestones, owners, risks, success measures, and review cadences while maintaining one dependable view of commitments.

  • Coordinate planning and execution across release integration testing, Core Infrastructure, feature teams, release owners, and partner organizations; synchronize roadmaps, staffing, and cross-team dependencies.

  • Build and maintain dashboards, SLA reporting, dependency maps, risk registers, decision logs, action tracking, and executive-ready status communication.

  • Drive timely decisions and follow-through on blocked or slipping work, surfacing trade-offs and escalating when teams cannot resolve issues at the working level.

  • Continuously improve the operating model using delivery data, retrospectives, recurring-failure patterns, stakeholder feedback, and changes in business priorities.

Minimum Skills & Qualifications
  • Significant experience leading complex technical programs across multiple engineering teams, ideally in infrastructure, distributed systems, platforms, release engineering, or AI systems.

  • Strong technical fluency and the ability to understand architecture, system dependencies, quality evidence, operational risk, and engineering trade-offs.

  • Proven ability to design and implement durable operating mechanisms, not merely report status or schedule meetings.

  • Experience defining goals, milestones, ownership, entry and exit criteria, SLAs, risk management, and executive review cadences.

  • Ability to influence technical leaders and teams without direct authority while preserving clear engineering ownership.

  • Strong analytical skills and experience using metrics, dashboards, and qualitative evidence to improve execution and decision quality.

  • Exceptional written and verbal communication, including concise executive synthesis and clear escalation under ambiguity or pressure.

Preferred Skills
  • Experience with AI infrastructure, model delivery, high-performance computing, distributed systems, or hardware/software platforms.

  • Experience supporting feature integration, release qualification, branch stability, developer infrastructure, or production readiness.

  • Familiarity with software quality, E2E testing, CI/CD, observability, incident learning, and reliability mechanisms.

  • Experience coordinating teams with distinct but overlapping ownership boundaries.

  • Experience in a startup or similarly fast-moving, resource-constrained engineering environment.

  • Track record of taking an operating model or cross-team program from zero to one and scaling it as the organization grows.

  • Technical or engineering background sufficient to build credibility with senior engineers and technical leaders.

What Success Looks Like
  • Adoption: 100% of release integration testing-engaged features have documented owners, entry criteria, dependencies, exception paths, and readiness evidence.

  • Stability: Pre-merge and post-merge E2E health and SLA breaches are visible, reviewed, assigned, and driven through closure.

  • Predictability: Goals, milestones, risks, and staffing dependencies are reviewed consistently, with fewer late surprises.

  • Decision quality: Important trade-offs, decisions, owners, and escalation outcomes are documented and easy to find.

  • Accountability: Blocked or slipping commitments have clear recovery plans and timely leadership visibility.

  • Leverage: Engineering leads spend less time on coordination and more time on technical strategy, infrastructure, and integration quality.

Location
  • This role follows a hybrid schedule and requires in-office presence three days per week. Fully remote work is not available.

  • Office locations: Sunnyvale, CA or Toronto, ON.

Why Join Cerebras

People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:

  1. Build a breakthrough AI platform beyond the constraints of the GPU.

  2. Publish and open source their cutting-edge AI research.

  3. Work on one of the fastest AI supercomputers in the world.

  4. Enjoy job stability with startup vitality.

  5. Our simple, non-corporate work culture that respects individual beliefs.

Find out more about what it's like to work at Cerebras here!

Apply today and become part of the forefront of groundbreaking advancements in AI!

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Core - Software Integration Engineer
AI Inference Core - Software Integration Engineer

Cerebras • United States

Hybrid
USD 150,000 - 190,000
AI Inference Core - Software Integration Engineer
AI Inference Core - Software Integration Engineer

Cerebras Systems • United States

Hybrid
USD 170,000 - 250,000
Sr. Member of Technical Staff
Sr. Member of Technical Staff

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 120,000 - 160,000
Inclusive work environment
Job stability with startup vitality
Opportunity for continuous learning
AI Inference Core - Senior SW Engineer for Platform & DevOps
AI Inference Core - Senior SW Engineer for Platform & DevOps

Cerebras • United States

On-site
USD 180,000 - 260,000
AI Inference Core - Senior SW Engineer for Platform & DevOps
AI Inference Core - Senior SW Engineer for Platform & DevOps

Cerebras Systems • United States

On-site
USD 180,000 - 270,000
ML Systems Integration Engineer
ML Systems Integration Engineer

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
AI Inference Core - Infrastructure SW Engineer
AI Inference Core - Infrastructure SW Engineer

Cerebras Systems • United States

On-site
USD 140,000 - 190,000
AI Inference Core - SDET Technical Lead, Release Integration Testing
AI Inference Core - SDET Technical Lead, Release Integration Testing

Cerebras Systems • United States

On-site
USD 150,000 - 210,000
AI Inference Core - SDET Technical Lead, Release Integration Testing
AI Inference Core - SDET Technical Lead, Release Integration Testing

Cerebras • United States

On-site
USD 140,000 - 210,000
AI Inference Core - Infrastructure SW Engineer
AI Inference Core - Infrastructure SW Engineer

Cerebras • United States

On-site
USD 140,000 - 190,000