Senior TPM - Global Reliability

Socket.dev

Rome (GA)

Hybrid

USD 150,000 - 210,000

Full time

7 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Slack seeks a seasoned Technical Program Manager to own and mature reliability programs across incident management, infrastructure resilience, and data residency. You’ll drive cross-functional alignment with CE, CIC, and Salesforce partner teams to deliver enterprise-grade availability for agentic workloads and AI services.

You will define SLOs, implement playbooks for AI failure scenarios, and shape capacity planning and post-incident analysis.

Qualifications

  • 8+ years leading technical programs in a dynamic product or engineering organization.
  • Strong verbal and written interpersonal skills, with ability to communicate with engineers and identify risks.
  • Ability to work independently and communicate across multiple time zones.
  • Excellent organizational and interpersonal skills, with experience managing multiple teams.
  • Ability to analyze large data sets and synthesize them into stories and slides.
  • Experience with SQL for extracting large data sets into executive dashboards.
  • Track record of delivering complex technical projects with multi-functional teams.
  • 3+ years in active development and management of programs within SRE/reliability or infrastructure.
  • Proficiency in AWS Cloud offerings or similar cloud services.
  • Direct experience with incident management programs and post-incident reviews.
  • Familiarity with AI/ML infrastructure or agentic systems is a plus.

Responsibilities

  • Own incident management end-to-end from detection to post-incident review.
  • Lead transition of incident response operations across CE and CIC with alignment.
  • Establish severity frameworks, escalation paths, and communication protocols.
  • Measure and reduce incident volume and MTTR through data-driven improvements.
  • Run reliability initiatives to track, prioritize, and deliver improvements across the platform.
  • Own load management and compute services, ensuring scalable performance.
  • Lead capacity planning and load-shedding strategies with infrastructure teams.
  • Track reliability metrics (SLOs, error budgets, incident trends) for governance.
  • Define reliability standards for agentic workloads and AI services.
  • Develop incident response playbooks for AI failure scenarios.
  • Promote reliability-as-a-feature in AI product development.

Skills

Technical program management
Cross-functional collaboration
Data analysis
SQL
Problem solving

Education

Related technical degree

Tools

AWS

Job description

Job Category

Program & Project Management

Job Details
About Salesforce

Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn’t a buzzword — it’s a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all.

Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re in the right place! Agentforce is the future of AI, and you are the future of Salesforce.

Slack is seeking an experienced Technical Program Manager to own and mature our reliability programs across incident management, infrastructure resilience, and data residency. This role sits at the center of Slack's Trust pillar — you will drive the evolution of how we prevent, detect, and respond to incidents while managing critical cross-functional programs spanning compute services, load management, and enterprise compliance.

You will take ownership of our incident management and response program, including the strategic handoff of incident response operations to Salesforce's Command Incident Center (CIC). You will run our reliability initiatives review, manage programs around load and compute services, and Enterprise Key Management (EKM) — complex, multi-region programs that span infrastructure, security, legal, and go-to-market teams.

As Slack's platform evolves to support agentic workloads — AI agents operating alongside people — this role will also shape how reliability engineering adapts: ensuring observability, SLOs, and incident response frameworks account for non-deterministic, LLM-powered services with new failure modes.

You are a systems thinker who can drive alignment across engineering, forward engineering, security, and Salesforce partner teams. You thrive when given ambiguous, high-stakes programs and the mandate to bring structure to them. You have a strong understanding of enterprise-grade availability, SLOs and error budgets, and you are a relentless advocate for the customer experience.

Responsibilities
  • Incident Management & Response: own and mature Slack's incident management program end-to-end — from detection and triage through response, resolution, and post-incident review.
  • Drive the strategic transition of incident response operations across the Customer Experience (CE) team and Salesforce's Command Incident Center (CIC), including process alignment, tooling integration, runbook handoff, and cross-org training.
  • Establish and continuously improve incident severity frameworks, escalation paths, and communication protocols across Slack and Salesforce.
  • Partner with Reliability leadership to measure and reduce customer-impacting incident volume and mean time to resolution through data-driven process improvements.
  • Reliability Programs & Infrastructure Resilience: Run Slack's reliability initiatives review — the operating rhythm for tracking, prioritizing, and delivering reliability improvements across the platform.
  • Own programs around load management and compute services, ensuring Slack can absorb traffic spikes and scale gracefully under peak demand.
  • Drive capacity planning and load-shedding strategy in partnership with infrastructure engineering teams.
  • Track and report reliability and availability metrics (SLOs, error budgets, incident trends) to drive accountability and inform investment decisions.
  • Reliability for an Agentic World: Define reliability standards and SLO frameworks for agentic workloads — AI agents that are non-deterministic, long-running, and chain multiple services.
  • Develop incident response playbooks for novel AI failure scenarios: model degradation, prompt injection, cascading agent failures, and provider outages.
  • Champion reliability-as-a-feature in AI product development, ensuring agentic services meet the same enterprise-grade availability bar as core Slack.
  • Cross-Functional Leadership: Serve as the connective tissue across Service Owner Platform and infrastructure, security, infrastructure, and Salesforce partner teams to deliver trust outcomes.
  • Design policies, processes, and operating rhythms that scale with Slack's growing complexity.
  • Build and maintain program artifacts (timelines, risk registers, dependency maps, executive dashboards) to keep stakeholders aligned and informed.
Requirements
  • 8+ years leading technical programs in a dynamic product or engineering organization, with progressive scope and complexity.
  • Strong verbal and written interpersonal skills, with sufficient level of technical capability to effectively communicate with engineers and identify technical risks.
  • Ability to work independently and communicate across multiple time zones.
  • Excellent organizational and interpersonal/social skills, and experience handling activities across multiple teams.
  • Ability to analyze large data sets and synthesize them into stories and slides
  • SQL experience extracting large data sets into executive level dashboards.
  • Consistent track record of delivering complex technical projects and programs with multi-functional teams.
  • 3+ years of experience actively developing and managing programs within an SRE, reliability, or infrastructure organization.
  • Proficient in AWS Cloud offerings (or similar cloud services)
  • Direct experience with incident management programs — building or maturing severity frameworks, escalation processes, and post-incident review practices.
  • Experience with data residency, compliance, or regulated infrastructure programs spanning multiple regions or jurisdictions is a strong plus.
  • Familiarity with AI/ML infrastructure, LLM serving platforms, or agentic systems is a plus — or a demonstrated ability to rapidly develop technical fluency in emerging domains.
  • Experience navigating large-org integrations (e.g., parent company partnerships, shared incident response, cross-org tooling) is highly valued.
  • A related technical degree required.

Slack is the collaboration hub of choice for companies of all sizes, all across the world. By using Slack, they ensure that the right people are always in the loop, that key information is always at their fingertips, and new team members can get up to speed easily. With Slack, teams are better connected.

Ensuring a diverse and inclusive workplace where we learn from each other is core to Slack's values. We welcome people of different backgrounds, experiences, abilities and perspectives. We are an equal opportunity employer and a pleasant and supportive place to work.

Come do the best work of your life here at Slack.

Unleash Your Potential

When you join Salesforce, you’ll be limitless in all areas of your life. Our benefits and resources support you to find balance and be your best, and our AI agents accelerate your impact so you can do your best. Together, we’ll bring the power of Agentforce to organizations of all sizes and deliver amazing experiences that customers love. Apply today to not only shape the future — but to redefine what’s possible — for yourself, for AI, and the world.

Accommodations

If you need a reasonable accommodation during the application or the recruiting process, please submit a request via this Accommodations Request Form.

Please note that Salesforce uses artificial intelligence (AI) tools to help our recruiters assess and evaluate candidates’ resumes and qualifications throughout the recruiting process. Humans will always make any candidate selection and hiring decisions. Please see our Candidate Privacy Statement for more information about how we use your personal data and your rights, including with regard to use of AI tools and opt out options.

Posting Statement

Salesforce is an equal opportunity employer and maintains a policy of non-discrimination with all employees and applicants for employment. What does that mean exactly? It means that at Salesforce, we believe in equality for all. And we believe we can lead the path to equality in part by creating a workplace that’s inclusive, and free from discrimination. Know your rights: workplace discrimination is illegal. Any employee or potential employee will be assessed on the basis of merit, competence and qualifications – without regard to race, religion, color, national origin, sex, sexual orientation, gender expression or identity, transgender status, age, disability, veteran or marital status, political viewpoint, or other classifications protected by law. This policy applies to current and prospective employees, no matter where they are in their Salesforce employment journey. It also applies to recruiting, hiring, job assignment, compensation, promotion, benefits, training, assessment of job performance, discipline, termination, and everything in between. Recruiting, hiring, and promotion decisions at Salesforce are fair and based on merit. The same goes for compensation, benefits, promotions, transfers, reduction in workforce, recall, training, and education.

In the United States, compensation offered will be determined by factors such as location, job level, job-related knowledge, skills, and experience. Certain roles may be eligible for incentive compensation, equity, and benefits. Salesforce offers a variety of benefits to help you live well including: time off programs, medical, dental, vision, mental health support, paid parental leave, life and disability insurance, 401(k), and an employee stock purchasing program. More details about company benefits can be found at the following link: https://www.salesforcebenefits.com.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior TPM - Global Reliability
Senior TPM - Global Reliability

salesforce.com, inc. • United States

On-site
USD 140,000 - 210,000
Senior TPM - Global Reliability
Senior TPM - Global Reliability

100 Salesforce, Inc. • Georgia

Hybrid
USD 140,000 - 200,000
Time off programs
Medical coverage
Dental insurance
+6
Senior TPM - Global Reliability
Senior TPM - Global Reliability

Salesforce, Inc. • Northern (KY)

Hybrid
USD 130,000 - 190,000
Software Engineer II (Full-Stack)
Software Engineer II (Full-Stack)

Salesforce • Raleigh (NC)

On-site
USD 120,000 - 180,000
Software Engineer II (Full-Stack)
Software Engineer II (Full-Stack)

salesforce.com, inc. • Raleigh (NC)

On-site
USD 130,000 - 190,000
Software Engineering PMTS, Slack Distributed Data Service
Software Engineering PMTS, Slack Distributed Data Service

Salesforce • Seattle (WA)

On-site
USD 197,000 - 314,000
Software Engineering PMTS, Slack Distributed Data Service
Software Engineering PMTS, Slack Distributed Data Service

salesforce.com, inc. • Atlanta (GA)

On-site
USD 197,000 - 314,000
Staff Software Engineer, Full Stack - Slack Enterprise Trust
Staff Software Engineer, Full Stack - Slack Enterprise Trust

Slack • Seattle (WA)

On-site
USD 150,000 - 210,000
Staff Software Engineer, Full Stack - Slack Enterprise Trust
Staff Software Engineer, Full Stack - Slack Enterprise Trust

salesforce.com, inc. • Seattle (WA)

On-site
USD 197,000 - 314,000
Staff Software Engineer, Full Stack - Slack Enterprise Trust
Staff Software Engineer, Full Stack - Slack Enterprise Trust

salesforce.com, inc. • San Francisco (CA)

On-site
USD 237,000 - 345,000