Major Incident Manager, Incident Management -TikTok USDS

TikTok USDS Joint Venture

Seattle (WA)

On-site

USD 130,000 - 246,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision insurance
401(k) match
Paid parental leave
Disability coverage
Wellbeing benefits
Paid holidays

Job summary

TikTok USDS Joint Venture LLC is seeking a Major Incident Manager to lead high-severity incident response and coordinate cross-functional teams across SRE, Infrastructure, and Platform. You will drive rapid restoration, craft stakeholder communications, and enable blameless post-mortems with actionable remediation.

Join a security-focused organization protecting U.S. user data, with responsibilities spanning dashboards, automation, and escalation processes to continuously improve reliability and

Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, or a related technical field, or equivalent practical experience.
  • 4+ years of Incident Management, Production Support, or SRE/Operations in large-scale SaaS or cloud environments.
  • Experience with cloud providers (AWS, GCP, Azure) and modern infra architectures (microservices, Kubernetes).
  • Experience configuring/optimizing alerts and monitoring templates with Grafana, Splunk, Prometheus, or New Relic.
  • Ability to write scripts (Bash/Python) or use low-code tools to automate tasks and alerts.
  • Excellent written and verbal communication translating technical issues to business impact.
  • Strong knowledge of ITIL incident management or modern DevOps/SRE practices.

Responsibilities

  • Lead end-to-end response for high-severity incidents and coordinate cross-functional teams.
  • Craft clear, timely updates to stakeholders, executives, and customer-facing teams.
  • Facilitate blameless post-mortems with root-cause analysis and actionable follow-ups.
  • Create real-time dashboards showing system health, incident trends, and reliability metrics.
  • Analyze incident data to identify trends and drive reduction in MTTD/MTTR.
  • Manage and optimize incident management rotations, playbooks, and escalation paths.
  • Refine alert thresholds with SRE and development teams to reduce alert fatigue.
  • Design and build automated incident workflows and notification pipelines.
  • Provide on-call support within a 24x5x365 schedule, Sunday-Thursday 5 PM-2 AM PT.

Skills

Incident management
Cloud concepts
Alerts & monitoring
Scripting
ITIL / DevOps
Communication
Post-mortems
Automation
Dashboards

Education

Bachelor’s degree in CS/IT or related field

Tools

Grafana
Splunk
Prometheus
New Relic
PagerDuty
Opsgenie
JIRA Service Management
Tableau

Job description

Responsibilities

The USDS JV Incident Management Team (IMT) is a critical pillar within the US Tech and Product organization, dedicated to ensuring the resilience and reliability of TikTok’s U.S. Data Security infrastructure. As a security-first division, we focus on providing specialized oversight and protection for U.S. user data and the platforms that support them.

While global teams monitor overall service health, the IMT is uniquely positioned to manage the "blast radius" within the USDS environment. We act as the bridge between technical engineering teams (SRE, Infrastructure, Platform) and business stakeholders, ensuring that every major incident is handled with the precision and urgency required by our unique operating environment.

As a Major Incident Manager, you will be at the forefront of protecting the TikTok experience for millions of users. You will join a cohesive, proficient team committed to upholding the highest standards of professionalism and technical expertise, directly contributing to the safety and security of the U.S. digital ecosystem.

  • Incident Command: Lead the end-to-end response for high-severity incidents, driving rapid restoration of service while coordinating across cross-functional engineering teams.
  • Stakeholder Communication: Craft and deliver clear, timely, and accurate communication updates to internal stakeholders, executive leadership, and customer-facing teams during critical events.
  • Post-Incident Reviews: Facilitate blameless post-mortems, ensuring root causes are thoroughly identified, actionable remediation items are tracked, and lessons learned are shared globally.
  • Dashboarding & Visibility: Create and maintain real-time operational dashboards that provide high visibility into system health, incident trends, and key reliability metrics (MTTD/MTTR).
  • Continuous Improvement: Analyze incident data and operational metrics to identify systemic trends, driving initiatives to reduce Mean Time to Detect (MTTD) and Mean Time to Resolution (MTTR).
  • Team Process Evolution: Manage and optimize the incident management rotation, refining playbooks, alerting thresholds, and escalation pathways for the SRE org.
  • Alerting & Monitoring Improvements: Partner with SRE and development teams to continuously refine alert thresholds, reduce alert fatigue, and ensure high-severity pages are highly actionable.
  • Automation Engineering: Design and build automated workflows to streamline incident response, such as automated stakeholder communications, auto-remediation scripts, and tool integrations.
  • Oncall Support: The Incident Management team operates 24x5x365 with on-call rotation during weekends. This position is for the Sunday to Thursday shift, from 5 PM PT to 2 AM PT.
Qualifications
Minimum Qualifications
  • Bachelor’s degree in Computer Science, Information Technology, or a related technical field (or equivalent practical experience) with 4+ years of experience in Incident Management, Production Support, or SRE/Operations within a large-scale SaaS or cloud environment.
  • Strong understanding of cloud computing concepts (AWS, GCP, or Azure) and modern infrastructure architectures (microservices, Kubernetes, CI/CD pipelines).
  • Hands-on experience configuring and refining alerts and monitoring templates using industry-standard tools (e.g., Grafana, Splunk, Prometheus, New Relic).
  • Ability to write scripts (e.g., Bash, Python) or use low-code integration tools to automate repetitive incident management tasks and notification pipelines.
  • Exceptional verbal and written communication skills, with a track record of translating deeply technical issues into clear business impact summaries for leadership.
  • Strong working knowledge of ITIL incident management frameworks or modern DevOps/SRE incident response practices.
Preferred Qualifications
  • 5+ years of experience in Incident Management, Production Support, or SRE/Operations within a large-scale SaaS or cloud environment.
  • Deep alignment with Google SRE principles, including Error Budgets, SLA/SLO/SLI management, and automation-first mindsets.
  • Proven experience implementing ChatOps, auto-remediation workflows, or Event-Driven Ansible/Runbook automation to programmatically resolve common alerts.
  • Advanced proficiency with modern incident orchestration platforms like PagerDuty, Opsgenie, JIRA Service Management, and Slack integrations.
  • Experience using SQL, Tableau, or similar data visualization tools to build operational dashboards and report on organizational reliability metrics.
  • Relevant industry certifications such as ITIL v4, AWS/GCP Certified Professional, or Certified Incident Commander.
  • Experience coaching and training engineering teams on best practices for on-call hygiene and conducting constructive post-mortems.
About USDS

TikTok USDS Joint Venture LLC is dedicated to the safety and security of millions of Americans who create, discover, and connect with what they love on the apps we operate. The Joint Venture has been established in compliance with the Executive Order signed by President Trump on September 25, 2025. Our foundation is a comprehensive data privacy and cybersecurity program we operate under defined safeguards to protect national security and secure U.S. user data, apps and the algorithm. We safeguard the U.S. content ecosystem, holding decision-making authority for trust and safety policies and moderation. USDS Joint Venture helps ensure Americans can continue to express their creativity, discover new hobbies and interests, and build thriving communities and businesses on a global scale.

On-site presence across teams allows the company to operate with greater speed, alignment, and agility — especially in areas like real-time decision-making, team development, and integrated execution. As such, the company is shifting from a hybrid work model to a fully in-person schedule up to 5 days a week.

Why Join Us

Inspiring creativity is at the core of TikTok's mission. Our innovative product is built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and bring joy - a mission we work towards every day.

We strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. Every challenge is an opportunity to learn and innovate as one team. We're resilient and embrace challenges as they come. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

TikTok is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At TikTok, our mission is to inspire creativity and bring joy. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

USDS Reasonable Accommodation

USDS is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at https://tinyurl.com/USDS-RA

Job Information

【For Pay Transparency】Compensation Description (Annually)

The base salary range for this position in the selected city is $129960 - $246240 annually.

Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

For Los Angeles County (unincorporated) Candidates:

  1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;
  2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems;
  3. Exercising sound judgment.

Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment:

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer, Reliability Team - USDS
Senior Site Reliability Engineer, Reliability Team - USDS

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 187,000 - 360,000
Tech Lead Manager, TikTok Core Product SRE
Tech Lead Manager, TikTok Core Product SRE

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 209,000 - 438,000
Medical insurance
401(k) with company match
Paid parental leave
Site Reliability Engineer, Tech Infrastructure - USDS
Site Reliability Engineer, Tech Infrastructure - USDS

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 130,000 - 246,000
Site Reliability Engineer, Tech Infra
Site Reliability Engineer, Tech Infra

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 130,000 - 246,000
Senior Site Reliability Engineer, Compute - USDS
Senior Site Reliability Engineer, Compute - USDS

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 178,000 - 342,000
Site Reliability Engineer, Compute - USDS
Site Reliability Engineer, Compute - USDS

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 137,000 - 360,000
Trust and Safety Rapid Response Lead - USDS
Trust and Safety Rapid Response Lead - USDS

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 117,000 - 217,000
Medical insurance
Dental insurance
Vision insurance
+9
Software Engineer, AI Data Application – USDS
Software Engineer, AI Data Application – USDS

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 130,000 - 246,000
Site Reliability Engineer, Platform Responsibility - USDS
Site Reliability Engineer, Platform Responsibility - USDS

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 123,000 - 259,000
Medical Insurance
401(k) with match
Parental Leave
+4
Site Reliability Engineer, Platform Responsibility - USDS
Site Reliability Engineer, Platform Responsibility - USDS

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 130,000 - 246,000
Medical Insurance
Dental Insurance
Vision Insurance
+8