Senior Manager, Platform Operations

Collibra

United States

On-site

USD 168,000 - 210,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Collibra is seeking a Senior Manager of Platform Operations to lead the 24x7 incident response, SaaS customer operations, and change management for enterprise customers on the Collibra Platform. You will oversee deployment across thousands of virtual machines and drive AI-powered automation while partnering with engineering, product, and security teams.

You will report to the Senior Director of Reliability Engineering & Operations, shaping development plans and nurturing talent across two

Qualifications

  • 7+ years of experience in engineering, with at least 3+ years in a leadership or management role, overseeing incident response, release, or deployment operations for customer-facing SaaS production environments
  • Experience managing 24x7 production operations for customer-facing systems, including on-call rotation and escalation models
  • Experience operating within a FedRAMP or comparable regulated compliance environment, including continuous monitoring and audit cycles
  • Experience with cloud infrastructure at enterprise scale, including AWS, AWS GovCloud and GCP. Azure experience is a plus.
  • Experience managing distributed infrastructure fleets, virtual machines, containers, or equivalent, supporting production SaaS environments
  • Demonstrated proficiency in leveraging AI tools (e.g., Claude, Gemini, ChatGPT, Copilot) to solve real-world business challenges, drive measurable outcomes, or streamline workflows
  • A bachelor's degree or equivalent related working experience is required
  • Because this role supports the US government, it is required that this candidate be a US citizen who resides on US soil

Responsibilities

  • Own 24x7 incident response, SaaS customer operations, and change management in order to deliver consistent platform reliability for every enterprise customer
  • Direct deployment and fleet operations across thousands of virtual machines underpinning the Collibra Platform, including weekly release delivery, in order to keep customer environments healthy, current, and secure
  • Build succession and development plans across the team's two technical domains in order to create space for the team's strengths to surface and grow
  • Advance AI-powered automation across incident remediation, release delivery, and internal tooling in order to reduce manual toil and continuously modernize how the team operates
  • Partner across engineering, product, security, finance, and support organizations in order to align priorities, resolve cross-team dependencies, and maintain compliance in a regulated environment

Skills

Leadership
Incident response
Cloud infrastructure
FedRAMP/compliance
AI tooling adoption
SaaS operations
Communication

Education

Bachelor's degree

Tools

AWS
GCP
Terraform
Ansible
Kubernetes
Datadog
Prometheus/Grafana

Job description

Joining Collibra's Platform Operations team

You'll manage the team that keeps the Collibra Platform running and current for every enterprise customer, around the clock, spanning incident response, change management, release execution, and the health of the fleet that powers it.

You'll report to the Senior Director of Reliability Engineering & Operations and oversee the team directly, partnering closely with two technical leads who anchor deep expertise across the team's two domains.

This is a highly visible, business-critical function. When this team is at its best, customers don't notice it. Their environments are healthy, current, and secure. When it isn't, the whole company feels it. We're looking for someone who treats that responsibility as a craft, not just a job.

Our customers are our true north. Every alert answered, every release shipped, and every patch applied is in direct service of the enterprise customers running on this platform.

Senior Managers of Platform Operations at Collibra are responsible for
  • Own 24x7 incident response, SaaS customer operations, and change management in order to deliver consistent platform reliability for every enterprise customer
  • Direct deployment and fleet operations across thousands of virtual machines underpinning the Collibra Platform, including weekly release delivery, in order to keep customer environments healthy, current, and secure
  • Build succession and development plans across the team's two technical domains in order to create space for the team's strengths to surface and grow
  • Advance AI-powered automation across incident remediation, release delivery, and internal tooling in order to reduce manual toil and continuously modernize how the team operates
  • Partner across engineering, product, security, finance, and support organizations in order to align priorities, resolve cross-team dependencies, and maintain compliance in a regulated environment
You have
  • 7+ years of experience in engineering, with at least 3+ years in a leadership or management role, overseeing incident response, release, or deployment operations for customer-facing SaaS production environments
  • Experience managing 24x7 production operations for customer-facing systems, including on-call rotation and escalation models
  • Experience operating within a FedRAMP or comparable regulated compliance environment, including continuous monitoring and audit cycles
  • Experience with cloud infrastructure at enterprise scale, including AWS, AWS GovCloud and GCP. Azure experience is a plus.
  • Experience managing distributed infrastructure fleets, virtual machines, containers, or equivalent, supporting production SaaS environments
  • Demonstrated proficiency in leveraging AI tools (e.g., Claude, Gemini, ChatGPT, Copilot) to solve real-world business challenges, drive measurable outcomes, or streamline workflows
  • A bachelor's degree or equivalent related working experience is required
  • Because this role supports the US government, it is required that this candidate be a US citizen who resides on US soil
You are able to
  • Balance hands-on technical depth with people leadership, staying credible in an incident or a release window while building growth paths for the team around you
  • Operate with urgency under production pressure while preserving rigor and rollback discipline
  • Think strategically about how the team and its practices should evolve, not only how to run today's process
  • Communicate operational health directly and blamelessly, even when the news is difficult
  • Apply infrastructure automation and observability tooling such as Ansible, Terraform, Python, Kubernetes, and platforms like Datadog or Prometheus/Grafana
Measures of success

Within your first month, you will:

  • Build working relationships with both technical leads and the full team across the two operational domains
  • Gain fluency in current incident response, release, and change management processes, including FedRAMP audit and control obligations
  • Get current on active projects and initiatives across both domains, and assess where team development, succession planning, and operational health reporting stand today

Within your third month, you will:

  • Actively participate in incident response and release cycles as a hands-on contributor
  • Have a development and succession plan in motion for both technical domains
  • Identify at least one workflow suited for AI-driven automation and begin scoping it
  • Stand up a regular operational health reporting cadence, covering MTTR, MTTD, and change failure rate, for leadership visibility
  • Build working partnerships across engineering, product, security, finance, and support organizations

Within your sixth month, you will:

  • Show measurable progress against a formal development and succession plan across both domains
  • Have implemented at least one AI-driven improvement to incident remediation, release delivery, or internal tooling
  • Demonstrate measurable improvement in at least one operational health metric, such as MTTR or change failure rate, in at least one domain
  • Shape a documented point of view on how the team's operating model evolves for an AI-first future
Compensation for this role

The standard base salary range for this position is 168,000 - 210,000 per year. This position is not eligible for additional commission-based compensation. Salary offers are based on a combination of factors, including, but not limited to, experience, skills, and location.In addition to base salary, we offer a competitive total rewards package, including bonus potential, equity for eligible roles, a Flex Fund monthly stipend, pension/401k plans, and more.

Benefits at Collibra

Collibra recognizes and values that everyone has different needs, interests, and life goals. We built our benefits program with flexibility in mind to support you and your loved ones through a diverse range of circumstances and life events. These flexible offerings sit on a foundation of competitive compensation, health coverage, and time off. Learn more about Collibra's benefits.

We create inclusion and belonging through how we onboard, meet, connect, engage, and communicate. Learn more about diversity, equity, and inclusion at Collibra.

At Collibra, we're proud to be an equal opportunity employer. We realize the key to creating a company with a world-class culture and employee experience comes from who we hire and creating a workplace that celebrates everyone.

With this, we proudly consider qualified applicants without regard to race, color, religion, creed, gender, national origin, age, disability, veteran status, sexual orientation, pregnancy, sex, gender identity, gender expression, genetic information, physical or mental disability, HIV status, registered domestic partner status, caregiver status, marital status, veteran or military status, citizenship status or any other legally protected category. If you have a need that requires accommodation, let us know by completing our Accommodations for Applicants form.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Manager, Platform Operations
Senior Manager, Platform Operations

Collibra • Raleigh (NC)

On-site
USD 168,000 - 210,000
Bonus potential
Equity
Flex Fund monthly stipend
+1
Senior Director, IT Systems
Senior Director, IT Systems

Collibra • Raleigh (NC)

Hybrid
USD 204,000 - 255,000
Senior Director, IT Systems
Senior Director, IT Systems

Collibra • New York (NY)

Hybrid
USD 204,000 - 255,000
Director, Field Security
Director, Field Security

Collibra Inc. • New York (NY)

Hybrid
USD 224,000 - 280,000
Senior AI Engineer
Senior AI Engineer

Collibra • Town of Brussels (WI)

Hybrid
USD 128,000 - 197,000
Senior AI Engineer
Senior AI Engineer

Socket.dev • New York (NY)

Hybrid
USD 204,000 - 255,000
Engagement Manager
Engagement Manager

Collibra • United States

Hybrid
USD 115,000 - 148,000
Senior AI Engineer, Unstructured AI
Senior AI Engineer, Unstructured AI

Collibra • New York (NY)

Hybrid
USD 204,000 - 255,000
Hybrid work model
Equity
Bonus potential
+2
Senior Product Security Engineer
Senior Product Security Engineer

Collibra • United States

On-site
USD 110,000 - 150,000
Flexible benefits program
Competitive compensation
Health coverage
+1
Senior Enterprise Account Executive I
Senior Enterprise Account Executive I

Socket.dev • United States

On-site
USD 130,000 - 160,000