Incident & Problem Manager - AI Operations

Sharon AI, Inc

Sydney

Hybrid

AUD 140,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Hybrid working
Birthday leave
Employee Assistance Program (EAP)
Exposure to AI technology
Learning & development
Novated leasing
Referral program
Employee of the month
Global footprint

Job summary

Sharon AI is seeking an experienced Incident & Problem Manager to join its NOC Service Management team in Australia. You will own end-to-end incident management, major incident handling, problem management and SLA governance across our AI Factory and Data Centre environments, coordinating cross-functional teams.

You will lead P1/P2 response, conduct RCA and PIR, implement corrective actions, and drive improvements through Lean Six Sigma techniques.

Qualifications

  • 5+ years' experience in service assurance, incident/major incident management and service-level governance.
  • Proven ability to lead high-priority incidents and coordinate cross-functional teams within SLA-driven environments.
  • Strong RCA/PIR experience and ITIL-based governance.

Responsibilities

  • Own end-to-end Incident and Major Incident Management, ensuring proper classification, prioritization and restoration.
  • Govern governance across technical bridges and engage L2/L3, Engineering, Build and Data Centre teams.
  • Lead P1/P2 incident coordination with clear ownership and communication cadence for leadership and customers.
  • Coordinate RCA and Post-Incident Reviews ensuring corrective actions are tracked to resolution.
  • Perform detailed analysis of incidents and recurring issues using structured problem-solving techniques.
  • Own SLA/OLA governance and drive actions to prevent breaches; monitor performance.
  • Maintain and improve incident, problem and service-level playbooks and escalation matrices.
  • Share lessons learned across Operations teams.

Skills

Incident Mgmt
Major Incident Mgmt
Problem Management
SLA Management
ITIL
Stakeholder Communication
Cross-team Coordination
Root Cause Analysis
Lean Six Sigma
Data Centre/Cloud Ops

Education

Bachelors in IT/Engineering

Job description

Sharon AI is an Australian neocloud, delivering trusted AI infrastructure organisations need to build, train and run AI at scale. We support customers across the full AI lifecycle, from training through to inference and agentic AI, drawing on a strong ecosystem of technology and co-location partners to deliver capability where it's needed. As the first neocloud to join NVIDIA's AI Compute Program, we're growing quickly, scaling our AI Factory platform to meet rising demand for advanced compute.

The Role

As Sharon AI continues to grow, we're looking for a talented Incident & Problem Manager to join our NOC Service Management team and lead the day-to-day execution of Incident Management, Major Incident Management, Problem Management and Service Level Management across our Australian and global AI Factory and Data Centre environments, reporting to our NOC Manager.

In this role, the Incident & Problem Manager will act as the primary focal point for major and ongoing incidents, coordinating customer teams, the NOC Service Desk, L2/L3 technical support, Engineering, Build, Service Operations, Data Centre Operations, SOC, Customer Success and technology vendors through to service restoration and closure. You'll own Post-Incident Reviews, Root Cause Analysis and SLA performance, working closely with our Data Analytics & Reporting team, and participate in an on-call roster to support after-hours P1 and escalated P2 incidents.

What You'll Be Doing
  • Own and execute end-to-end Incident and Major Incident Management, ensuring incidents are classified, prioritised, escalated and driven through to service restoration and closure
  • Maintain strong governance across technical bridges, assessing customer, service and business impact and rapidly engaging L2/L3, Engineering, Build, Service Operations, Data Centre and vendor teams
  • Lead P1/P2 incident coordination, establishing ownership, restoration priorities and communication cadence for leadership, stakeholders and customers
  • Coordinate and govern Root Cause Analysis and Post-Incident Reviews, ensuring corrective and preventive actions are documented, assigned and tracked to resolution
  • Perform detailed analysis of incidents, root causes and recurring issues, applying Lean Six Sigma and structured problem-solving techniques such as DMAIC, 5 Whys, Fishbone and Pareto analysis
  • Own operational governance of SLA and OLA commitments, proactively monitoring performance and driving corrective actions ahead of breach risk
  • Maintain and continuously improve Incident, Major Incident, Problem and Service Level Management processes, playbooks and escalation matrices, including periodic ticket quality audits
  • Use incident, problem and SLA insights to identify recurring risks and improvement opportunities, sharing lessons learned across L1/L2/L3 Operations
What We're Looking For
  • 5+ years' relevant experience in Service Assurance, Incident/Major Incident Management, Problem Management and Service Level Management within complex, business-critical infrastructure and operational environments
  • Proven experience leading high-priority incidents and coordinating technical teams, stakeholders, escalations and communications under SLA-driven conditions
  • Strong experience in Problem Management, RCA/PIR, corrective actions, known errors and ITIL-based operational governance
  • Ability to lead and coordinate major incidents, make timely decisions and drive actions and communications with wider teams operating across multiple time zones
  • Strong technical understanding and analytical problem-solving ability, with strong attention to detail and accuracy
  • Customer-first mindset with a clear focus on service availability, SLA performance and customer outcomes
  • Strong stakeholder communication, collaboration and active listening skills across customers, leadership and technical teams
  • Ability to remain calm, focused, composed and effective in a fast-paced, complex and highly technical environment
Nice to have:
  • Experience within Data Centre, Cloud, Infrastructure, Network or AI/HPC operational environments
  • Lean Six Sigma or other relevant Service Management / Quality certifications
  • Tertiary qualification in Information Technology, Engineering, Telecommunications, Business or a related discipline
Why Join Sharon AI?
  • Hybrid working – flexibility between our office and working from home
  • Birthday leave – take some extra time to celebrate your day
  • Employee Assistance Program (EAP) – confidential support when you need it
  • Exposure to AI and next-generation technology – work in one of the fastest-moving areas of technology
  • Learning & development – we support your career growth with approved conferences, professional memberships & courses
  • Novated leasing – a tax-effective way to finance and run your car
  • Bounty referral program – generous rewards for successfully referring new talent to Sharon AI
  • Employee of the month – recognition plus a $500 gift card
  • Growing global business – be part of an Australian technology company with an expanding international footprint
Our Values

Integrity | Innovation | Collaboration | Wellbeing | Inclusion

Due to the high volume of applications we receive, we're unfortunately not always able to provide individual feedback to unsuccessful candidates. We appreciate your understanding and want to assure you that every application will be reviewed with care.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Security Engineer
Security Engineer

Sharon AI • Sydney

Hybrid
AUD 130,000 - 170,000
Hybrid working
Birthday leave
Employee Assistance Program (EAP)
+4
Change Manager
Change Manager

Sharon AI, Inc • Sydney

Hybrid
AUD 120,000 - 160,000
Hybrid working
Birthday leave
Employee Assistance Program (EAP)
+1
Operational Tools Administrator
Operational Tools Administrator

Sharon AI, Inc • Sydney

Hybrid
AUD 90,000 - 130,000
Hybrid working
Birthday leave
Employee Assistance Program (EAP)
+3
Platform Engineer (GPUaaS – AI Neocloud)
Platform Engineer (GPUaaS – AI Neocloud)

Sharon AI, Inc • Sharon

Hybrid
AUD 120,000 - 180,000
Hybrid work option
Birthday leave
Employee Assistance Program (EAP)
+2
Data Centre Technician
Data Centre Technician

Sharon AI • Sydney

On-site
AUD 70,000 - 95,000
Senior Incident & Problem Manager
Senior Incident & Problem Manager

Sharon AI, Inc • Sydney

Hybrid
AUD 140,000 - 190,000
Hybrid working
Birthday leave
Employee Assistance Program (EAP)
+6
Incident & Change Manager
Incident & Change Manager

First Focus AU • City of Brisbane

On-site
AUD 108,000 - 132,000
Competitive pay up to $120,000 + super
Flexible hybrid working
Never Stop Growing program
+4
Incident & Change Manager
Incident & Change Manager

First Focus • City of Brisbane

Hybrid
AUD 108,000 - 132,000
Hybrid working
Uprise 1:1 coaching
Certification study days & exam fees
+5
Principal AI Architect
Principal AI Architect

NCS Group Australia • Melbourne

On-site
AUD 120,000 - 150,000
Paid parental leave
Well-being initiatives
Discounted health insurance
+1
Senior / Tech Lead AI Engineer
Senior / Tech Lead AI Engineer

Datacom • Sydney

On-site
AUD 150,000 - 210,000
Flexible hybrid work options
Parental leave
Employee assistance program
+2