Sr NOC Engineer NOC SRO Scripting Observability AWS Monitoring Tools

Vertafore

Hyderabad

Hybrid

INR 900,000 - 1,300,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Vertafore is seeking a Senior Reliability Operations Engineer to own day-to-day production operations, service continuity, and restoration excellence across AWS, hybrid data centers, and customer-hosted environments. You will partner with SRE, engineering, security, and support teams to ensure observability, resilience, and operational maturity of critical services.

As part of the Global Command Center, you will lead incident management, runbooks, and change governance, while driving automation,

Qualifications

  • 3.5 - 5 years of hands-on experience in Production Operations, Service Reliability Operations, Infrastructure Operations, NOC, or related roles.
  • Experience with observability, monitoring, alert tuning, dashboards on modern platforms.
  • Hands-on experience with AWS, Kubernetes, CI/CD pipelines, and hybrid environments.
  • Strong knowledge of Linux/Windows systems and relational databases.
  • Exposure to ITIL or GCC/shared-services models is preferred.

Responsibilities

  • Own day-to-day service health, monitoring, triage, escalation, and restoration coordination across cloud and hybrid environments.
  • Lead operational readiness reviews for new services, releases, and infrastructure changes with runbooks and escalation paths.
  • Tune monitoring and dashboards; drive telemetry and instrumentation standards with SRE/engineering.
  • Track SLO/SLA, error budgets, and KPI trends; recommend guardrails when thresholds are breached.
  • Identify and address recurring operational failure patterns to drive preventive controls.

Skills

Incident management
SRE collaboration
Automation & scripting
Communication

Education

Bachelor's degree in CS/Information Systems

Tools

Datadog
Splunk
Grafana
CloudWatch
Dynatrace

Job description

Job Description

We are seeking a Senior Reliability Operations Engineer to own day-to-day production operations, service continuity, operational readiness, and restoration excellence for critical production services. This role is responsible for monitoring effectiveness, incident coordination, production change execution, operational governance, runbook maturity, automation adoption, and continuous operational improvement across AWS, hybrid data centers, and customer-hosted environments.

As part of the Global Command Center (GCC), this role partners with SRE, engineering, platform, security, product, support, and business teams to ensure services are observable, supportable, resilient, and operationally mature. The role focuses on operational execution and service reliability outcomes while partnering with engineering on systemic reliability improvements.

Key Responsibilities
Service Reliability & Production Operations
  • Own operational execution for day-to-day service health, including monitoring, triage, escalation, restoration coordination, operational risk tracking, and service continuity across cloud and hybrid environments.
  • Lead operational readiness reviews for new services, releases, infrastructure changes, and production onboarding; ensure support models, runbooks, dashboards, alerts, escalation paths, rollback procedures, access, and validation checks are ready before handoff.
  • Operate and tune monitoring, alerting, dashboards, and operational observability practices aligned with the Four Golden Signals; partner with SRE and engineering on telemetry and instrumentation standards.
  • Track SLO attainment, SLA risk, error budget burn, operational KPIs, recurring incidents, and reliability trends; elevate service risk and recommend operational guardrails when thresholds are breached.
  • Monitor service performance, infrastructure utilization, capacity, saturation, and customer-impacting degradation to proactively identify and escalat operational risk.
Operational Excellence & Automation
  • Maintain an operational toil backlog and reduce repetitive manual work through automation, scripting, AI-assisted operations, self-healing workflows, AIOps capabilities and process simplification.
  • Plan, execute, coordinate, and validate production changes including patching, certificate renewals, software releases, infrastructure updates, and maintenance using standardized change governance, risk assessment, rollback readiness, stakeholder communication, and post-change verification.
  • Investigate and troubleshoot complex production issues across applications, infrastructure, middleware, databases, and platform services; restore service while partnering with SRE/engineering on permanent remediation.
  • Develop and improve runbooks, SOPs, escalation frameworks, recovery procedures, production validation checks and knowledge documentation to ensure consistent support execution.
  • Identify recurring operational failure patterns and partner with SRE, engineering, platform, and application teams to drive preventive controls and permanent corrective actions.
Incident Management, Reporting & GCC Collaboration
  • Act as the operational incident lead during major incidents, managing bridges, stakeholder communications, escalation paths, restoration coordination, event timelines, and post-incident follow-up.
  • Facilitate blameless post-incident reviews and own corrective-action tracking for remediation items, runbook updates, automation opportunities, preventive controls, and operational improvements.
  • Own incident, problem, and service reliability trend reporting, operational dashboards, SLA/SLO tracking, error budget burn visibility and leadership reviews.
  • Support the GCC operational model by driving globally standardized operational practices, governance, runbook maturity, escalation frameworks, and service management processes.
  • Collaborate with globally distributed SRE, engineering, cloud operations, security, support, product, and business teams to align operational priorities with reliability risks and service-continuity needs.
  • Mentor junior engineers and promote knowledge sharing while maintaining a customer-first mindset focused on service reliability, responsiveness, operational quality, and business continuity.
Qualifications
  • 3.5 - 5 years of hands-on experience in Production Operations, Service Reliability Operations, Infrastructure Operations, NOC, or related operational engineering roles.
  • Proven experience managing production environments with accountability for operational stability, incident response, and service continuity.
  • Strong understanding of operational readiness, incident management, production support, operational governance, change management, problem management, and reliability practices.
  • Practical experience with observability, monitoring, alert tuning, dashboarding and platforms such as Datadog, Splunk, Grafana, CloudWatch, Dynatrace, or similar tools.
  • Experience using SLO/SLA metrics, error budget burn, incident trends, and operational KPIs to manage service risk and drive continuous improvement.
  • Hands-on experience supporting AWS, Kubernetes, CI/CD pipelines, infrastructure platforms, and hybrid environments.
  • Strong knowledge of Linux and Windows systems, application hosting platforms, middleware, and relational databases.
  • Familiarity with automation and scripting using PowerShell, Python, Bash, or similar technologies.
  • Exposure to ITIL, operational governance, globally distributed support, or GCC/shared-services models is preferred.
  • Strong communication, stakeholder management, collaboration, and problem-solving skills; bachelor's/master's degree in computer science, information systems, or equivalent experience; participation in an on-call rotation for 24x7 GCC support is required.
Why Vertafore is the place for you: Canada Only
  • The opportunity to work in a space where modern technology meets a stable and vital industry
  • Medical, vision & dental plans
  • Life, AD&D
  • Short Term and Long Term Disability
  • Pension Plan & Employer Match
  • Maternity, Paternity and Parental Leave
  • Employee and Family Assistance Program (EFAP)
  • Education Assistance
  • Additional programs - Employee Referral and Internal Recognition
Why Vertafore is the place for you: US Only
  • The opportunity to work in a space where modern technology meets a stable and vital industry
  • We have a Flexible First work environment! Our North America team members use our offices for collaboration, community and team-building, with members asked to sometimes come into an office and/or travel depending on job responsibilities. Other times, our teams work from home or a similar environment.
  • Medical, vision & dental plans
  • PPO & high-deductible options
  • Health Savings Account & Flexible Spending Accounts Options:
  • Health Care FSA
  • Dental & Vision FSA
  • Dependent Care FSA
  • Commuter FSA
  • Life, AD&D (Basic & Supplemental), and Disability
  • 401(k) Retirement Savings Plain & Employer Match
  • Supplemental Plans - Pet insurance, Hospital Indemnity, and Accident Insurance
  • Parental Leave & Adoption Assistance
  • Employee Assistance Program (EAP)
  • Education & Legal Assistance
  • Additional programs - Tuition Reimbursement, Employee Referral, Internal Recognition, and Wellness
  • Commuter Benefits (Denver)

The selected candidate must be legally authorized to work in the United States.

The above statements are intended to describe the general nature and level of work being performed by people assigned to this job. They are not intended to be an exhaustive list of all the job responsibilities, duties, skill, or working conditions. In addition, this document does not create an employment contract, implied or otherwise, other than an "at will" relationship.

Vertafore strongly supports equal employment opportunity for all applicants regardless of race, color, religion, sex, gender identity, pregnancy, national origin, ancestry, citizenship, age, marital status, physical disability, mental disability, medical condition, sexual orientation, genetic information, or any other characteristic protected by state or federal law.

The Professional Services (PS) and Customer Success (CX) bonus plans are a quarterly monetary bonus plan based upon individual and practice performance against specific business metrics. Eligibility is determined by several factors including: start date, good standing in the company, and actives status at time of payout.

The Vertafore Incentive Plan (VIP) is an annual monetary bonus for eligible employees based on both individual and company performance. Eligibility is determined by several factors including: start date, good standing in the company, and actives status at time of payout.

Commission plans are tailored to each sales role but common components include quota, MBO's and ABPMs. Salespeople receive their formal compensation plan within 30 days of hire.

Vertafore is a drug free workplace and conducts preemployment drug and background screenings.

We do not accept resumes from agencies, headhunters or other suppliers who have not signed a formal agreement with us.

We want to make sure our recruiting process is accessible for everyone. if you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please contact recruiting@vertafore.com

Just a note, this contact information is for accommodation requests only.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. Technical Lead, Software Engineer (Java, React, Oracle, AI skills)
Sr. Technical Lead, Software Engineer (Java, React, Oracle, AI skills)

Vertafore • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Medical plans
Education assistance
Employee referral program
Sr. Information Security Analyst (SOC/SIEM (Splunk, CrowdStrike, Scripting)
Sr. Information Security Analyst (SOC/SIEM (Splunk, CrowdStrike, Scripting)

Vertafore • Hyderabad

On-site
INR 2,500,000 - 4,500,000
Medical plans
Dental plans
Pension Plan & Employer Match
+2
Database Architect(Oracle, PostgreSQL, AWS, Performance Tuning, Terraform)
Database Architect(Oracle, PostgreSQL, AWS, Performance Tuning, Terraform)

Vertafore • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Medical, vision & dental plans
Life, AD&D and disability insurance
Pension Plan with employer match
+1
QA Analyst II Manual Testing And AI Knowledge
QA Analyst II Manual Testing And AI Knowledge

Vertafore • Hyderabad

On-site
INR 600,000 - 1,200,000
Implementation Engineer II MS-SQL Developer RDBMS T-SQL
Implementation Engineer II MS-SQL Developer RDBMS T-SQL

Vertafore • Ahmedabad District

On-site
INR 8,637,000 - 11,516,000
Medical, vision & dental plans
Life, AD&D
401(k) Retirement Savings Plan
+3
Implementation Engineer II (MS-SQL Developer, RDBMS,T-SQL)
Implementation Engineer II (MS-SQL Developer, RDBMS,T-SQL)

Vertafore • Ahmedabad District

On-site
INR 600,000 - 900,000
Implementation Engineer II
Implementation Engineer II

Vertafore Career Center • Hyderabad

On-site
INR 900,000 - 1,200,000
Application Testing and Business Process Analyst
Application Testing and Business Process Analyst

Vertiv Energy Pvt Ltd • Thane

On-site
INR 600,000 - 1,000,000
Principal Engineer Software Engineering XIII
Principal Engineer Software Engineering XIII

Vertiv • Maharashtra

On-site
INR 2,800,000 - 5,200,000
Sr. Technical Lead, Software Engineer (OOPs, C#, JavaScript, ASP.NET, MVC, WebAPI/Rest API, Angular / React Js, Design patte)
Sr. Technical Lead, Software Engineer (OOPs, C#, JavaScript, ASP.NET, MVC, WebAPI/Rest API, Angular / React Js, Design patte)

Vertafore Career Center • Hyderabad

On-site
INR 1,800,000 - 3,200,000