Vice President, Reliability Engineering & Technology Operations

Antares Capital LP

Chicago (IL)

On-site

USD 175,000 - 225,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical benefits
Dental benefits
Vision benefits
Disability insurance
Life insurance
401(k)
Profit sharing
Paid time off
Maven family benefits
Parental leave

Job summary

Antares Capital LP seeks a VP of Reliability Engineering & Technology Operations to lead production operations, reliability engineering, observability, and automation across cloud platforms. You will partner with Engineering, Infrastructure, Cybersecurity, Data, and Business teams to ensure systems are secure, observable, scalable, and operationally mature.

This leadership role drives AI-enabled operational capabilities, service ownership, and KPI-driven improvements, shaping the future of

Qualifications

  • 7+ years of experience in software engineering, platform engineering, cloud engineering, infrastructure engineering, DevOps, technology operations, or related technical disciplines.
  • 3-5+ years of experience leading engineering, reliability, platform, cloud, DevOps, or technology operations teams.
  • Strong experience operating and supporting applications within Microsoft Azure environments.
  • Hands-on experience with Kubernetes and container-based platforms.
  • Experience supporting distributed systems and cloud-native architectures.
  • Strong understanding of application architecture, system dependencies, and production support models.
  • Experience leading major incident response activities and outage management processes.
  • Experience implementing monitoring, observability, and alerting solutions using tools such as Datadog, Grafana, Splunk, Dynatrace, New Relic, or similar platforms.
  • Experience building automation solutions using scripting languages, APIs, orchestration tools, and workflow platforms.
  • Strong understanding of DevOps, CI/CD, release management, and software delivery practices.
  • Demonstrated ability to define, measure, and improve KPIs related to reliability, availability, operational efficiency, and service quality.
  • Excellent communication and stakeholder management skills with the ability to influence technical and business leaders.
  • Preferred: Experience building or leading Reliability Engineering, Production Engineering, DevOps, or Platform Engineering organizations.
  • Experience implementing AI-assisted operational workflows, AIOps platforms, AI agents, or intelligent automation solutions.
  • Experience with ServiceNow, Control-M, or equivalent enterprise operational platforms.
  • Experience operating within highly regulated environments.
  • Financial services experience, including asset management, lending, banking, private credit, or investment management.

Responsibilities

  • Lead Reliability Engineering function; define strategies to improve reliability, resiliency, scalability, and operational excellence.
  • Oversee production operations; establish support models, ownership boundaries, and service management processes.
  • Provide Azure Cloud & Kubernetes operations leadership; optimize reliability and capacity.
  • Be executive incident leader during major production events and coordinate cross-functional teams.
  • Champion automation and AI-enabled operations; drive self-healing capabilities.
  • Mentor high-performing teams and build partnerships across Engineering, Infrastructure, Data, and Business units.
  • Manage vendor relationships and ensure service commitments.
  • Ensure disaster recovery and business continuity readiness; develop recovery automation.
  • Define and report operational KPIs; drive continuous improvement across the org.

Skills

Leadership
Reliability engineering
Cloud computing
Automation
Observability
Incident management

Tools

Azure
Kubernetes
Datadog
Grafana
Splunk
Dynatrace
New Relic
ServiceNow
Control-M

Job description

About Antares Capital Antares Capital is a leading alternative credit manager and a trusted financing partner to private equity sponsors and middle-market companies.

We are committed to building resilient, scalable, and modern technology platforms that support our business and clients.

As part of our continued technology transformation, we are seeking a Vice President, Reliability Engineering & Technology Operations to lead the evolution of our production operations, reliability engineering, observability, and operational automation capabilities.

This is a strategic leadership role for an engineering-minded leader who thrives at the intersection of software engineering, cloud infrastructure, platform operations, and operational excellence.

The ideal candidate combines strong technical depth with exceptional execution skills and has experience building highly reliable systems while leading teams through modernization and transformation initiatives.

The Opportunity Technology is central to Antares' growth strategy. We are investing heavily in cloud platforms, engineering excellence, AI-enabled workflows, automation, and modern operational practices.

As the leader of Reliability Engineering & Technology Operations, you will be responsible for the availability, performance, scalability, and resilience of critical business platforms.

You will partner closely with Engineering, Infrastructure, Cybersecurity, Data, and Business stakeholders to ensure our systems remain secure, observable, scalable, and operationally mature.

You will help shape the future of technology operations by introducing reliability engineering practices, expanding observability, leveraging AI-driven operational capabilities, and reducing operational overhead through automation.

This role is ideal for someone who has grown through engineering, cloud, platform, infrastructure, or DevOps leadership roles and understands how to bridge engineering and operations to deliver exceptional business outcomes.

Key Responsibilities
Reliability Engineering Leadership

Reliability Engineering Leadership: Establish and lead Antares' Reliability Engineering function. Define and implement strategies that improve system reliability, resiliency, scalability, and operational excellence. Partner with engineering teams to embed reliability practices throughout the software development lifecycle. Drive adoption of modern operational practices including service ownership, operational readiness reviews, SLOs, error budgets, and post-incident learning.

Technology Operations Lead

Technology Operations Lead: Lead production operations across critical business applications and technology platforms. Establish clear support models, escalation paths, ownership boundaries, and service management processes. Oversee operational readiness, release support, change management, and platform health. Continuously improve operational maturity through metrics, automation, process simplification, and engineering collaboration.

Azure Cloud & Kubernetes Operations

Azure Cloud & Kubernetes Operations: Provide technical leadership for cloud-based platforms running in Microsoft Azure. Partner with Infrastructure and Engineering teams to optimize reliability, scalability, and operational efficiency. Support containerized workloads and Kubernetes-based environments. Drive best practices around cloud architecture, capacity planning, platform resilience, security, governance, and cost optimization. Ensure cloud platforms are designed and operated to meet business continuity and availability objectives.

Incident Response & Problem Management

Incident Response & Problem Management: Serve as the executive incident leader during major production events. Coordinate cross-functional teams during outages and high-severity incidents. Manage communications with business stakeholders and technology leadership. Establish disciplined root cause analysis processes and ensure corrective actions are executed. Drive long-term reduction in recurring incidents and operational risk.

AI-Powered Operations & Automation

AI-Powered Operations & Automation: Champion an automation-first and AI-enabled approach to technology operations. Identify opportunities to leverage AI for incident triage, alert correlation, knowledge management, operational analytics, and runbook execution. Partner with engineering teams to develop intelligent automation and self-healing capabilities. Evaluate emerging AIOps, agentic AI, and automation technologies and drive adoption where appropriate. Reduce manual operational effort through scripting, workflow automation, orchestration platforms, and AI-assisted tooling.

Strategic Delivery & Organizational Leadership

Strategic Delivery & Organizational Leadership: Lead and mentor high-performing technology operations and reliability engineering teams. Build strong partnerships across Engineering, Infrastructure, Cybersecurity, Architecture, Data, and Business teams. Translate operational challenges into actionable roadmaps and measurable initiatives. Drive accountability, execution excellence, and continuous improvement across the organization. Present operational trends, risks, recommendations, and performance metrics to technology leadership.

Vendor & Partner Management

Vendor & Partner Management: Manage strategic relationships with technology vendors and managed service providers. Ensure vendor accountability against service commitments and contractual obligations. Lead vendor escalations during service-impacting events. Integrate third-party support processes into internal operational workflows.

Resiliency & Disaster Recovery

Resiliency & Disaster Recovery: Partner with Engineering and Infrastructure teams to ensure systems meet recovery objectives. Improve operational readiness for disaster recovery and business continuity events. Support development of recovery automation, replay capabilities, and resilient platform architectures. Promote documentation and operational knowledge sharing to reduce dependency on tribal knowledge.

Qualifications Required

7+ years of experience in software engineering, platform engineering, cloud engineering, infrastructure engineering, DevOps, technology operations, or related technical disciplines.

3-5+ years of experience leading engineering, reliability, platform, cloud, DevOps, or technology operations teams.

Strong experience operating and supporting applications within Microsoft Azure environments.

Hands-on experience with Kubernetes and container-based platforms.

Experience supporting distributed systems and cloud-native architectures.

Strong understanding of application architecture, system dependencies, and production support models.

Experience leading major incident response activities and outage management processes.

Experience implementing monitoring, observability, and alerting solutions using tools such as Datadog, Grafana, Splunk, Dynatrace, New Relic, or similar platforms.

Experience building automation solutions using scripting languages, APIs, orchestration tools, and workflow platforms.

Strong understanding of DevOps, CI/CD, release management, and software delivery practices.

Demonstrated ability to define, measure, and improve KPIs related to reliability, availability, operational efficiency, and service quality.

Excellent communication and stakeholder management skills with the ability to influence technical and business leaders.

Preferred: Experience building or leading Reliability Engineering, Production Engineering, DevOps, or Platform Engineering organizations.

Experience implementing AI-assisted operational workflows, AIOps platforms, AI agents, or intelligent automation solutions.

Experience with ServiceNow, Control-M, or equivalent enterprise operational platforms.

Experience operating within highly regulated environments.

Financial services experience, including asset management, lending, banking, private credit, or investment management.

The Fine Print

Must have unrestricted authorization to work in the United States.

Must be willing to comply with pre-employment screening, including but not limited to drug testing, reference verification, background check.

Must be willing to work from the Chicago or New York office. #LI-hybrid

A reasonable estimate of the current base salary range at the time of posting is below.

Base salary does not include other forms of compensation or benefits.

Actual base salary within the specified range is comprised of several components, including but not limited to applicant's skill, prior relevant experience, specific degrees and certifications, job responsibilities, market considerations and the location of the position.

This role is eligible for a discretionary annual bonus (based on company, business unit and individual performance).

Our benefit offerings include medical, dental and vision coverage, employer paid short & long-term disability and life insurance, 401(k), profit sharing, paid time off, Maven family & fertility benefit, parental leave (including adoption, surrogacy, and foster placement), as well as other voluntary benefits.

Base Salary Range $175,000 - $225,000

To learn more, visit www.antares.com.

Antares is an Equal Opportunity Employer.

Employment decisions are made without regard to race, color, religion, national or ethnic origin, sex, sexual orientation, gender identity or expression, age, disability, protected veteran status or other characteristics protected by law.

To learn more about Antares, please visit http://www.antares.com.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Solutions Architect, Technology Operations & Service Delivery
Solutions Architect, Technology Operations & Service Delivery

Antares Capital LP • Chicago (IL)

On-site
USD 160,000 - 225,000
Medical, dental, and vision coverage
401(k) and profit sharing
Parental leave
Vice President, Lead Engineer - Infrastructure Operations
Vice President, Lead Engineer - Infrastructure Operations

Antares Capital LP • Illinois

On-site
USD 175,000 - 190,000
Enterprise Software Engineer
Enterprise Software Engineer

re-zoo-me • Torrance (CA), Northern (KY)

Hybrid
USD 135,000 - 160,000
Staff-Principal Reliability Engineer
Staff-Principal Reliability Engineer

Antares • Los Angeles (CA)

On-site
USD 175,000 - 250,000
Enterprise Software Engineer
Enterprise Software Engineer

Antares • Los Angeles (CA)

On-site
USD 135,000 - 160,000
Enterprise Software Engineer
Enterprise Software Engineer

53 Stations • Los Angeles (CA)

On-site
USD 120,000 - 190,000
Senior-Staff Reliability Engineer
Senior-Staff Reliability Engineer

Antares • Los Angeles (CA)

On-site
USD 140,000 - 210,000
Solutions Architect, Technology Operations & Service Delivery
Solutions Architect, Technology Operations & Service Delivery

Antares Capital LP. • New York (NY)

On-site
USD 160,000 - 225,000
Medical, dental, and vision coverage
401(k) and profit sharing
Paid time off
+2
Analyst, Loan Administration
Analyst, Loan Administration

Antares Capital LP • Atlanta (GA), Northern (KY)

Hybrid
USD 75,000 - 90,000
Discretionary annual bonus
Medical, dental and vision coverage
401(k) profit sharing
+2
Assistant Vice President, Fund Accounting Technology Engineer
Assistant Vice President, Fund Accounting Technology Engineer

Antares Capital LP • Chicago (IL)

On-site
USD 135,000 - 165,000
Medical coverage
Dental and vision coverage
Employer paid short & long-term dis­能力
+5