AVP SRE, Cloud Solutions

Jobtailor

Arlington (TX)

On-site

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Jobtailor seeks an experienced Site Reliability and Platform Engineering leader to drive PSRE strategy and governance across reliability, AI-enabled operations, and non-functional requirements. You’ll build resilient, self-healing systems, establish SLOs/SLIs, and partner with Cybersecurity, Architecture, and Product to reduce technology risk.

You will lead high-performing teams, champion automation, and advance observability to improve MTTR, service availability, and operational efficiency in a

Qualifications

  • Experience leading Site Reliability Engineering, Platform Engineering, Operational Excellence, or comparable technology organizations.
  • Demonstrated leadership in agentic AI, AIOps, intelligent automation, or autonomous operational capabilities, including governance, evaluation, security, human oversight, and auditability.
  • Deep understanding of observability, incident and problem management, resiliency engineering, operational readiness, technology risk, SLOs, SLIs, error budgets, and toil reduction.
  • Experience improving incident trends, recurring incidents, MTTR, service availability, alert quality, customer impact, and operational efficiency.
  • Experience with CI/CD, Infrastructure as Code (IaC), DevOps automation, policy as code, test automation, and modern software delivery practices.
  • Experience leading vulnerability remediation, technology debt reduction, non-functional requirement automation, and cloud cost optimization initiatives.
  • Strong software engineering foundation and technical judgment related to automation, APIs, debugging, performance, maintainability, and scalable solution design.
  • Demonstrated success developing engineering talent, growing future technical leaders, and building high-performing organizations focused on continuous learning and technical excellence.
  • Strong executive communication, stakeholder management, collaboration, and organizational influence skills.
  • Ability to operate effectively in a fast-paced environment, provide quality service to internal and external customers, work a flexible schedule when business needs require, and travel on a limited basis.
  • 7-10 years IT industry experience required
  • 3-5 years technical leadership in large enterprises required
  • 2-4 years early career experience as a software developer required
  • High School Diploma required
  • Bachelor’s Degree in Computer Science or equivalent work experience required
  • Experience in financial services industry preferred
  • Master’s Degree preferred

Responsibilities

  • Define and execute the PSRE strategy, roadmap, standards, and operating model aligned with business and technology objectives.
  • Lead the adoption and responsible governance of agentic AI across reliability engineering, incident management, monitoring, observability, technology risk, vulnerability and technology debt remediation, cloud cost optimization, and non-functional requirements.
  • Establish automated controls and intelligent agents that validate reliability, security, operational readiness, and non-functional requirements throughout the delivery lifecycle.
  • Drive resilient, self-healing systems that detect, diagnose, remediate, and recover from known failure conditions with safeguards, human oversight, and auditability.
  • Champion engineering practices that improve availability, scalability, recoverability, performance, maintainability, and operational sustainability while reducing toil.
  • Establish observability strategies across metrics, logs, traces, application performance, service health, and user-experience signals.
  • Drive improvements in incident trends, recurring-issue elimination, MTTR, service availability, alert quality, customer impact, and operational efficiency.
  • Establish operational governance for incident and escalation management, operational readiness, service ownership, documentation, runbooks, supportability, and service health.
  • Partner with Cybersecurity, Architecture, Product, and Engineering teams to reduce technology risk and improve vulnerability remediation, technology debt, resiliency, and engineering standards.
  • Establish SLOs, SLIs, error budgets, reliability KPIs, operational metrics, and executive reporting.
  • Build and lead high-performing teams through coaching, mentorship, career development, succession planning, and technical upskilling.
  • Establish a culture of knowledge sharing, innovation, engineering excellence, accountability, and continuous improvement.
  • Provide executive leadership during major incidents and influence technology strategy, investment decisions, and engineering priorities.

Skills

SRE Leadership
Platform Engineering
Observability
Automation
Incident Management
Code & Debugging
Cloud Cost Optimization
Non-functional Requirements
Technical Leadership
Auditing & Governance

Education

Bachelor’s Degree in Computer Science
Master’s Degree Preferred

Tools

IaC
DevOps Automation
Policy as Code
Test Automation
Modern Software Delivery

Job description

  • Define and execute the PSRE strategy, roadmap, standards, and operating model aligned with business and technology objectives
  • Lead the adoption and responsible governance of agentic AI across reliability engineering, incident management, monitoring, observability, technology risk, vulnerability and technology debt remediation, cloud cost optimization, and non-functional requirements
  • Establish automated controls and intelligent agents that validate reliability, security, operational readiness, and non-functional requirements throughout the delivery lifecycle
  • Drive resilient, self-healing systems that detect, diagnose, remediate, and recover from known failure conditions with safeguards, human oversight, and auditability
  • Champion engineering practices that improve availability, scalability, recoverability, performance, maintainability, and operational sustainability while reducing toil
  • Establish observability strategies across metrics, logs, traces, application performance, service health, and user-experience signals
  • Drive improvements in incident trends, recurring-issue elimination, MTTR, service availability, alert quality, customer impact, and operational efficiency
  • Establish operational governance for incident and escalation management, operational readiness, service ownership, documentation, runbooks, supportability, and service health
  • Partner with Cybersecurity, Architecture, Product, and Engineering teams to reduce technology risk and improve vulnerability remediation, technology debt, resiliency, and engineering standards
  • Establish SLOs, SLIs, error budgets, reliability KPIs, operational metrics, and executive reporting
  • Build and lead high-performing teams through coaching, mentorship, career development, succession planning, and technical upskilling
  • Establish a culture of knowledge sharing, innovation, engineering excellence, accountability, and continuous improvement
  • Provide executive leadership during major incidents and influence technology strategy, investment decisions, and engineering priorities
Requirements
  • Experience leading Site Reliability Engineering, Platform Engineering, Operational Excellence, or comparable technology organizations
  • Demonstrated leadership in agentic AI, AIOps, intelligent automation, or autonomous operational capabilities, including governance, evaluation, security, human oversight, and auditability
  • Deep understanding of observability, incident and problem management, resiliency engineering, operational readiness, technology risk, SLOs, SLIs, error budgets, and toil reduction
  • Experience improving incident trends, recurring incidents, MTTR, service availability, alert quality, customer impact, and operational efficiency
  • Experience building or enabling resilient, self-healing systems and automation that reduce manual intervention
  • Experience with CI/CD, Infrastructure as Code (IaC), DevOps automation, policy as code, test automation, and modern software delivery practices
  • Experience leading vulnerability remediation, technology debt reduction, non-functional requirement automation, and cloud cost optimization initiatives
  • Strong software engineering foundation and technical judgment related to automation, APIs, debugging, performance, maintainability, and scalable solution design
  • Demonstrated success developing engineering talent, growing future technical leaders, and building high-performing organizations focused on continuous learning and technical excellence
  • Strong executive communication, stakeholder management, collaboration, and organizational influence skills
  • Ability to operate effectively in a fast-paced environment, provide quality service to internal and external customers, work a flexible schedule when business needs require, and travel on a limited basis
  • 7-10 years IT industry experience required
  • 3-5 years technical leadership in large enterprises required
  • 2-4 years early career experience as a software developer required
  • High School Diploma required
  • Bachelor’s Degree in Computer Science or equivalent work experience required
  • Experience in financial services industry preferred
  • Master’s Degree preferred
Core Competencies

Demonstrates expertise in Site Reliability Engineering and Platform Engineering, with a strong focus on agentic AI, observability, and automation. Proven ability to lead high-performing teams, drive operational excellence, and establish governance frameworks that enhance reliability and reduce technology risk.

Highest-signal resume keywords
  • Site Reliability Engineering Leadership
  • Agentic AI and AIOps Expertise
  • Observability and Incident Management
  • CI/CD and DevOps Automation
  • Technical Leadership in Large Enterprises
Hard Skills
  • Automation
  • APIs
  • Debugging
  • Performance Optimization
  • Scalable Solution Design
  • Incident Management
  • Operational Readiness
  • SLOs and SLIs
  • Cloud Cost Optimization
  • Vulnerability Remediation
Soft Skills
  • Executive Communication
  • Stakeholder Management
  • Collaboration
  • Organizational Influence
  • Coaching and Mentorship
Certifications & Qualifications
  • Bachelor’s Degree in Computer Science
  • Master’s Degree Preferred
Industry Keywords
  • Financial Services Industry
  • Operational Excellence
  • Technology Risk
  • Non-Functional Requirements
  • Continuous Improvement
Tools & Technologies
  • Infrastructure as Code (IaC)
  • DevOps Automation
  • Policy as Code
  • Test Automation
  • Modern Software Delivery Practices
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer
Principal Site Reliability Engineer

Jobtailor • Arizona

On-site
USD 180,000 - 240,000
Distinguished Software Engineer – AI/ML Engineer
Distinguished Software Engineer – AI/ML Engineer

Jobtailor • Sunnyvale (CA)

On-site
USD 230,000 - 350,000
Senior Site Reliability Engineer – Digital Assets
Senior Site Reliability Engineer – Digital Assets

Jobtailor • Arizona

On-site
USD 120,000 - 170,000
Principal Platform Engineer, AI – Automation
Principal Platform Engineer, AI – Automation

Jobtailor • Phoenix (AZ)

On-site
USD 170,000 - 210,000
Senior Technology Manager – AI, Intelligent Automation Platforms
Senior Technology Manager – AI, Intelligent Automation Platforms

Jobtailor • Charlotte (NC)

On-site
USD 170,000 - 210,000
Senior Software Engineering Manager
Senior Software Engineering Manager

Jobtailor • North Carolina

On-site
USD 190,000 - 230,000
Lead Software Engineer
Lead Software Engineer

Jobtailor • Colorado

On-site
USD 150,000 - 190,000
Senior Engineering Manager – Hands-on
Senior Engineering Manager – Hands-on

Jobtailor • Santa Clara (CA)

On-site
USD 180,000 - 250,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Jersey

On-site
USD 120,000 - 180,000
Senior Manager, Staff Engineering – Software Development, Microservices
Senior Manager, Staff Engineering – Software Development, Microservices

Jobtailor • Maryland

On-site
USD 180,000 - 240,000