Overview
About Us
Working across the globe, V2X builds smart solutions designed to integrate physical and digital infrastructure from base to battlefield. We bring 120 years of successful mission support to improve security, streamline logistics, and enhance readiness. Aligned around a shared purpose, our $4.5B company and 16,000 people work alongside our clients, here and abroad, to tackle their most complex challenges with integrity, respect, responsibility, and professionalism.
Responsibilities
What You'll Do:
- The AI-Enabled Manager, Monitoring & Workload Health is a technical leadership role responsible for the continuous observability, performance, and resilience of the enterprise IT ecosystem. Reporting to the Director of Adaptive Infrastructure Network & Security, this leader will drive the transition from reactive infrastructure monitoring to proactive, self-healing operations. By leveraging agentic AI alongside enterprise monitoring platforms such as Microsoft System Center Operations Manager (SCOM) and SolarWinds Orion, the Manager will optimize the health of cloud workloads, on-premises infrastructure, network devices, and SaaS applications. This role is critical to establishing a resilient AIOps environment where predictive failure analysis and automated remediations eliminate operational downtime and reduce manual toil for the engineering teams.
Key Responsibilities:
- AIOps Strategy & Advanced Monitoring
- Design, deploy, and manage enterprise monitoring solutions, specifically focusing on optimizing SolarWinds Orion and Microsoft SCOM across a hybrid IT environment.
- Integrate network telemetry, system logs, and application performance data into a centralized AIOps platform to support intelligent operations and anomaly detection.
- Define and enforce monitoring baselines, thresholds, and intelligent alerting rules that prioritize business-critical SaaS applications and AI workloads.
- Develop observability strategies that provide end-to-end visibility into complex infrastructure paths, including Zero Trust network segments and multi-cloud environments.
- Predictive Failure & Agentic Remediation
- Implement agentic AI workflows to identify degradation patterns and execute predictive failure analysis before system outages occur.
- Build and maintain automated remediation runbooks utilizing tools like Ansible, Terraform, and Python to allow AI agents to execute self-healing actions on infrastructure and network devices.
- Establish strict governance and human-in-the-loop oversight thresholds for agent-driven automated remediations, ensuring safe and compliant execution within the production environment.
- Lead post-incident reviews (RCA) leveraging AI-generated timelines and telemetry data to continuously tune predictive algorithms and prevent recurring issues.
- Workload Health & Performance Optimization
- Monitor AI model inference traffic, LLM API calls, and agent-to-agent communication, ensuring the underlying infrastructure meets required Quality of Service (QoS) and latency SLAs.
- Collaborate with infrastructure engineering and cloud teams to right-size compute workloads based on automated capacity planning and performance trend analysis.
- Develop and deliver performance-based KPIs to IT leadership, highlighting system uptime, mean time to remediate (MTTR), and the effectiveness of automated resolution rates.
- Team Leadership & Continuous Improvement
- Manage the day-to-day tasking, performance, and effectiveness of a team of monitoring engineers and AIOps developers.
- Foster a culture of automation-first thinking, guiding the team to codify operational policies and build "Compliance-as-Code" and "Self-healing" network capabilities.
- Manage enterprise vendor and supplier agreements related to monitoring platforms, ensuring technical debt is minimized and platforms are optimized for value.
Qualifications
Minimum Requirements
- Education:
- Bachelor’s degree in Information Technology, Computer Science, Systems Engineering, or a related field; OR an equivalent combination of education and experience from which comparable knowledge and job skills can be obtained. (One year related experience may be substituted for one year of education, if degree is required).
- Experience:
- 5-10 years of experience designing and operating complex enterprise monitoring, infrastructure, or reliability engineering solutions.
- Leadership:
- Proven experience as a manager or team leader overseeing technical personnel and large-scale automation projects.
- Other Requirements:
- Required Skills & Competencies
- Deep technical expertise in administering and scaling Microsoft SCOM and SolarWinds Orion in enterprise environments.
- Strong proficiency in network automation, Infrastructure-as-Code (IaC), and scripting languages (PowerShell, Python, YAML, JSON) to drive automated remediation workflows.
- Experience integrating AIOps tools and agentic AI models with ITSM platforms (e.g., ServiceNow) for closed-loop incident resolution.
- Understanding of modern cloud architectures (Azure, AWS), virtualization (VMware, Hyper‑V), and Zero Trust networking principles.
- Exceptional collaboration skills to partner with service desk, security, and application teams to build a cohesive, automated operational fabric.
What We Bring
- At V2X we strive to be market competitive in our total reward offerings.
- The successful candidate’s starting pay will be based on, but not limited to, their job related skills, experience, qualifications, work location, and market conditions.
- The following salary range is intended to display the value of the company’s base pay compensation and may be modified at the discretion of the company.
- USD $ 140,000 - 200,000
- Provided salary range minimum and maximum values correspond to variances between regional/geographic locations across the United States.
- Employee benefits include the following:
- Healthcare coverage
- Retirement plan
- Life insurance, AD&D, and disability benefits
- Wellness programs
- Paid time off, including holidays
- Learning and Development resources
- Employee assistance resources
- Pay and benefits are subject to change at any time and may be modified at the discretion of the company, consistent with the terms of any applicable compensation or benefit plans.
At V2X, we are deeply committed to both equal employment opportunity, including protection for Veterans and individuals with disabilities, and fostering an inclusive and diverse workplace. We ensure all individuals are treated with fairness, respect, and dignity, recognizing the strength that comes from a workforce rich in diverse experiences, perspectives, and skills. This commitment, aligned with our core Vision and Values of Integrity, Respect, and Responsibility, allows us to leverage differences, encourage innovation, and expand our success in the global marketplace, ultimately enabling us to best serve our clients.