Capacity & Performance Management Engineer

Subway

Shelton (CT)

On-site

USD 120,000 - 180,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Subway is seeking an Engineer, Capacity & Performance Management to ensure our cloud estate (Azure, AWS, Databricks) runs reliably and cost-effectively. You will forecast demand, model capacity, tune performance, and optimize consumption using telemetry from Dynatrace, ServiceNow, and platform metrics.

You will collaborate with Data Engineering, EOC, and the ServiceNow team to translate telemetry into dashboards, recommendations, and cost-optimized strategies across reliability operations.

Qualifications

  • 4-8 years in capacity/performance engineering, cloud infrastructure, or IT operations with a data/analytics component.
  • Strong data/analytics skills including SQL, BI dashboards, trend analysis, and forecasting.
  • Experience with Databricks and monitoring workload performance and consumption.
  • Hands-on capacity planning and performance monitoring in Azure and/or AWS.
  • Observability tooling experience (Dynatrace) and turning telemetry into insight.
  • Bachelor’s in IT/CS/Data or related field, or equivalent combination.

Responsibilities

  • Own cloud capacity planning and forecasting across Azure and AWS.
  • Partner with data platform and EOC teams to optimize Databricks workload performance and cost.
  • Define and track performance baselines, KPIs, and SLAs across systems.
  • Analyze telemetry to set baselines and surface anomalies.
  • Build dashboards in Power BI and ServiceNow Performance Analytics.

Skills

SQL & BI dashboards
Databricks experience
Cloud infrastructure monitoring
Dynatrace observability
Python scripting

Education

Bachelor’s in IT, CS, Data/Analytics

Tools

Databricks
Dynatrace
ServiceNow
Power BI
Azure/AWS

Job description

Engineer, Capacity & Performance Management

Job Description

Job Code: (as applicable)

FLSA Status: Exempt

Reports to: Director, Reliability & Operations

Department: Infrastructure, Reliability & Operations

Location: Hybrid / Office (Shelton, CT) per policy

Position Summary

The Engineer, Capacity & Performance Management keeps Subway’s cloud estate performing reliably and scaling cost-effectively with business demand. That estate spans Azure and AWS infrastructure, managed SaaS platforms, and our strategic Databricks data platform, which is Subway’s standard for enterprise data and analytics and a particular focus of this role. Using operational telemetry from Dynatrace, ServiceNow, and the cloud and data platforms themselves, the engineer forecasts demand, models capacity, tunes performance, and optimizes consumption and cost.

This is a data-driven engineering role, and the data in question is operational: the analytics are of the infrastructure, platforms, and their consumption (how they perform and what they cost), not the business data stored within them. The engineer pairs strong analytics and BI skills with hands-on cloud and platform performance expertise, turning that telemetry into the forecasts, dashboards, and recommendations that guide capacity, performance, and cost decisions across Reliability & Operations. The work is done in partnership with Data Engineering (which owns Databricks), the Enterprise Operations Center (EOC), and the ServiceNow platform team; success is measured in avoided saturation, reliable performance, and optimized spend, not in operating or administering any single platform.

Essential Functions

Approximate time allocation shown per area; priorities shift with business demand.

Cloud Infrastructure Capacity Planning & Forecasting (primary focus, ~25%)

  • Own capacity planning and demand forecasting across Azure and AWS (compute, storage, and network).
  • Analyze utilization and saturation telemetry; set baselines and thresholds; model growth and project future needs.
  • Right-size cloud resources to balance performance, reliability, and cost.
  • Produce capacity forecasts and lead capacity reviews so resources are provisioned ahead of demand.

Data Platform Capacity & Performance (Databricks) (strategic, growing focus, ~15%)

  • Partner with Data Engineering (platform owner) to optimize Databricks workload performance and consumption.
  • Monitor cluster and SQL warehouse utilization, job/query performance, and DBU consumption via Databricks system tables (billable usage, query history, list prices); track spend using Databricks budgets, alerts, and prebuilt usage dashboards.
  • Surface top cost/performance offenders (long-running jobs, oversized clusters, idle warehouses) and partner on remediation such as right-sizing, warehouse scaling, and job scheduling.
  • Improve cost attribution via tagging across classic and serverless compute (resource tags and serverless usage policies); forecast DBU, compute, and storage demand; recommend compute policies, right-sizing, autoscaling, and idle shutdown (auto-termination / auto-stop).

Performance Engineering & Tuning (~15%)

  • Define and track performance baselines, KPIs, and proposed SLAs across applications, databases, and infrastructure; watch for deviation and anomalies.
  • Recommend monitoring thresholds and success criteria; partner with performance testers to validate scalability ahead of releases and demand peaks.
  • Diagnose bottlenecks across application, database, infrastructure, and network layers; recommend and validate tuning.
  • Analyze database performance across the estate (SQL Server and managed cloud databases such as Azure SQL and AWS RDS / Aurora) for query performance, execution plans, and contention, and recommend remediation to the owning teams.

Analytics, Dashboards & Reporting (the analytical core, ~20%)

  • Analyze operational telemetry from Dynatrace and the platforms’ own metrics (utilization, performance, and consumption/cost, not the business data the platforms store) to set baselines, surface anomalies, and feed capacity and performance models.
  • Query telemetry, metering, and cost data with SQL; build capacity, performance, and cost dashboards in Power BI and ServiceNow Performance Analytics.
  • Apply trend analysis and forecasting to predict saturation and inform demand planning.
  • Translate telemetry into clear narratives and executive reporting; flag emerging capacity, performance, and cost risks early.

Cost & Consumption Optimization (FinOps) (~10%)

  • Identify idle, over-provisioned, and inefficient resources across cloud and data platforms; drive right-sizing and optimization.
  • Improve cost attribution through tagging standards and enforcement; reduce untagged and unallocated spend.
  • Detect and investigate cost anomalies and usage spikes against utilization and operational events.
  • Produce showback and unit-cost views (e.g., cost per application/workload); support chargeback with Finance if adopted.
  • Recommend Reserved Instance, Savings Plan, and Azure Reservation coverage; track commitment usage and expirations via Microsoft Cost Management and AWS Cost Explorer.

Monitoring, Alerting & Incident Support (~10%)

  • Provide capacity/performance monitoring in support of the Enterprise Operations Center (EOC); recommend alerting and threshold changes to cut false positives and sharpen signal.
  • Support major incident management for capacity and performance events during business hours; contribute root-cause analysis and permanent fixes.
  • Feed recurring issues into problem management to prevent repeats.
  • Use ServiceNow (Incident, Problem, Change; ITOM / CMDB; Performance Analytics) as the system of record. (Administration not required; the ServiceNow team owns the platform.)

AI & Automation (~5%)

  • Apply AI-powered monitoring and analytics to surface signals and anomalies that manual monitoring misses.
  • Automate recurring capacity and cost reporting, alerting, and anomaly detection (light scripting, platform APIs, ServiceNow / observability integrations).
  • Evaluate and pilot AI / AIOps features within existing observability and FinOps tooling; document outcomes and recommend adoption.
Required Qualifications
  • 4-8 years in capacity/performance engineering, cloud infrastructure, or IT operations with a strong data/analytics component.
  • Strong data/analytics skills, including SQL, BI/dashboards (Power BI or equivalent), trend analysis, and forecasting, applied to infrastructure, platform, and cost telemetry rather than the business data within the platforms.
  • Experience with Databricks (or a comparable data platform): monitoring workload performance and DBU/consumption, and forecasting capacity.
  • Hands-on capacity planning and performance monitoring in a cloud environment (Azure and/or AWS).
  • Observability / APM tooling (e.g., Dynatrace) and turning telemetry into actionable insight.
  • Database performance with SQL Server and/or managed cloud databases (Azure SQL, AWS RDS / Aurora).
  • Clear communication, with the ability to translate technical capacity, performance, and cost data into decision-ready insight for engineering leaders and non-technical stakeholders.
  • Scripting for automation (Python or PowerShell).
  • Bachelor’s in Information Technology, Computer Science, Data/Analytics, or a related field, or an equivalent combination of education and experience.
Preferred Qualifications
  • Deeper Databricks skills: compute policies, Unity Catalog, and Spark tuning.
  • Cloud cost/FinOps experience: showback, tagging/allocation, Reserved Instances/Savings Plans, Microsoft Cost Management or AWS Cost Explorer; FinOps Certified Practitioner a plus.
  • Performance and load/stress testing.
  • ServiceNow (ITSM; ITOM/CMDB; Performance Analytics) and ITIL 4 capacity, problem, and service-level practices.
  • Certifications: ITIL 4 Foundation; Microsoft Azure (AZ-104) or AWS Certified CloudOps Engineer (Associate); Dynatrace (Associate/Professional); ServiceNow Certified System Administrator; Databricks Certified Data Engineer Associate or Databricks Fundamentals.
  • Integrating AI tools to optimize workflows and drive measurable impact.
Core Competencies
  • Self-starter with a healthy curiosity who takes initiative, investigates the unknowns, works independently, and makes sound decisions with minimal day-to-day direction.
  • Analytical rigor and data storytelling that turns operational telemetry into clear, decision-ready insight.
  • Strong written and executive-level communication.
  • Influence and drive outcomes without direct authority; effective cross-team collaboration.
  • Prioritization under competing demands; bias toward measurable, actionable deliverables.
  • Cost and business acumen.
Accountability / Scope

People Management: No

Direct Reports: None

Reports to: Director, Reliability & Operations

Scope: Capacity & performance across cloud infrastructure (Azure + AWS; compute, storage, network; databases) and the Databricks platform (consumption & performance, growing); analytics, alerting recommendations, and ServiceNow reporting; capacity/performance support to the EOC.

Financial Authority: No budget authority; influences cloud and Databricks spend through analysis, forecasting, and recommendations.

Decision Making: Individual contributor who makes technical and analytical decisions and recommendations within the capacity, performance, and cost-optimization domain, working with a high degree of autonomy.

Hours: Standard business hours, with occasional off-hours support for major incidents or planned capacity events.

Other: Other duties may be assigned as business needs evolve.

Travel Requirements: Minimal (less than 5%).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Capacity & Performance Management Engineer
Capacity & Performance Management Engineer

Franchise World Headquarters, LLC • Shelton (CT)

Hybrid
USD 120,000 - 150,000
Capacity & Performance Management Engineer
Capacity & Performance Management Engineer

Pho Prime, LLC • Shelton (CT)

On-site
USD 120,000 - 160,000
Staff+ Software Engineer, Capacity Engineering
Staff+ Software Engineer, Capacity Engineering

United States Digital Space LLC • San Francisco (CA)

On-site
USD 190,000 - 250,000
Cloud Capacity & Performance Engineer
Cloud Capacity & Performance Engineer

Pho Prime, LLC • Shelton (CT)

On-site
USD 120,000 - 160,000
Senior Manager Software Engineering - Cloud Engineering (Remote)
Senior Manager Software Engineering - Cloud Engineering (Remote)

The Home Depot • Atlanta (GA)

On-site
USD 180,000 - 280,000
Principal Software Engineer
Principal Software Engineer

PowerPlan, Inc. • Atlanta (GA)

On-site
USD 180,000 - 240,000
IT Engineer I
IT Engineer I

Allen Distribution • Carlisle

On-site
USD 70,000 - 90,000
Infrastructure Engineer
Infrastructure Engineer

Dexian • New York (NY)

On-site
USD 150,000 - 190,000
Lead Business Operations Data Engineer
Lead Business Operations Data Engineer

Jobtailor • Colorado

On-site
USD 140,000 - 170,000
Director of Infrastructure Operations
Director of Infrastructure Operations

HomeServe • Norwalk (CT)

On-site
USD 157,000 - 211,000
Annual bonus