We are seeking a Technical Project Manager (TPM) to lead the Tier 3 Production Support function for a critical Backend-for-Frontend (BFF) platform. This high-impact role sits at the intersection of technical leadership, production operations, incident management, multi-vendor coordination, and engineering excellence.
The ideal candidate will have strong hands-on technical expertise, proven experience managing complex production environments, and the ability to drive Agentic Solutioning and automation initiatives while coordinating multiple engineering and vendor teams.
Primary Skills: Agentic Solutioning, Technical Product Expertise, Production Support, Incident Management
Key Responsibilities
- Own the Tier 3 escalation process and serve as the final point of resolution for complex production issues.
- Lead P1/P2 incident war rooms and bridges, coordinating cross‑functional teams through resolution.
- Perform hands‑on API triage and troubleshooting using Splunk, including log analysis, request tracing, and root‑cause identification across BFF and downstream microservices.
- Drive Root Cause Analysis (RCA) and blameless Post‑Incident Reviews (PIRs).
- Track corrective actions through closure and ensure SLA compliance.
- Build and maintain Splunk dashboards for API health, error rates, and latency monitoring.
- Coordinate production support activities across multiple development pods and external vendors.
- Establish governance processes including daily standups, weekly reviews, escalation protocols, and performance tracking.
- Drive accountability for production fixes and resolution timelines.
- Manage dependencies and handoffs between platform engineering and vendor development teams.
- Track vendor performance metrics including MTTR, incident recurrence, and fix quality.
Technical Product & Stakeholder Management
- Partner with Product Management to balance production stability, technical debt, and feature delivery.
- Provide data‑driven updates to leadership on production health, risks, incidents, and capacity.
- Manage scope, priorities, schedules, resources, dependencies, and risks across multiple initiatives.
- Establish capacity allocation between production support and feature development.
- Drive structured change control and protect engineering teams from unplanned scope creep.
Operational Excellence & Continuous Improvement
- Identify recurring production issues and drive permanent corrective solutions.
- Champion automation and Agentic Solutioning to reduce manual support and incident‑management activities.
- Establish and monitor DORA metrics, incident trends, SLA performance, and team health indicators.
- Implement effective on‑call rotation and workload management practices.
- Promote engineering best practices such as canary deployments, feature flags, pre‑production validation, and automated monitoring.
- Develop and maintain operational runbooks, SOPs, knowledge bases, and support documentation.
- Lead, mentor, and guide production support teams to ensure timely issue resolution.
- Manage resource planning, shift coverage, rotations, and support requirements.
- Conduct peer reviews and promote high‑quality engineering standards.
- Facilitate knowledge transfer and support team development.
- Coordinate distributed teams across multiple time zones.
Required Qualifications
- Strong experience in Agentic Solutioning and applying modern AI/automation approaches to technical operations.
- Hands‑on technical expertise with APIs, microservices, application troubleshooting, and production environments.
- Proven experience in Tier 3 Production Support and complex incident resolution.
- Strong Incident Management, Problem Management, and Escalation Management experience.
- Experience with Splunk, log analysis, monitoring, and production diagnostics.
- Experience coordinating multiple engineering teams and external vendors.
- Strong understanding of SLA, MTTR, RCA, PIR, DORA metrics, and operational KPIs.
- Experience working with Agile methodologies and distributed teams.
- Excellent communication, stakeholder management, presentation, and conflict‑resolution skills.
- Ability to work effectively under pressure during critical production incidents.
Preferred Qualifications
- PMP, PRINCE2, CSM, or other relevant management certification.
- Knowledge of ITIL / IT Service Management standards.
- Experience supporting enterprise‑scale applications and customer‑facing platforms.
- Experience with scripting, automation, modern development tools, and emerging technologies.
- Strong understanding of software development lifecycle, release management, change management, QA, DBA, infrastructure, and application development processes.
Success Metrics
Success in this role will be measured by:
- Improved incident resolution and MTTR
- Reduction in recurring incidents through permanent fixes and automation
- Strong adherence to SLAs and engineering standards
- Effective management of critical escalations
- Improved application availability and production stability
- Vendor performance and accountability
ABOUT BRICKRED SYSTEMS
BrickRed Systems is a global leader in next‑generation technology consulting and workforce solutions, specializing in delivering high‑quality talent across digital, engineering, marketing, analytics, finance, operations, and business transformation domains. With a strong emphasis on innovation, scalability, and client success, BrickRed Systems helps organizations solve complex business challenges by providing skilled professionals across strategy, technology, creative, and operational functions. BrickRed fosters a culture of continuous learning, collaboration, and excellence, enabling professionals to contribute to high‑impact global initiatives while advancing their careers.