Technical Lead-Cloud & Infra Engg

Birlasoft

New Jersey

On-site

USD 140,000 - 190,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Birlasoft is seeking a Technical Lead-Cloud & Infra Engg to guide SRE/DevOps efforts, drive automation, and manage middleware, CDN, and incident response. You will lead a team across incident management, change control, and production reliability while collaborating with product, QA, and development teams.

The role emphasizes leadership, governance, and hands-on optimization of Java EE environments, edge platform management, and comprehensive monitoring to ensure production resilience.

Qualifications

  • Proven experience in Site Reliability Engineering (SRE) and cloud/infrastructure environments.
  • Strong leadership and mentoring abilities for technical teams.
  • Experience with incident management, RCA, and ITIL-aligned processes.

Responsibilities

  • Ensure high availability, performance, and resilience of production systems.
  • Implement SRE best practices: error budgets, SLIs/SLOs, capacity planning, chaos testing, runbook creation.
  • Drive automation to reduce manual operational tasks and improve MTTR.
  • Conduct post‑incident reviews (PIRs) and implement long‑term corrective actions.
  • Manage deployments, rollbacks, and environment synchronization across Dev, QA, UAT, and Production.
  • Install, configure, upgrade, and maintain WebLogic, Tomcat, Apache, and Nginx servers; perform JVM tuning, thread pool optimization, connection pool management, and performance tuning.
  • Troubleshoot middleware issues related to memory leaks, thread contention, SSL, certificates, and clustering.
  • Configure dashboards, alerts, and performance insights using New Relic and Splunk.
  • Develop log‑based monitoring strategies and anomaly detection.
  • Implement proactive monitoring to reduce downtime and improve reliability.
  • Configure Akamai caching rules, WAF policies, edge redirects, and performance optimizations.
  • Troubleshoot CDN‑related latency, caching, and routing issues.
  • Collaborate with Akamai support for advanced troubleshooting.
  • Lead major incident bridges, coordinate cross functional teams, and provide timely updates.
  • Manage problem tickets, root cause analysis, and preventive action plans.
  • Ensure compliance with ITIL processes for change, release, and incident management.
  • Lead and mentor a team of SRE/DevOps engineers.
  • Provide technical guidance, training, and performance feedback.
  • Act as a customer facing technical SME for escalations and production issues.
  • Collaborate with product, QA, development, and business teams to ensure smooth delivery.
  • Ownership & Accountability: Takes responsibility for production stability and issue resolution.
  • Leadership: Guides team members, manages workload, and drives operational excellence.
  • Communication: Clear, structured communication with customers and internal teams.
  • Problem Solving: Strong analytical skills and ability to troubleshoot complex issues.
  • Collaboration: Works effectively across engineering, QA, product, and business teams.
  • Calm Under Pressure: Handles critical incidents with composure and clarity.

Job description

Select how often (in days) to receive an alert:

Title: Technical Lead-Cloud & Infra Engg

Description:

Area(s) of responsibility
Site Reliability Engineering (SRE)
  • Ensure high availability, performance, and resilience of production systems.
  • Implement SRE best practices: error budgets, SLIs/SLOs, capacity planning, chaos testing, runbook creation.
  • Drive automation to reduce manual operational tasks and improve MTTR.
  • Conduct post‑incident reviews (PIRs) and implement long‑term corrective actions.
Middleware & Application Platform Management
  • Manage deployments, rollbacks, and environment synchronization across Dev, QA, UAT, and Production.
  • Install, configure, upgrade, and maintain WebLogic, Tomcat, Apache, and Nginx servers. Perform JVM tuning, thread pool optimization, connection pool management, and performance tuning.
  • Troubleshoot middleware issues related to memory leaks, thread contention, SSL, certificates, and clustering.
Monitoring, Logging & Observability
  • Configure dashboards, alerts, and performance insights using New Relic and Splunk.
  • Develop log‑based monitoring strategies and anomaly detection.
  • Implement proactive monitoring to reduce downtime and improve reliability.
CDN & Edge Platform Management (Akamai)
  • Configure Akamai caching rules, WAF policies, edge redirects, and performance optimizations.
  • Troubleshoot CDN‑related latency, caching, and routing issues.
  • Collaborate with Akamai support for advanced troubleshooting.
Incident, Problem & Change Management
  • Lead major incident bridges, coordinate cross functional teams, and provide timely updates.
  • Manage problem tickets, root cause analysis, and preventive action plans.
  • Ensure compliance with ITIL processes for change, release, and incident management.
Leadership & Stakeholder Management
  • Lead and mentor a team of SRE/DevOps engineers.
  • Provide technical guidance, training, and performance feedback.
  • Act as a customer facing technical SME for escalations and production issues.
  • Collaborate with product, QA, development, and business teams to ensure smooth delivery.
Behavioral Competencies
  • Ownership & Accountability: Takes responsibility for production stability and issue resolution.
  • Leadership: Guides team members, manages workload, and drives operational excellence.
  • Communication: Clear, structured communication with customers and internal teams.
  • Problem Solving: Strong analytical skills and ability to troubleshoot complex issues.
  • Collaboration: Works effectively across engineering, QA, product, and business teams.
  • Calm Under Pressure: Handles critical incidents with composure and clarity.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Quality Engineering Manager
Quality Engineering Manager

mtb • Buffalo (NY)

On-site
USD 90,000 - 130,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

On-site
USD 120,000 - 150,000
Technical Engineer (Production Liability Engineer)
Technical Engineer (Production Liability Engineer)

mtb • Buffalo (NY)

On-site
USD 110,000 - 150,000
Full-Time Lead Site Reliability Engineer
Full-Time Lead Site Reliability Engineer

TSP talent • O’Fallon (MO)

On-site
USD 120,000 - 190,000
Lead SRE
Lead SRE

JPMorgan Chase & Co. • Plano (TX)

On-site
USD 150,000 - 190,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Ll Oefentherapie • Reston (VA)

On-site
USD 50,000 - 70,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Plano (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer (SRE) – Production Services
Site Reliability Engineer (SRE) – Production Services

TechDigital Group • Pittsburgh

On-site
USD 120,000 - 160,000