Senior Cloud Platform & Reliability Architect

AT&T

City of Middletown (NY)

On-site

USD 155,000 - 261,000

Full time

10 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement program
Paid Time Off and Holidays
Paid Parental Leave
Adoption Reimbursement
Life and AD&D insurance
Employee discounts

Job summary

AT&T is seeking a Principal Systems Engineer to lead the reliability, security, and scalability of enterprise platforms and cloud services. You will design and evolve critical applications and integrations, enabling event-driven architectures with Kafka and IXBUS across cloud and on‑prem environments.

You will mentor engineers, drive automation, establish governance, and own disaster recovery, CI/CD, and observability initiatives.

Qualifications

  • Deep experience designing, operating, and governing enterprise-scale cloud platforms and infrastructure.
  • 7+ years of cloud platforms, infrastructure services, platform engineering, and cloud governance practices.
  • Experience with containerized workloads, orchestration platforms, microservices architectures, and cloud-native deployment patterns.
  • Strong understanding of CI/CD pipelines, release automation, artifact management, infrastructure automation, and deployment governance.
  • Experience with monitoring, alerting, observability, logging, service health, and operational telemetry strategies.
  • Strong knowledge of security controls, access governance, role-based access management, separation of duties, and compliance practices.
  • Experience leading platform upgrades, lifecycle management, migration planning, disaster recovery, and resiliency improvements.
  • Strong troubleshooting skills across infrastructure, application platforms, integrations, networking, and operational workflows.
  • Ability to define standards for SOPs, runbooks, operational playbooks, technical documentation, and support readiness.
  • Working knowledge of cost optimization, capacity planning, license management, and technology lifecycle planning.

Responsibilities

  • Serve as the senior technical leader and primary platform steward for critical enterprise platforms and cloud services.
  • Define and influence platform strategy, technical direction, architecture alignment, and long-term modernization roadmaps.
  • Lead cross-functional initiatives across Engineering, Operations, Architecture, Cybersecurity, Support, and business stakeholders.
  • Establish and promote engineering standards, operational best practices, automation principles, and platform governance models.
  • Own platform reliability, scalability, resiliency, lifecycle planning, operational readiness, and business continuity posture.
  • Lead compliance, cybersecurity, access governance, risk-management, and audit-readiness initiatives.
  • Drive automation-first approaches across infrastructure, deployments, operational workflows, support processes, and self-service capabilities.
  • Provide strategic oversight for CI/CD pipeline governance, release standards, artifact management, infrastructure automation, and deployment maturity.
  • Lead observability, monitoring, alerting, logging, and service health strategies to improve operational visibility and reduce incident impact.
  • Act as the highest-level escalation point for complex technical issues, major incidents, and platform-impacting events.
  • Lead root cause analysis, problem management, corrective action planning, and prevention of recurring issues.
  • Guide capacity planning, cloud cost optimization, licensing strategy, technology lifecycle management, and platform sustainability.
  • Own disaster recovery strategy, planning, exercises, documentation, and continuous improvement.
  • Mentor senior engineers and technical leads, helping raise the engineering maturity of the broader organization.
  • Influence decisions without direct authority by building alignment, communicating tradeoffs, and driving consensus.

Skills

Cloud platforms
Containerization
CI/CD pipelines
Observability/monitoring
Security/compliance
Troubleshooting
Leadership by influence
Incident response

Tools

Kafka
IXBUS

Job description

AT&T is seeking a Principal Systems Engineer to lead the reliability, security, and scalability of enterprise platforms and cloud services. You will design and evolve critical applications and integrations, enabling event-driven architectures with Kafka and IXBUS across cloud and on‑prem environments.

You will mentor engineers, drive automation, establish governance, and own disaster recovery, CI/CD, and observability initiatives.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Reliability Engineer – Cloud & Kafka
Senior Platform Reliability Engineer – Cloud & Kafka

AT&T • Dallas (TX)

On-site
USD 155,000 - 261,000
Medical/Dental/Vision coverage
401(k) plan
Paid Time Off and Holidays
+6
Senior Platform & Cloud Reliability Architect
Senior Platform & Cloud Reliability Architect

AT&T • Plano (TX)

On-site
USD 155,000 - 261,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement
+5
Senior Platform Architect – Cloud, Kafka & Automation
Senior Platform Architect – Cloud, Kafka & Automation

AT&T • Atlanta (GA)

On-site
USD 155,000 - 261,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement program
+11
Senior Platform Engineer – Enterprise Cloud & Reliability
Senior Platform Engineer – Enterprise Cloud & Reliability

AT&T • Middletown (NJ)

On-site
USD 155,000 - 261,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement program
+5
Senior Platform Reliability Architect
Senior Platform Reliability Architect

AT&T • Alpharetta (GA)

On-site
USD 155,000 - 261,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement
+7
Principal Systems Engineer: Enterprise Platform & Cloud
Principal Systems Engineer: Enterprise Platform & Cloud

AT&T • Bedminster Township (NJ)

On-site
USD 155,000 - 261,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement program
+7
Senior Cloud Platform Architect & Automation Leader
Senior Cloud Platform Architect & Automation Leader

The Depository Trust & Clearing Corporation (DTCC) • Jersey City (NJ)

Hybrid
USD 150,000 - 190,000
Base pay and annual incentive
Health and life insurance
Pension / Retirement benefits
+2
Senior Cloud Platform Engineer - Automation & Reliability
Senior Cloud Platform Engineer - Automation & Reliability

Alliant Energy Corp Serv Inc • United States

On-site
USD 103,000 - 141,000
Lead Cloud-Native Platform Engineer (Kubernetes)
Lead Cloud-Native Platform Engineer (Kubernetes)

AT&T • Plano (TX)

On-site
USD 141,000 - 237,000
Medical/Dental/Vision coverage
401(k) plan
Paid Time Off and Holidays
+1
Senior Cloud Reliability & Observability Engineer
Senior Cloud Reliability & Observability Engineer

Socket.dev • Charlotte (NC), Northern (KY)

Hybrid
USD 140,000 - 190,000
Medical insurance
Dental insurance
Vision insurance
+2