Senior Cloud Operations Engineer

ADP, Inc.

Columbus (OH)

On-site

USD 120,000 - 150,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

On-site in Columbus, OH
AI-enabled automation initiatives
Mentor junior engineers

Job summary

ADP, Inc. in Columbus, OH is seeking a Senior Cloud Operations Engineer to lead enterprise cloud platform operations across AWS, with a focus on reliability, security, and cost optimization. You will mentor junior engineers, develop AI-assisted automation, and partner with Cloud Engineering, Security, and IT to modernize while preserving stability.

You will own incident response, runbooks, on-call coordination, and ensure audit evidence readiness for SOC 2 and PCI DSS compliance.

Qualifications

  • 5+ years in cloud operations, SRE, DevOps, or infrastructure engineering roles.
  • Deep hands-on experience with a major cloud provider (AWS preferred).
  • Strong troubleshooting skills across networking, compute, storage, and containerized workloads.
  • Experience owning incident response and writing postmortems.
  • Comfort reading and operating Terraform managed infrastructure.
  • Experience with monitoring/observability tooling (CloudWatch, Datadog, Grafana, Prometheus).
  • Familiarity with SOC 2, PCI DSS and security collaboration.
  • Excellent written communication for runbooks, postmortems, and audits.
  • Hands-on experience using AI tools in daily technical work.
  • Experience using AI to automate operational workflows.

Responsibilities

  • Monitor production infrastructure health, availability, and performance across cloud environments; own alerting and dashboards.
  • Lead incident response for production issues — triage, root-cause analysis, and postmortems.
  • Operate and troubleshoot Kubernetes clusters in production — node health, scheduling, resources, upgrades, scaling.
  • Manage patching, OS/dependency updates, configs, and vulnerability remediation with Security.
  • Own capacity planning and scaling decisions ahead of peak events.
  • Drive cloud cost optimization and monthly spend reviews.
  • Manage backups, disaster recovery testing, and runbooks.
  • Execute infrastructure changes defined by IaC (Terraform).
  • Build and refine operational runbooks, on-call procedures, escalation paths.
  • Mentor Junior Cloud Operations Engineers; pair programming and code/review.
  • Flag operational gaps to Engineering and coordinate handoffs from build to run.
  • Support audit and compliance evidence requests (SOC 2, PCI DSS).
  • Participate in an on-call rotation.
  • Build AI agents and AI-assisted automations to reduce toil.
  • Evaluate and integrate AI/LLM-powered tools into operations workflow.

Skills

Cloud Ops
SRE/DevOps
AWS
Troubleshooting
Incident response
Terraform
Monitoring
SOC 2/PCI DSS
Communication
AI tools
Runbooks

Tools

CloudWatch
Datadog
Grafana
Prometheus
Terraform

Job description

If you are unable to complete this application due to a disability, contact this employer to ask for an accommodation or an alternative application process.

Senior Cloud Operations Engineer

Columbus, OH, US

2 days ago Requisition ID: 1061

Senior Cloud Operations Engineer Role: Responsibilities, Qualifications, and AI Integration
Senior Cloud Operations Engineer

Role summary The Senior Cloud Operations Engineer leads enterprise cloud platform operations, governance, security, reliability, and optimization across Amazon Web Services (AWS) and related cloud technologies. This person serves as the technical anchor of the CloudOps function, mentors the junior team members, and ensures cloud services remain stable, secure, scalable, cost-effective, and aligned with business priorities.

This role establishes platform standards, operational practices, service objectives, and accountability across the cloud environment. This person partners closely with Cloud Engineering, Security, and Enterprise IT to enable modernization while protecting operational stability and business continuity. AI fluency is core to this role: Jumpmind expects its CloudOps team to actively use AI tools and build lightweight agents to reduce toil, speed up diagnosis, and automate operational workflows, not just run infrastructure manually.

Day-to-day responsibilities
  • Monitor production infrastructure health, availability, and performance across cloud environments; own alerting and dashboards (uptime, latency, error rates, capacity)
  • Lead incident response for production issues — triage, coordinate, drive root-cause analysis, and own postmortems/corrective actions
  • Operate and troubleshoot Kubernetes clusters in production — node health, pod scheduling issues, resource limits/requests , cluster upgrades, and workload scaling
  • Manage patching, OS/dependency updates, configs, and vulnerability remediation timelines for infrastructure in coordination with Security
  • Own capacity planning and scaling decisions ahead of peak retail traffic events
  • Drive cloud cost optimization (FinOps) — rightsizing, reserved capacity, waste elimination, monthly spend reviews
  • Manage backups, disaster recovery testing, and recovery runbooks
  • Execute infrastructure changes defined by Cloud Engineering's IaC (Terraform) — apply, validate, and operate what Engineering builds
  • Build and refine operational runbooks, on-call procedures, and escalation paths
  • Serve as the technical mentor for the Junior Cloud Operations Engineers — pairing, code/config review, on-call shadowing
  • Partner with Cloud Engineering on handoffs from build to run; flag operational gaps (missing alerting, fragile deploy patterns) back to Engineering
  • Support audit and compliance evidence requests (SOC 2, PCI DSS) related to operational controls, access, and change management
  • Participate in an on-call rotation
  • Build and maintain AI agents and AI-assisted automations that handle operational toil — e.g., alert triage/summarization, log analysis, runbook execution, auto-remediation of known failure patterns
  • Evaluate and integrate AI/LLM-powered tools into the operations workflow (incident summarization, on-call copilots, ChatOps assistants) and drive adoption across the team
  • Use AI coding/agent tools day-to-day to write and maintain operational scripts, IaC change validation, and internal tooling faster
Required qualifications
  • 5+ years in cloud operations, SRE, DevOps, or infrastructure engineering roles
  • Deep hands-on experience with a major cloud provider (AWS preferred)
  • Strong troubleshooting skills across networking, compute, storage, and containerized workloads
  • Experience owning incident response and writing postmortems
  • Comfort reading and operating Terraform managed infrastructure (not necessarily authoring modules from scratch)
  • Experience with monitoring/observability tooling (e.g., CloudWatch, Datadog, Grafana, Prometheus)
  • Familiarity with compliance-driven environments (SOC 2, PCI DSS) and working alongside a security team
  • Excellent written communication — this role will produce runbooks, postmortems, and audit evidence
  • Hands-on experience using AI tools in daily technical work, and demonstrated experience building or configuring AI agents/automations
  • Experience leveraging AI to automate operational workflows — examples might include auto-generated incident summaries, AI-assisted root cause analysis, or agentic remediation scripts
Preferred qualifications
  • Experience operating infrastructure for a SaaS or retail/commerce platform with high-availability requirements
  • Scripting ability (Python, Bash, or Go) for operational tooling and automation
  • Prior experience mentoring or leading a junior engineer
  • Exposure to SIEM or security monitoring tooling
  • Experience with agent frameworks or protocols (e.g., MCP) and integrating AI agents with internal systems (Slack, ticketing, monitoring)
  • Familiarity with AI governance/security concepts (model access controls, data exposure risks in AI tooling)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cloud Operations Engineer I (Junior)
Cloud Operations Engineer I (Junior)

ADP, Inc. • Columbus (OH)

On-site
USD 65,000 - 90,000
Staff Software Engineer(Cloudops)
Staff Software Engineer(Cloudops)

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 190,000 - 230,000
Senior DevOps Engineer
Senior DevOps Engineer

P\\S\\L Group • United States

Remote
USD 120,000 - 160,000
Senior Cloud Engineer
Senior Cloud Engineer

Compunnel, Inc. • McLean (VA)

On-site
USD 120,000 - 150,000
Cloud Ops Engineer II
Cloud Ops Engineer II

Signet Jewelers • Irving (TX)

On-site
USD 110,000 - 150,000
Cloud Engineer
Cloud Engineer

Confidential Company • Atlanta (GA)

On-site
USD 120,000 - 150,000
Senior Cloud Engineer
Senior Cloud Engineer

Pointwest Technologies Corp • Roseville (MN)

On-site
USD 120,000 - 150,000
Sr. Systems Engineer - AI
Sr. Systems Engineer - AI

dfa • Kansas City (KS)

On-site
USD 150,000 - 190,000
Cloud & AI Platform Engineer
Cloud & AI Platform Engineer

Dice • Atlanta (GA), Northern (KY)

Hybrid
USD 83,000 - 152,000
Senior Cloud DevOps Engineer
Senior Cloud DevOps Engineer

JRN Associates • New York (NY)

On-site
USD 120,000 - 150,000