Senior Cloud Operations Engineer

JumpMind, LLC

Columbus (OH)

On-site

USD 120,000 - 180,000

Full time

11 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

JumpMind, LLC is seeking a Senior Cloud Operations Engineer in Columbus, OH to lead enterprise cloud platform operations across AWS and related technologies. You will guide reliability, security, and optimization, mentor junior staff, and champion AI-driven automation to reduce toil.

The role partners with Cloud Engineering and Security to maintain platform standards, service objectives, and incident postmortems, while supporting audit evidence requests and on-call duties.

Qualifications

  • 5+ years in cloud operations, SRE, DevOps, or infrastructure engineering roles.
  • Deep hands-on experience with a major cloud provider (AWS preferred).
  • Strong troubleshooting skills across networking, compute, storage, and containerized workloads.
  • Experience owning incident response and writing postmortems.
  • Comfort reading and operating Terraform managed infrastructure.

Responsibilities

  • Monitor production infrastructure health, availability, and performance across cloud environments; own alerting and dashboards (uptime, latency, error rates, capacity).
  • Lead incident response for production issues — triage, coordinate, drive root-cause analysis, and own postmortems/corrective actions.
  • Operate and troubleshoot Kubernetes clusters in production — node health, pod scheduling issues, resource limits/requests, cluster upgrades, and workload scaling.
  • Manage patching, OS/dependency updates, configs, and vulnerability remediation timelines in coordination with Security.
  • Own capacity planning and scaling decisions ahead of peak retail traffic events.
  • Drive cloud cost optimization (FinOps) — rightsizing, reserved capacity, waste elimination, monthly spend reviews.
  • Manage backups, disaster recovery testing, and recovery runbooks.
  • Execute infrastructure changes defined by Cloud Engineering's IaC (Terraform) — apply, validate, and operate what Engineering builds.
  • Build and refine operational runbooks, on-call procedures, and escalation paths.
  • Serve as the technical mentor for the Junior Cloud Operations Engineers.
  • Partner with Cloud Engineering on handoffs from build to run; flag operational gaps back to Engineering.
  • Support audit and compliance evidence requests (SOC 2, PCI DSS) related to operational controls, access, and change management.
  • Participate in an on-call rotation.
  • Build and maintain AI agents and AI-assisted automations that handle operational toil.
  • Evaluate and integrate AI/LLM-powered tools into the operations workflow and drive adoption across the team.
  • Use AI coding/agent tools day-to-day to write and maintain operational scripts, IaC change validation, and internal tooling faster.

Skills

Cloud operations
SRE/DevOps
Troubleshooting
Incident response
AI integration

Tools

AWS
Terraform
CloudWatch
Datadog
Grafana
Prometheus
Kubernetes

Job description

If you are unable to complete this application due to a disability, contact this employer to ask for an accommodation or an alternative application process.

Senior Cloud Operations Engineer

Columbus, OH, US

5 days ago Requisition ID: 1061

Senior Cloud Operations Engineer Role: Responsibilities, Qualifications, and AI Integration

Senior Cloud Operations Engineer

Role summary The Senior Cloud Operations Engineer leads enterprise cloud platform operations, governance, security, reliability, and optimization across Amazon Web Services (AWS) and related cloud technologies. This person serves as the technical anchor of the CloudOps function, mentors the junior team members, and ensures cloud services remain stable, secure, scalable, cost-effective, and aligned with business priorities.

This role establishes platform standards, operational practices, service objectives, and accountability across the cloud environment. This person partners closely with Cloud Engineering, Security, and Enterprise IT to enable modernization while protecting operational stability and business continuity. AI fluency is core to this role: Jumpmind expects its CloudOps team to actively use AI tools and build lightweight agents to reduce toil, speed up diagnosis, and automate operational workflows, not just run infrastructure manually.

Day-to-day responsibilities
  • Monitor production infrastructure health, availability, and performance across cloud environments; own alerting and dashboards (uptime, latency, error rates, capacity)
  • Lead incident response for production issues — triage, coordinate, drive root-cause analysis, and own postmortems/corrective actions
  • Operate and troubleshoot Kubernetes clusters in production — node health, pod scheduling issues, resource limits/requests , cluster upgrades, and workload scaling
  • Manage patching, OS/dependency updates, configs, and vulnerability remediation timelines for infrastructure in coordination with Security
  • Own capacity planning and scaling decisions ahead of peak retail traffic events
  • Drive cloud cost optimization (FinOps) — rightsizing, reserved capacity, waste elimination, monthly spend reviews
  • Manage backups, disaster recovery testing, and recovery runbooks
  • Execute infrastructure changes defined by Cloud Engineering's IaC (Terraform) — apply, validate, and operate what Engineering builds
  • Build and refine operational runbooks, on-call procedures, and escalation paths
  • Serve as the technical mentor for the Junior Cloud Operations Engineers — pairing, code/config review, on-call shadowing
  • Partner with Cloud Engineering on handoffs from build to run; flag operational gaps (missing alerting, fragile deploy patterns) back to Engineering
  • Support audit and compliance evidence requests (SOC 2, PCI DSS) related to operational controls, access, and change management
  • Participate in an on-call rotation
  • Build and maintain AI agents and AI-assisted automations that handle operational toil — e.g., alert triage/summarization, log analysis, runbook execution, auto-remediation of known failure patterns
  • Evaluate and integrate AI/LLM-powered tools into the operations workflow (incident summarization, on-call copilots, ChatOps assistants) and drive adoption across the team
  • Use AI coding/agent tools day-to-day to write and maintain operational scripts, IaC change validation, and internal tooling faster
Required qualifications
  • 5+ years in cloud operations, SRE, DevOps, or infrastructure engineering roles
  • Deep hands-on experience with a major cloud provider (AWS preferred)
  • Strong troubleshooting skills across networking, compute, storage, and containerized workloads
  • Experience owning incident response and writing postmortems
  • Comfort reading and operating Terraform managed infrastructure (not necessarily authoring modules from scratch)
  • Experience with monitoring/observability tooling (e.g., CloudWatch, Datadog, Grafana, Prometheus)
  • Familiarity with compliance-driven environments (SOC 2, PCI DSS) and working alongside a security team
  • Excellent written communication — this role will produce runbooks, postmortems, and audit evidence
  • Hands-on experience using AI tools in daily technical work, and demonstrated experience building or configuring AI agents/automations
  • Experience leveraging AI to automate operational workflows — examples might include auto-generated incident summaries, AI-assisted root cause analysis, or agentic remediation scripts
Preferred qualifications
  • Experience operating infrastructure for a SaaS or retail/commerce platform with high-availability requirements
  • Scripting ability (Python, Bash, or Go) for operational tooling and automation
  • Prior experience mentoring or leading a junior engineer
  • Exposure to SIEM or security monitoring tooling
  • Experience with agent frameworks or protocols (e.g., MCP) and integrating AI agents with internal systems (Slack, ticketing, monitoring)
  • Familiarity with AI governance/security concepts (model access controls, data exposure risks in AI tooling)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Cloud Operations Engineer
Senior Cloud Operations Engineer

ADP, Inc. • Columbus (OH)

On-site
USD 120,000 - 150,000
On-site in Columbus, OH
AI-enabled automation initiatives
Mentor junior engineers
Cloud Operations Engineer I (Junior)
Cloud Operations Engineer I (Junior)

JumpMind, LLC • Columbus (OH)

On-site
USD 50,000 - 75,000
Cloud Operations Engineer I (Junior)
Cloud Operations Engineer I (Junior)

ADP, Inc. • Columbus (OH)

On-site
USD 65,000 - 90,000
Cloud Engineer
Cloud Engineer

Confidential Company • Atlanta (GA)

On-site
USD 120,000 - 150,000
Staff Software Engineer(Cloudops)
Staff Software Engineer(Cloudops)

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 190,000 - 230,000
Principal Dev-Ops Architect
Principal Dev-Ops Architect

Zohorecruit • United States

Remote
USD 180,000 - 240,000
Sr. Systems Engineer - AI
Sr. Systems Engineer - AI

Dairy Farmers of America • Kansas City (KS)

On-site
USD 130,000 - 170,000
Sr. Systems Engineer - AI
Sr. Systems Engineer - AI

Kansas Ag Connection • Kansas City (KS)

On-site
USD 140,000 - 190,000
Senior Platform Engineer (Cloud & AI Platform)
Senior Platform Engineer (Cloud & AI Platform)

OEC • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Mid-to-Senior Cloud SRE & DevOps Engineer
Mid-to-Senior Cloud SRE & DevOps Engineer

Mondrian Alpha • New York (NY)

On-site
USD 170,000 - 230,000