Senior Cloud Operations Engineer

PTC

San Ramon (CA)

Hybrid

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health benefits
401(k) with employer match
Commuter benefits
Learning & career growth

Job summary

ServiceMax is seeking a Senior Cloud Operations Engineer, AI Ops (Cloud) to help run and optimize cloud infrastructure across AWS in a hybrid setting out of San Ramon, CA. You will manage Kubernetes clusters, automate infrastructure with Terraform/Ansible, and collaborate with security and development teams to improve CI/CD and observability.

The role emphasizes AI-assisted incident triage, root-cause analysis, and reducing toil, with on-call rotation and capacity planning.

Qualifications

  • 3–5 years of experience in cloud operations, DevOps, or infrastructure engineering.
  • Ability to commute to the office in San Ramon, CA 2 days a week.
  • Strong understanding of AWS cloud services.
  • Experience working in Linux/Unix environments.
  • Familiarity with container technologies (Docker and/or Kubernetes).
  • Exposure to infrastructure-as-code tools (Terraform, Ansible, or similar).
  • Understanding of CI/CD concepts and tools.
  • Familiarity with monitoring and logging tools (CloudWatch, ELK, Prometheus, Grafana).

Responsibilities

  • Support the operation and maintenance of cloud infrastructure across AWS environments.
  • Deploy, manage, and operate applications on Kubernetes clusters (EKS or similar).
  • Perform Kubernetes troubleshooting including pod failures, scaling issues, networking, and cluster health.
  • Assist in managing production systems including monitoring, alerting, and incident response.
  • Troubleshoot and resolve infrastructure and application issues in a timely manner.
  • Work with engineering teams to support deployments and improve CI/CD workflows, including container-based releases.
  • Contribute to infrastructure automation using Terraform and Ansible.
  • Help maintain system reliability, availability, and performance through day-to-day operations.
  • Participate in on-call rotation and incident response process.
  • Assist with capacity planning and cost optimization efforts.
  • Partner with security teams to support compliance and security best practices.
  • Contribute to improving observability via logs, metrics, and dashboards.
  • Actively leverage AI/ML tools to improve cloud operations efficiency.

Skills

Cloud operations
DevOps
Infrastructure engineering
AWS
Linux/Unix
Containers

Tools

Docker
Kubernetes
Terraform
Ansible
CloudWatch
ELK
Prometheus
Grafana

Job description

Senior Cloud Operations Engineer, AI Ops (Cloud)

Hybrid-San Ramon, CA (2 days in office)

What We Do

ServiceMax is the global leader in Service Execution Management, delivering cloud-based software that helps companies maintain and service complex equipment at scale. Our platform powers mission-critical operations for customers around the world.

We foster a collaborative, #wintogether and #customerobsessed culture, where engineers learn, grow, and contribute to building reliable, secure, and intelligent cloud system.

What You Will
  • Support the operation and maintenance of cloud infrastructure across AWS environments
  • Deploy, manage, and operate applications on Kubernetes clusters (EKS or similar)
  • Perform Kubernetes troubleshooting including pod failures, scaling issues, networking, and cluster health
  • Assist in managing production systems including monitoring, alerting, and incident response
  • Troubleshoot and resolve infrastructure and application issues in a timely manner
  • Work with engineering teams to support deployments and improve CI/CD workflows, including container-based releases
  • Contribute to infrastructure automation using tools like Terraform and Ansible
  • Help maintain system reliability, availability, and performance through day-to-day operations
  • Participate in on-call rotation and incident response process
  • Assist with capacity planning and cost optimization efforts
  • Partner with security teams to support compliance and security best practices
  • Contribute to improving observability via logs, metrics, and dashboards
AI-Driven Operations
  • Actively leverage AI/ML tools and copilots to improve cloud operations efficiency and effectiveness
  • Use AI-assisted solutions for incident triage, root cause analysis, alert noise reduction, and automation of repetitive tasks
  • Continuously identify opportunities to apply AI to reduce operational toil and improve MTTR
  • Collaborate with teams to integrate AI-driven insights into monitoring, logging, and operational workflows
What You Bring to ServiceMax
  • 3–5 years of experience in cloud operations, DevOps, or infrastructure engineering
  • Ability to commute to the office in San Ramon, CA 2 days a week
  • Strong understanding of AWS cloud services
  • Experience working in Linux/Unix environments
  • Familiarity with container technologies (Docker and/or Kubernetes)
  • Exposure to infrastructure-as-code tools (Terraform, Ansible, or similar)
  • Understanding of CI/CD concepts and tools
  • Familiarity with monitoring and logging tools (e.g., CloudWatch, ELK, Prometheus, Grafana)
  • Strong problem-solving skills and willingness to learn
  • Ability to work collaboratively in a team environment
  • Good communication skills and attention to detail
Required AI
  • Hands-on experience using AI-powered tools such as GitHub Copilot, AWS Kiro, Claude Code, CodeX, etc.
  • Understanding of AI applications in cloud operations (anomaly detection, automation, incident analysis)
  • Ability to use AI tools for troubleshooting, log analysis, and workflow improvement
  • Demonstrated curiosity and willingness to adopt AI-driven approaches
Nice to Have Qualifi
  • Python or Shell scripting experience
  • Experience in SaaS or cloud-based production environments
  • Exposure to AIOps or observability platforms with built-in AI capabilities
What ServiceMax Offers You
  • Competitive health and wellness benefits (Medical, Dental, Vision, Life Insurance)
  • Flexible Spending Accounts
  • Flexible Time Off
  • 401(k) with employer match
  • Commuter Benefits
  • Opportunities for learning, mentorship, and career growth
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud AI Ops Engineer – Hybrid (San Ramon)
Senior Cloud AI Ops Engineer – Hybrid (San Ramon)

PTC • San Ramon (CA)

Hybrid
USD 140,000 - 190,000
Health benefits
401(k) with employer match
Commuter benefits
+1
Senior Cloud Operations Engineer - AI Cloud Ops
Senior Cloud Operations Engineer - AI Cloud Ops

PTC • San Ramon (CA)

On-site
USD 130,000 - 170,000
Competitive health benefits
401(k) with employer match
Flexible Time Off
Senior Cloud Operations Manager
Senior Cloud Operations Manager

PTC • San Ramon (CA)

Hybrid
USD 175,000 - 210,000
Wellness benefits
Commuter benefits
401(k) with employer match
+2
Senior Cloud Operations Engineer - AI Cloud Ops
Senior Cloud Operations Engineer - AI Cloud Ops

PTC • San Jose (CA)

Hybrid
USD 130,000 - 170,000
Health and wellness benefits
401(k) with employer match
Flexible Time Off
Senior IT Ops Automation Engineer (Hybrid, AI-Driven)
Senior IT Ops Automation Engineer (Hybrid, AI-Driven)

CloudZero • Boston (MA)

On-site
Manager-Cloud Operations
Manager-Cloud Operations

WellSpan Health • York

On-site
USD 90,000 - 120,000
Comprehensive health benefits
Retirement savings plan
Paid time off (PTO)
+3
Senior IT Operations Engineer
Senior IT Operations Engineer

CloudZero • Boston (MA)

Hybrid
USD 150,000 - 190,000
Equity options
Senior Cloud Security Engineer
Senior Cloud Security Engineer

BackOps AI • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Hybrid work model
Remote-friendly options
Equity opportunity
Cloud Operations Manager - EC2, VPC, IAM, S3, RDS, EKS
Cloud Operations Manager - EC2, VPC, IAM, S3, RDS, EKS

Resiliency LLC • San Francisco (CA)

On-site
USD 120,000 - 150,000
Staff Engineer, Cloud
Staff Engineer, Cloud

BrightAI Corporation • Palo Alto (CA)

On-site
USD 120,000 - 160,000