Senior AI-Enabled Platform / SRE Engineer

ZipStaff Inc.

Dallas (TX)

Hybrid

USD 124,000 - 207,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ZipStaff Inc. seeks a Senior AI-Enabled Platform / SRE Engineer to advance platform reliability in a hybrid Dallas/Scottsdale setting. The role emphasizes automation, AI-driven operations, and cross-functional collaboration to support a major healthcare organization.

The ideal candidate has 5+ years in Kubernetes and SRE, with strong cloud and automation experience across GCP, Terraform, Helm, and CI/CD, plus hands-on observability and API/microservices reliability.

Qualifications

  • The candidate has 5+ years of Kubernetes experience with GKE and Rancher RKE2 across multi-cluster setups.
  • 5+ years in Python and Java; Node.js for integrations and automation is a plus.
  • Strong SRE background with focus on reliability, SLO/SLI and incident response.
  • Hands-on experience with GCP, Terraform, Helm, GitHub, CI/CD pipelines.
  • Observability experience: Splunk, Grafana, Datadog, AppDynamics.
  • Experience with API/microservices reliability (Apigee X, REST, GraphQL, traffic routing, canary, failover).
  • Experience applying LLMs (Gemini, Llama, Mistral, Qwen) to operations and alerting.
  • Hybrid work in Dallas, TX or Scottsdale, AZ; US work authorization required.

Responsibilities

  • Build automation and tooling using Java, Python, and Node.js to improve efficiency and scale.
  • Leverage Generative AI to automate alert analysis, incident triage, and runbooks.
  • Implement API and microservices reliability with Apigee X, REST, GraphQL, traffic routing, and failover.
  • Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster admin and tuning.
  • Support active-active deployments, DR readiness, and multi-datacenter environments.
  • Develop observability using Splunk, Grafana, Datadog, and AppDynamics.
  • Drive SRE best practices by partnering with cross-functional teams on reliability and security.

Skills

Kubernetes
GKE
Rancher RKE2
Python
Java
Node.js
GCP
Terraform
Helm
GitHub
CI/CD
Splunk
Grafana
Datadog
AppDynamics
Apigee
GraphQL
REST
LLMs

Tools

Apigee X
GitHub Actions

Job description

Job Title: Senior AI-Enabled Platform / SRE Engineer.
Location: Dallas, TX or Scottsdale, AZ (Hybrid).
Employment Type: Contract (W2 through ZipStaff).

About the Opportunity:

ZipStaff is seeking a Senior Kubernetes-focused SRE with strong cloud automation and software engineering skills to support a major healthcare organization.

This hybrid role (Dallas, TX or Scottsdale, AZ) focuses on platform reliability at scale and using AI/LLMs to automate operations. The initial assignment is approximately one year, with potential extension.

What You'll Do:
  • Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.
  • Leverage Generative AI (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.
  • Implement API and microservices reliability solutions using Apigee / Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.
  • Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
  • Support active-active deployments, disaster-recovery readiness, and multi-datacenter Kubernetes environments.
  • Develop observability and monitoring using Splunk, Grafana, Datadog, and AppDynamics.
  • Drive SRE best practices by partnering with cross-functional teams on reliability, security, incident management, and continuous improvement.
Required Qualifications:
  • 5+ years of strong hands-on Kubernetes platform experience with GKE and Rancher RKE2, including multi-cluster management, troubleshooting, and performance optimization.
  • 5+ years of advanced programming in Python and Java (Node.js preferred for integrations and automation).
  • Strong SRE background: reliability, availability, incident management, SLO/SLI monitoring, and operational excellence.
  • Strong experience in GCP, Terraform, Helm, GitHub, CI/CD, and production-grade automation.
  • Hands-on observability experience with Splunk, Grafana, Datadog, and/or AppDynamics.
  • Experience with API and microservices reliability (Apigee / Apigee X, REST, GraphQL, traffic routing, canary, failover).
  • Experience applying LLMs (Gemini, Llama, Mistral, Qwen, or similar) to alert analysis, incident triage, automation, or operational workflows.
  • Ability to work hybrid in Dallas, TX or Scottsdale, AZ.
  • Must be legally authorized to work in the United States without sponsorship now or in the future.
Preferred Qualifications:
  • Experience supporting highly available, multi-datacenter production platforms.
  • Prior healthcare or regulated-enterprise SRE experience.
  • Demonstrated AIOps / GenAI-for-operations implementations in production.
About ZipStaff:

ZipStaff partners with leading organizations to connect skilled technology professionals with high-impact contract opportunities. We focus on quality matches and long-term success.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI-Driven Platform SRE - Kubernetes & Cloud Automation
Senior AI-Driven Platform SRE - Kubernetes & Cloud Automation

ZipStaff Inc. • Dallas (TX)

Hybrid
USD 124,000 - 207,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

ZipStaff Inc. • Scottsdale (AZ)

Hybrid
USD 120,000 - 160,000
Kubernetes / GCP Engineer
Kubernetes / GCP Engineer

ZipStaff Inc. • Scottsdale (AZ)

Hybrid
USD 140,000 - 190,000
Sr. SRE – AI Platforms
Sr. SRE – AI Platforms

Mainz Brady Group • Portland (OR)

On-site
USD 140,000 - 180,000
Technical Program Manager SRE Kubernetes Cloud AI
Technical Program Manager SRE Kubernetes Cloud AI

IPolarity LLC • Whippany (NJ)

Hybrid
USD 130,000 - 170,000
Flexible work options
SRE/Platform Engineer
SRE/Platform Engineer

Stash Talent Services • Virginia (MN)

Remote
CloudOps / SRE Engineer
CloudOps / SRE Engineer

Autonomize AI • Austin (TX)

On-site
USD 140,000 - 190,000
Health, Vision, Dental insurance
401(k) retirement plan
Disability insurance
+1
SRE/Devops Engineer
SRE/Devops Engineer

INSPYR Solutions • Sunnyvale (CA)

Hybrid
USD 120,000 - 180,000
Work-life balance
No on-call requirements
Standard business hours
Senior Forward Deployed Engineer (DevOps/SRE)
Senior Forward Deployed Engineer (DevOps/SRE)

LeoForce • Pleasanton (CA)

On-site
USD 300,000 - 350,000
Medical benefits
401(k) plan
Equity
+1
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000