Platform Reliability Engineer - AI, Kubernetes & GitOps

Nasdaq

Philadelphia (Philadelphia County)

Hybrid

USD 140,000 - 200,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

401(k) with company match
Employee stock purchase program (ESPP)
Health, dental, vision coverage
Generous PTO and paid time off
Hybrid work schedule

Job summary

Nasdaq is seeking a Sr. Specialist - Platform Operations to ensure reliable, scalable operation of Nasdaq's platforms, including AI systems.

You will manage Kubernetes clusters, design automated delivery pipelines with ArgoCD, and drive IaC with Terraform/OpenTofu/Pulumi while improving observability and incident response. You will collaborate across teams for developer enablement, lead deployment and modernization projects, and apply AI-focused automation to reduce toil in a fast-paced

Qualifications

  • 2+ years of hands-on experience managing Kubernetes workloads, ingress controllers, storage, and networking.
  • Proven experience implementing GitOps workflows, specifically configuring and troubleshooting ArgoCD at scale.
  • Strong familiarity with continuous integration tools like GitHub Actions, GitLab CI, or Jenkins.
  • Proficiency in programming languages like Go or Python, alongside strong Bash scripting skills.
  • Deep understanding of Linux internals, container runtimes (Docker, containerd), and core networking concepts (DNS, TCP/IP).
  • Hands-on experience with cloud infrastructure and services, with strong expertise in AWS architecture.
  • Familiarity with AI concepts and tooling and a strong interest in operationalizing AI systems.

Responsibilities

  • Platform Operations: Manage, scale, and optimize production-grade Kubernetes clusters across multi-cloud or hybrid environments.
  • GitOps & Deployment: Design and maintain automated application delivery pipelines using ArgoCD to ensure declarative environment states.
  • Infrastructure as Code (IaC): Provision and manage cloud infrastructure using tools like Terraform, OpenTofu, or Pulumi.
  • Observability: Implement and maintain robust monitoring, logging, and alerting systems using Prometheus, Grafana, and ELK/OpenSearch stacks.
  • Reliability Engineering: Participate in on-call rotations, conduct blameless post-mortems, and minimize operational toil through automation.
  • Developer Experience: Collaborate with software engineering teams to streamline onboarding and reduce friction in the software development lifecycle.
  • AI and agentic workflows: Design, plan, and deploy changes to existing systems, driving operational excellence and introducing new solutions, including.
  • Lead and contribute to implementation projects, from requirements through to launch and ongoing improvement, with a focus on deploying, scaling, and monitoring AI and agentic services.
  • Identify and implement automation and process improvements, leveraging AI and agentic tooling to enhance how systems are tested, deployed, and maintained.
  • Support the operational reliability, observability, and responsible/governed use of AI models and agents, including performance, cost, and safety monitoring.

Skills

Kubernetes
GitOps
CI/CD
Automation & Scripting
Linux fundamentals
AWS
AI concepts

Tools

Go
Python

Job description

Nasdaq is seeking a Sr. Specialist - Platform Operations to ensure reliable, scalable operation of Nasdaq's platforms, including AI systems.

You will manage Kubernetes clusters, design automated delivery pipelines with ArgoCD, and drive IaC with Terraform/OpenTofu/Pulumi while improving observability and incident response. You will collaborate across teams for developer enablement, lead deployment and modernization projects, and apply AI-focused automation to reduce toil in a fast-paced

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Operations Engineer for AI-Driven Automation
Platform Operations Engineer for AI-Driven Automation

Nasdaq, Inc. • Philadelphia

Hybrid
USD 120,000 - 150,000
401(k) with employer match
Employee stock purchase program
Student loan repayment
+2
Platform Reliability Engineer - Cloud Infra & CI/CD
Platform Reliability Engineer - Cloud Infra & CI/CD

crewAI, Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
Platform Engineer — Scalable AI Deployment & CI/CD
Platform Engineer — Scalable AI Deployment & CI/CD

Baseten • San Francisco (CA)

On-site
USD 165,000 - 330,000
Equity
Health insurance
Flexible PTO
+4
Senior Staff Platform Engineer — AI-Driven Reliability
Senior Staff Platform Engineer — AI-Driven Reliability

ServiceNow • California (MO)

On-site
USD 180,000 - 270,000
Health plans
401(k) Plan with company match
ESPP
+3
Senior Platform Reliability Engineer (Kubernetes & CI/CD)
Senior Platform Reliability Engineer (Kubernetes & CI/CD)

Optomi • Orlando (FL)

Hybrid
USD 120,000 - 180,000
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Platform Engineer – AI Infra, CI/CD & Observability
Platform Engineer – AI Infra, CI/CD & Observability

Outmarket AI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Site Reliability Engineer: AI-Driven Kubernetes Automation
Site Reliability Engineer: AI-Driven Kubernetes Automation

Kindredventures • United States

Remote
USD 8,000 - 15,000
Senior Platform Engineer — Hybrid Infra, Cloud & Automation
Senior Platform Engineer — Hybrid Infra, Cloud & Automation

Socket.dev • Pennsylvania

Hybrid
USD 155,000 - 287,000
401(k) match
Employee Stock Purchase Plan
Student loan repayment
+4
Lead AI-Driven DevOps Engineer
Lead AI-Driven DevOps Engineer

Nasdaq • New York (NY)

Hybrid
USD 139,000 - 243,000
Hybrid work model