Senior DevOps Engineer

PLP Group

New York, Northern (NY, KY)

Hybrid

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Domyn is seeking a Senior DevOps Engineer in the United States to lead deployment and scaling of AI solutions for financial workflows. You will drive infrastructure across cloud and on‑prem environments, ensuring secure, reliable, and high‑performing systems that power enterprise AI.

You will collaborate with engineering, product, and customer teams to streamline continuous delivery of software and AI models using Argo CD, GitHub Actions, and robust monitoring.

Qualifications

  • 6+ years of DevOps or SRE experience in distributed systems.
  • Expert-level Kubernetes (EKS, GKE, AKS) and cloud experience (AWS/GCP/Azure).
  • Proficiency with Terraform or Ansible for reproducible infra.
  • Experience building complex pipelines with Jenkins, Argo CD, or GitLab CI.
  • Scripting in Bash and Python; strong communication skills.

Responsibilities

  • Infrastructure leadership across cloud and on‑prem environments.
  • Build and support HPC clusters for model training and inference.
  • Design and implement robust CI/CD pipelines (Argo CD, GitHub Actions).
  • Define SLOs/SLIs and implement monitoring (Datadog, Prometheus, Grafana).
  • Enforce DevSecOps practices, IAM, network security, and compliance automation.
  • Oversee production databases (Postgres, Vector/Graph DB).

Skills

Kubernetes
Docker
Cloud platforms
Terraform/Ansible
CI/CD tooling
Monitoring/Observability
Networking/IAM
Scripting (Bash/Python)
Security/DevSecOps

Tools

Argo CD
GitHub Actions
Jenkins
GitLab CI
Datadog
Prometheus
Grafana
ELK

Job description

We are hiring a Senior DevOps Engineer to join our US team to lead the deployment and scaling of AI solutions for financial workflows. As our company delivers full-stack AI applications to some of the world’s leading financial institutions, your work will be critical in ensuring these solutions run flawlessly in the most demanding environments.

You will provide technical leadership and hands-on expertise to keep our infrastructure secure, reliable, and high performing. This role requires optimizing deployments for massive scale across both cloud and on-prem environments, implementing rigorous security controls, and managing the unique challenges of AI/ML workloads.

You’ll collaborate closely with our engineering, product, and customer teams to streamline continuous delivery of software and AI applications. As we scale deployments across Google Cloud Platform, Microsoft Azure, Amazon Web Services, and on-premises environments, you’ll play a central role in maintaining, optimizing, and supporting the systems that make it all possible.

This is an opportunity to work at the forefront of AI adoption in financial services, solving complex challenges and shaping the infrastructure behind innovative enterprise solutions.

Key Responsibilities

Infrastructure Leadership: Architect and maintain scalable, secure software-stack and infrastructure on major cloud providers (AWS, GCP, Azure) and on-premises environments.

AI Operations (AIOps): Build and support high-performance computing clusters for model training and inference.

Continuous Delivery: Design and implement robust CI/CD pipelines (Argo CD, GitHub Actions) to streamline the delivery of AI models and applications.

Reliability Engineering: Define SLOs/SLIs and implement comprehensive monitoring and alerting systems (Datadog, Prometheus, Grafana) to ensure high availability.

Security & Compliance: Enforce DevSecOps best practices, managing IAM policies, network security, and compliance automation for regulated financial environments.

Database Management: Oversee the deployment and maintenance of production databases, including Postgres, Vector Stores and Graph Databases.

What You Have

6+ years of DevOps or SRE experience, with a strong background in supporting distributed systems at scale.

Expert-level knowledge of Kubernetes (EKS, GKE, AKS) and Docker. Deep proficiency with at least one major cloud provider (AWS, GCP, or Azure) and hybrid/on-prem deployments.

Advanced skills in Terraform or Ansible for reproducible infrastructure.

Experience building complex pipelines with tools like Jenkins, Argo CD, or GitLab CI. Strong scripting skills in Bash and Python.

Hands-on experience with modern monitoring stacks (Datadog, ELK, Prometheus/Grafana) and distributed tracing.

Solid understanding of network security, IAM, VPC peering, and encryption standards.

Excellent communication and collaboration abilities.

What Would Be Nice to have

Experience deploying and scaling ML models (vLLM, Ray, Kubeflow) or managing GPU clusters.

Prior experience working in highly regulated industries (Investment Banks, Fund Managers, Custodian Banks, etc.).

Knowledge of managing stateful workloads on K8s or optimizing PostgreSQL/Vector DB/Graph DB performance.

Ability to interact with clients as needed.

Compensation

Domyn offers a competitive compensation structure, including salary, performance-based bonuses, and additional components based on experience. All roles include comprehensive benefits as part of the total compensation package.

About Domyn

Domyn is a company specializing in the research and development of Responsible AI for regulated industries, including financial services, government, and heavy industry. It supports enterprises with proprietary, fully governable solutions based on a composable AI architecture — including LLMs, AI agents, and one of the world’s largest supercomputers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Engineer
Senior AI Engineer

Domyn • New York (NY)

On-site
USD 140,000 - 230,000
AI Engineering and Research Intern
AI Engineering and Research Intern

PLP Group • New York (NY), Northern (KY)

Hybrid
USD 25,000 - 36,000
AI Infra Lead: Scalable FinTech DevOps
AI Infra Lead: Scalable FinTech DevOps

PLP Group • New York (NY), Northern (KY)

Hybrid
USD 140,000 - 210,000
AI / ML Engineer
AI / ML Engineer

Domify AI • New York (NY)

On-site
USD 150,000 - 210,000
Equity stake in startup
Health benefits
Competitive base salary
Senior AI Engineer: Scalable Agentic Finance Systems
Senior AI Engineer: Scalable Agentic Finance Systems

Domyn • New York (NY)

On-site
USD 140,000 - 230,000
Finance AI Research Engineer Intern
Finance AI Research Engineer Intern

Domyn • New York (NY)

On-site
DevOps
DevOps

Complexio • Warsaw (IN)

On-site
USD 100,000 - 130,000
Opportunity for professional growth
Collaborative team environment
Continuous learning in a dynamic field
Senior MLOps Engineer
Senior MLOps Engineer

Harnham • New York (NY)

On-site
USD 140,000 - 190,000
Base salary + bonus
Comprehensive benefits package
Senior DevOps Engineer
Senior DevOps Engineer

Gormat • Corridor North (MD)

On-site
USD 150,000 - 210,000
Senior DevOps Engineer
Senior DevOps Engineer

Gormat • Maryland

On-site
USD 120,000 - 180,000