Senior DevOps Engineer

Domyn

New York (NY)

On-site

USD 150,000 - 190,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Performance bonuses
Comprehensive benefits

Job summary

Domyn is seeking a Senior DevOps Engineer to lead deployment and scaling of AI solutions for financial workflows in the United States. You will drive secure, reliable infrastructure across cloud providers and on‑prem environments to support massive AI workloads.

You will collaborate with engineering, product, and customer teams to accelerate continuous delivery using Argo CD, GitHub Actions, and other modern tools while ensuring SLOs/SLIs, monitoring, and IAM security are in place.

Qualifications

  • 6+ years of DevOps or SRE experience with distributed systems at scale.
  • Expert-level Kubernetes and Docker knowledge with a major cloud (AWS, GCP or Azure) and hybrid/on-prem deployments.
  • Advanced skills in Terraform or Ansible for reproducible infrastructure.
  • Experience building CI/CD pipelines with Jenkins, Argo CD, or GitLab CI; strong Bash and Python.
  • Experience with monitoring stacks (Datadog, ELK, Prometheus/Grafana) and distributed tracing.
  • Solid understanding of network security, IAM, VPC peering, and encryption standards.

Responsibilities

  • Infrastructure Leadership: Architect and maintain scalable, secure infra on cloud and on-prem.
  • AI Operations: Build HPC clusters for model training and inference.
  • Continuous Delivery: Design CI/CD pipelines (Argo CD, GitHub Actions) for AI models and apps.
  • Reliability Engineering: Define SLOs/SLIs; implement monitoring and alerting.
  • Security & Compliance: Enforce DevSecOps, IAM, network security, and compliance automation.
  • Database Management: Oversee production databases (Postgres, Vector Stores, Graph DB).

Skills

Distributed systems
Kubernetes
Scripting (Bash, Python)
Cloud & on-prem
Security & IAM
Collaboration

Tools

Kubernetes (EKS, GKE, AKS)
Docker
Terraform
Ansible
Jenkins
Argo CD
GitHub Actions
Datadog
Prometheus/Grafana

Job description

We are hiring a Senior DevOps Engineer to join our US team to lead the deployment and scaling of AI solutions for financial workflows. As our company delivers full-stack AI applications to some of the worlds leading financial institutions, your work will be critical in ensuring these solutions run flawlessly in the most demanding environments.

You will provide technical leadership and hands-on expertise to keep our infrastructure secure, reliable, and high performing. This role requires optimizing deployments for massive scale across both cloud and on-prem environments, implementing rigorous security controls, and managing the unique challenges of AI/ML workloads.

You will collaborate closely with our engineering, product, and customer teams to streamline continuous delivery of software and AI applications. As we scale deployments across Google Cloud Platform, Microsoft Azure, Amazon Web Services, and on-premises environments, you will play a central role in maintaining, optimizing, and supporting the systems that make it all possible.

This is an opportunity to work at the forefront of AI adoption in financial services, solving complex challenges and shaping the infrastructure behind innovative enterprise solutions.

Key Responsibilities
  • Infrastructure Leadership: Architect and maintain scalable, secure software-stack and infrastructure on major cloud providers (AWS, GCP, Azure) and on-premises environments.
  • AI Operations (AIOps): Build and support high-performance computing clusters for model training and inference.
  • Continuous Delivery: Design and implement robust CI/CD pipelines (Argo CD, GitHub Actions) to streamline the delivery of AI models and applications.
  • Reliability Engineering: Define SLOs/SLIs and implement comprehensive monitoring and alerting systems (Datadog, Prometheus, Grafana) to ensure high availability.
  • Security & Compliance: Enforce DevSecOps best practices, managing IAM policies, network security, and compliance automation for regulated financial environments.
  • Database Management: Oversee the deployment and maintenance of production databases, including Postgres, Vector Stores and Graph Databases.
What You Have
  • 6+ years of DevOps or SRE experience, with a strong background in supporting distributed systems at scale.
  • Expert-level knowledge of Kubernetes (EKS, GKE, AKS) and Docker. Deep proficiency with at least one major cloud provider (AWS, GCP, or Azure) and hybrid/on-prem deployments.
  • Advanced skills in Terraform or Ansible for reproducible infrastructure.
  • Experience building complex pipelines with tools like Jenkins, Argo CD, or GitLab CI. Strong scripting skills in Bash and Python.
  • Hands-on experience with modern monitoring stacks (Datadog, ELK, Prometheus/Grafana) and distributed tracing.
  • Solid understanding of network security, IAM, VPC peering, and encryption standards.
  • Excellent communication and collaboration abilities.
What Would Be Nice to have
  • Experience deploying and scaling ML models (vLLM, Ray, Kubeflow) or managing GPU clusters.
  • Prior experience working in highly regulated industries (Investment Banks, Fund Managers, Custodian Banks, etc.).
  • Knowledge of managing stateful workloads on K8s or optimizing PostgreSQL/Vector DB/Graph DB performance.
  • Ability to interact with clients as needed.
Benefits

Domyn offers a competitive compensation structure, including salary, performance-based bonuses, and additional components based on experience. All roles include comprehensive benefits as part of the total compensation package.

About Domyn

Domyn is a company specializing in the research and development of Responsible AI for regulated industries, including financial services, government, and heavy industry. It supports enterprises with proprietary, fully governable solutions based on a composable AI architecture - including LLMs, AI agents, and one of the worlds largest supercomputers.

Please review our Privacy Policy here https://bit.ly/4tndszN.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer
Senior DevOps Engineer

PLP Group • New York (NY), Northern (KY)

Hybrid
USD 140,000 - 210,000
Senior AI Engineer
Senior AI Engineer

Domyn • New York (NY)

On-site
USD 140,000 - 230,000
Senior AI-Scale DevOps Engineer (Cloud & On-Prem)
Senior AI-Scale DevOps Engineer (Cloud & On-Prem)

Domyn • New York (NY)

On-site
USD 150,000 - 190,000
Performance bonuses
Comprehensive benefits
AI Infra Lead: Scalable FinTech DevOps
AI Infra Lead: Scalable FinTech DevOps

PLP Group • New York (NY), Northern (KY)

Hybrid
USD 140,000 - 210,000
AI / ML Engineer
AI / ML Engineer

Domify AI • New York (NY)

On-site
USD 150,000 - 210,000
Equity stake in startup
Health benefits
Competitive base salary
Senior AI Engineer: Scalable Agentic Finance Systems
Senior AI Engineer: Scalable Agentic Finance Systems

Domyn • New York (NY)

On-site
USD 140,000 - 230,000
AI Engineering and Research Intern
AI Engineering and Research Intern

Domyn • New York (NY)

On-site
Finance AI Research Engineer Intern
Finance AI Research Engineer Intern

Domyn • New York (NY)

On-site
DevOps Engineer
DevOps Engineer

High 5 Games • New York (NY), Northern (KY)

Hybrid
USD 120,000 - 155,000
DevOps
DevOps

Complexio • Warsaw (IN)

On-site
USD 100,000 - 130,000
Opportunity for professional growth
Collaborative team environment
Continuous learning in a dynamic field