Senior MLOps Platform Engineer {S}

Danbury Mission Technologies

Washington

On-site

USD 100,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive medical/vision/dental insurance packages
401k retirement plan
3 weeks vacation plus holidays
Annual bonus program
Upfront tuition assistance

Job summary

A leading technology company is looking for a MLOps Engineer to develop a unified MLOps platform for Kubernetes and AWS. You will design CI/CD pipelines and monitoring stacks for AI services. Ideal candidates have 5+ years in software infrastructure, expertise in Kubernetes, AWS, and proficiency in Python. The role is remote, supporting various U.S. locations, and offers flexible scheduling, competitive benefits including comprehensive insurance, 401k, and generous vacation policies.

Qualifications

  • 5+ years of experience building and operating production grade software infrastructure, preferably in a hybrid onprem / cloud environment.
  • Deep expertise with Kubernetes and container runtimes.
  • Hands on experience with AWS services and bridging onprem resources.

Responsibilities

  • Design and operate a unified MLOps platform for Kubernetes clusters and AWS.
  • Develop CI/CD pipelines for model packaging and testing.
  • Create monitoring stacks for real-time and batch workloads.

Skills

Kubernetes
AWS services
Python programming
CI/CD tooling
Distributed systems
Data engineering frameworks
Observability stacks

Education

BS in computer science or related engineering field

Tools

GitLab CI
Docker
Prometheus
Grafana

Job description

Description

ARKA Group L.P. (ARKA) is an advanced technologies company serving the U.S. military, intelligence community, and commercial space industry delivering next-generation solutions to support the national security space enterprise. Built on more than six decades of excellence, ARKA brings modern approaches and a culture of innovation to the challenges of today.

Position Overview

Our AI Center of Excellence builds the next generation of Agentic AI products that autonomously reason, plan, and act on behalf of our customers. To deliver these capabilities at scale, we need a platform engineering group that provides a robust, secure, and highly available MLOps foundation across both on premise clusters and AWS. The team works closely with data scientists, product engineers, and SREs to turn experimental models into reliable services that power mission critical applications.

In support of work/life balance, many positions are available for a flexible schedule within the pay period. Ask us about the opportunity for flex scheduling if that’s of interest to you.

Why join us
  • Shape the end-to-end lifecycle of cutting-edge AI services—from model training to production inference.
  • Influence architecture decisions for a hybrid cloud environment that will serve thousands of concurrent agents.
  • Collaborate with world-class researchers and product teams while enjoying a strong engineering culture focused on automation, observability, and reliability.
Responsibilities
  • Design, implement, and operate a unified MLOps platform that supports both on-premise Kubernetes clusters and AWS. The platform should enable rapid onboarding of new Agentic AI services and provide consistent governance across environments.
  • Develop reusable CI/CD pipelines (GitLab CI) for model packaging, containerization, automated testing, canary releases, and rollbacks.
  • Build observability, monitoring, and alerting stacks (Prometheus, Grafana, OpenTelemetry, CloudWatch) to track inference latency, throughput, resource utilization, and data drift for real time and batch workloads.
  • Create self-service tooling (CLI, SDKs, UI dashboards) that allows data science and product teams to register models, define inference endpoints, and manage versioning without deep DevOps involvement.
  • Architect and maintain data pipelines that feed training data, model artifacts, and inference logs into a governed data lake (S3, on prem object store).
  • Collaborate with research and product engineers to translate experimental Agentic AI prototypes into production grade services, ensuring reproducibility, security, and compliance.
  • Drive performance optimization for inference workloads (GPU/CPU scaling, model quantization, batching strategies).
  • Champion best practices in security (IAM, network policies, secret management), cost efficiency, and disaster recovery for the hybrid infrastructure.
  • Mentor junior engineers and contribute to internal knowledge bases, upskilling, and review processes.
Required Qualifications
  • BS in computer science or related engineering field
  • 5+ years of experience building and operating production grade software infrastructure, preferably in a hybrid onprem / cloud environment
  • Deep expertise with Kubernetes (cluster provisioning, Helm, operators, custom resources) and container runtimes (Docker, OCI)
  • Hands on experience with AWS services (EKS, SageMaker, S3, IAM, CloudWatch, Step Functions) and the ability to bridge onprem resources with AWS via VPN/Direct Connect
  • Strong software engineering skills in Python and at least one compiled language (Go, Rust, or Java) for building platform components and SDKs
  • Proficiency with CI/CD and GitOps tooling (Argo CD, Flux, Gitlab, GitHub Actions, or similar)
  • Solid understanding of distributed systems (consensus, fault tolerance, load balancing) and experience tuning high throughput, low latency inference pipelines
  • Experience with data engineering frameworks (Airflow, Prefect, Kafka, Spark, Flink) and building robust, versioned data pipelines
  • Familiarity with observability stacks (Prometheus, Grafana, OpenTelemetry, ELK) and the ability to define meaningful SLIs/SLOs for AI services
  • Track record of collaborating with research or product teams to move prototypes to production, translating experimental code into maintainable services
  • Strong problem solving mindset, excellent written and verbal communication, and a passion for building scalable AI platforms
Preferred Qualifications
  • Working knowledge of Scrum and Agile software development methodology
Location

Remote. This is a remote position that will primarily be supporting our Aurora, CO and King of Prussia, PA locations. Due to contract requirements, the job has to be performed from a remote location in the United States.

What We Offer
  • Comprehensive medical/vision/dental insurance packages
  • Company contributions to qualified HSA accounts
  • 401k retirement plan with industry leading company contributions
  • 3 weeks of vacation accrual per year plus time off for sick leave and unscheduled life events
  • 13 paid holidays
  • Upfront tuition assistance for approved degree programs
  • Annual bonus program based on company and employee performance
  • Company paid life insurance, AD&D, Short-Term and Long-Term disability insurance
  • 4 weeks paid Parental Leave
  • Employee assistance program (EAP)
EHS/Environmental Requirements

This job operates alongside a professional office environment. While performing the duties of this job, the employee routinely is required to use hands to keyboard, communicate, listen to, and interpret instructions and remain stationary for extended periods of the time. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions of the job.

ITC & Security Clearance Requirements

This position requires the incumbent to access export-controlled information. If you are not a U.S. Person, any offer is contingent upon the Company\'s ability to obtain a special license granting you access. This could take several months. You will not be able to begin employment until such license is obtained.

Visa Restrictions

No visa sponsorship is available for this position.

Pre-employment Screenings

Employment with any ARKA companies in the U.S. is contingent upon satisfactory completion of several pre-employment requirements to include a credit check, background check, and drug screen.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior MLOps Platform Engineer {S}
Senior MLOps Platform Engineer {S}

ARKA • Colorado Springs (CO)

Remote
USD 100,000 - 130,000
Company contributions to HSA accounts
401k retirement plan
3 weeks of vacation
+2
Staff AI/ML Engineer (TS/SCI) {S}
Staff AI/ML Engineer (TS/SCI) {S}

ARKA • Aurora (CO)

On-site
USD 155,000 - 200,000
Company contributions to qualified HSA accounts
401k retirement plan with industry leading contributions
3 weeks of vacation accrual per year
+2
Staff AI/ML Engineering Manager (TS/SCI) {S}
Staff AI/ML Engineering Manager (TS/SCI) {S}

TSG • Aurora (CO)

On-site
USD 160,000 - 200,000
401k retirement plan
3 weeks vacation per year
Annual bonus program
+1
Staff AI/ML Engineer (Large Language Model) (TS/SCI) {S}
Staff AI/ML Engineer (Large Language Model) (TS/SCI) {S}

TSG • Aurora (CO)

On-site
USD 150,000 - 200,000
401k retirement plan with industry leading company contributions
3 weeks of vacation accrual per year
Annual bonus program based on performance
AI/ML Engineer (TS/SCI with CI Poly) {S}
AI/ML Engineer (TS/SCI with CI Poly) {S}

TSG • King of Prussia (PA)

On-site
USD 100,000 - 140,000
Company contributions to qualified HSA accounts
401k retirement plan with leading contributions
3 weeks of vacation plus sick leave
+6
Staff AI/ML Engineer (Large Language Model) (TS/SCI) {S}
Staff AI/ML Engineer (Large Language Model) (TS/SCI) {S}

ARKA • King of Prussia (PA)

On-site
USD 120,000 - 160,000
401k retirement plan
Paid vacation and holidays
Tuition assistance
+2
Staff AI/ML Engineer (TS/SCI) {S}
Staff AI/ML Engineer (TS/SCI) {S}

TSG • Aurora (CO)

On-site
USD 155,000 - 200,000
401(k) retirement plan
3 weeks of vacation accrual
Annual bonus program
+2
Staff AI/ML Engineer (Large Language Model) (TS/SCI) {S}
Staff AI/ML Engineer (Large Language Model) (TS/SCI) {S}

TSG • King of Prussia (PA)

On-site
USD 100,000 - 150,000
401k retirement plan with industry-leading contributions
Annual bonus program
3 weeks of vacation accrual
+2
Staff AI/ML Engineer (TS/SCI) {S}
Staff AI/ML Engineer (TS/SCI) {S}

ARKA • King of Prussia (PA)

On-site
USD 130,000 - 170,000
401k retirement plan
3 weeks of vacation accrued annually
Company contributions to HSA accounts
+4
Principal AI/ML Engineer (Large Language Model) (TS/SCI) {S}
Principal AI/ML Engineer (Large Language Model) (TS/SCI) {S}

TSG • Aurora (CO)

On-site
USD 180,000 - 210,000
401k retirement plan with company contributions
3 weeks of vacation plus sick leave
Tuition assistance for approved programs
+3