MLOps Architect - Gen Al

Kapitus

Arlington (VA)

On-site

USD 117,800 - 189,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) with company match
Tuition reimbursement
Paid maternity and parental leave
Flexible Spending Account

Job summary

Kapitus is seeking a Senior MLOps Architect in Arlington, VA, to design a modern ML and Generative AI platform on AWS. The ideal candidate will lead the architecture of ML pipelines, ensuring reliability and security while optimizing costs.

This role requires extensive experience in ML engineering and AWS, with a strong focus on implementing scalable solutions for Generative AI. Competitive salary and benefits include health insurance, retirement plans, and tuition reimbursement.

Qualifications

  • 6+ years of experience in ML engineering, data engineering, or MLOps.
  • Proven experience architecting ML platforms in AWS.
  • Strong hands-on experience with SageMaker.

Responsibilities

  • Design and implement scalable ML infrastructure on AWS.
  • Architect end-to-end ML and Generative AI lifecycle workflows.
  • Define standards for CI/CD/CT pipelines.

Skills

ML engineering
MLOps
Data engineering
AWS
SageMaker
Generative AI
CI/CD
Docker
Kubernetes
Infrastructure as code

Tools

Databricks
Terraform
CloudFormation

Job description

Overview

We are seeking a Senior MLOps Architect to design and scale a modern ML and Generative AI platform across AWS. This role will own the architecture for traditional ML and LLM/Generative AI pipelines, ensuring production reliability, governance, cost optimization, and enterprise‑grade security.

Responsibilities
  • Design and implement scalable ML and LLM infrastructure on AWS (SageMaker, EKS, S3, IAM, Lambda, Step Functions, CloudWatch).
  • Architect end‑to‑end ML and Generative AI lifecycle workflows:
    • Data ingestion & preprocessing, feature engineering / embedding generation, model training & fine‑tuning (traditional ML + foundation models).
    • Model evaluation & validation.
    • Deployment (real‑time, batch, streaming).
    • Monitoring & retraining.
  • Integrate LLM pipelines (prompt workflows, RAG architectures, fine‑tuning flows) into the enterprise MLOps stack.
  • Define standards for CI/CD/CT pipelines across ML and GenAI workloads.
  • Architect Retrieval‑Augmented Generation (RAG) pipelines:
    • Embedding generation workflows.
    • Vector database integration.
    • Document ingestion and chunking strategies.
    • Retrieval evaluation and monitoring.
  • Design and deploy LLM‑based services using:
    • Managed services (e.g., SageMaker endpoints, Bedrock‑style APIs).
    • Containerized custom inference services.
  • Establish prompt versioning, evaluation frameworks, and experiment tracking for LLM systems.
  • Implement guardrails for hallucination control, safety monitoring, bias detection, and usage logging.
  • Define architecture for LLM fine‑tuning workflows (including data curation, evaluation, and cost controls).
  • Implement scalable orchestration of LLM pipelines using workflow engines and event‑driven patterns.
  • Architect scalable inference patterns for traditional ML models, LLM APIs, and RAG systems.
  • Define latency and token usage metrics, SLAs/SLOs, and safe deployment strategies (blue/green, canary, shadow testing).
  • Establish logging, observability, and traceability standards for GenAI systems.
  • Implement cost tracking for training workloads (GPU utilization), inference endpoints (token consumption), and vector database storage.
  • Optimize LLM workloads for cost‑performance tradeoffs (model size, batching, caching strategies). Design autoscaling and compute optimization strategies for GPU and CPU‑based inference.
  • Partner with finance and engineering teams to forecast ML/GenAI infrastructure spend.
  • Provide experiment tracking; give architectural guidance to data science, AI, and engineering teams; evaluate and recommend tooling across the ML/GenAI stack (MLflow, feature stores, vector databases, orchestration tools). Drive documentation and reusable patterns.
Qualifications
  • 6+ years of experience in ML engineering, data engineering, or MLOps roles.
  • Proven experience architecting ML platforms in AWS.
  • Strong hands‑on experience with SageMaker (training, pipelines, deployment).
  • Experience operationalizing LLM or Generative AI systems in production.
  • Experience building RAG pipelines and integrating vector databases.
  • Experience working with Databricks in production.
  • Experience implementing data governance and catalog systems (e.g., Atlan).
  • Strong understanding of CI/CD principles for ML and GenAI.
  • Experience with containerization (Docker) and orchestration (Kubernetes/EKS).
  • Deep knowledge of infrastructure‑as‑code (Terraform, CloudFormation).
  • Strong understanding of observability and monitoring for ML systems.
  • Experience implementing cloud cost optimization strategies (FinOps).
  • Experience with foundation model fine‑tuning and parameter‑efficient methods.
  • Experience implementing model registries and experiment tracking tools.
  • Experience designing feature stores and embedding stores.
  • Familiarity with AI risk management, bias mitigation, and safety controls.
  • Experience supporting regulated or data‑sensitive environments.
  • Platform‑level architectural thinking.
  • Deep understanding of how to integrate GenAI into enterprise ML ecosystems.
  • Ability to balance scalability, governance, security, performance, and cost.
  • Strong technical leadership and cross‑functional collaboration skills.
  • Hands‑on ability to move from architecture design to implementation.
Benefits
  • Competitive Base Salary Range: $117,800 – $189,000.
  • Annual Incentive Compensation: up to 10% of base.
  • Health, dental, vision insurance (UnitedHealthcare).
  • Flexible Spending Account, Lifestyle Spending Account.
  • Fully paid disability insurance.
  • Paid maternity and parental leave beyond state‑mandated policies.
  • Commuter benefits.
  • LifeBalance membership discounts.
  • Plum Benefits discount program.
  • Tuition reimbursement up to $5,000 annually.
  • Travel reimbursement for work‑related travel.
  • Paid time off and sick time.
  • 401(k) plan with 25% match up to 6% of salary.
EEO Statement

As set forth in Kapitus’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AWS Gen AI / ML Engineer - Plano, TX
AWS Gen AI / ML Engineer - Plano, TX

Photon • United States

Hybrid
USD 48,000 - 168,000
Medical, vision, and dental benefits
401k retirement plan
Paid time off
+1
Senior AI Engineer
Senior AI Engineer

Hophr • Los Angeles (CA)

On-site
USD 120,000 - 160,000
Staff MLOps Engineer – ML Platform
Staff MLOps Engineer – ML Platform

BrightAI Corporation • Palo Alto (CA)

On-site
USD 150,000 - 190,000
Architect ML/GenAI IRC296972
Architect ML/GenAI IRC296972

GlobalLogic • Town of Poland (NY)

On-site
USD 120,000 - 150,000
Comprehensive benefits package
Career development opportunities
Flexible work arrangements
AI/ML Engineer 2
AI/ML Engineer 2

Day & Zimmermann Company • Philadelphia

On-site
USD 101,000 - 166,000
Medical/Rx coverage
Dental and vision coverage
100% paid maternity leave
+2
Senior AI Engineer - GenAI + Data Platform - AWS
Senior AI Engineer - GenAI + Data Platform - AWS

Compunnel, Inc. • Los Angeles (CA)

On-site
USD 120,000 - 160,000
AI/ML Engineer
AI/ML Engineer

CCS INC • Plano (TX)

On-site
USD 120,000 - 160,000
Bonus based on performance
Dental insurance
Health insurance
+1
Senior ML OPs Engineer
Senior ML OPs Engineer

Glocomms • California (MO)

Hybrid
USD 198,000 - 230,000
Meal stipends for remote work days
Generous paid time off
Comprehensive health coverage
+1
MLOps Engineer
MLOps Engineer

ERT, Inc. • Arlington (VA)

On-site
USD 120,000 - 150,000
Architect, GenAI
Architect, GenAI

Lovelytics • Chicago (IL), Arlington (VA)

Hybrid
USD 170,000 - 210,000