MLOps Architect

Kapitus

Virginia (MN)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Kapitus is seeking a Senior MLOps Architect to design and scale a modern ML and Generative AI platform across AWS. You will own the architecture for traditional ML and LLM pipelines, ensuring production reliability, governance, cost optimization, and enterprise-grade security in a large-scale environment.

You will lead end-to-end ML/GenAI workflows, integrate RAG pipelines and vector databases, define CI/CD standards, and enable secure deployment, monitoring, and retraining of models using

Qualifications

  • 6+ years of experience in ML engineering, data engineering, or MLOps roles.
  • Proven experience architecting ML platforms in AWS.
  • Strong hands-on experience with SageMaker (training, pipelines, deployment).

Responsibilities

  • Design and implement scalable ML and LLM infrastructure on AWS (SageMaker, EKS, S3, IAM, Lambda, Step Functions, CloudWatch).
  • Architect end-to-end ML and Generative AI lifecycle workflows (data ingestion, feature engineering, model training, evaluation, deployment).
  • Integrate LLM pipelines (RAG, prompts) into the enterprise MLOps stack.
  • Define CI/CD/CT pipelines across ML and GenAI workloads.
  • Establish guardrails for monitoring, governance, cost controls, and security.

Skills

AWS
SageMaker
Databricks
MLOps
GenAI
RAG architectures
Security
FinOps
CI/CD
Kubernetes

Tools

SageMaker
Databricks
EKS
IAM
CloudWatch
Terraform

Job description

Job Description

We are seeking a senior MLOps Architect to design and scale a modern ML and Generative AI platform across AWS. This role will own the architecture for traditional ML and LLM/Generative AI pipelines, ensuring production reliability, governance, cost optimization (FinOps), and enterprise-grade security. The ideal candidate has deep expertise in AWS, SageMaker, Databricks, Atlan (data catalog/governance), and modern MLOps tooling, and understands how to operationalize LLMs, RAG systems, and foundation models within a governed, scalable MLOps stack. This is a strategic, hands‑on architecture role responsible for integrating GenAI capabilities into an enterprise ML platform.

What you’ll Do:

MLOps & GenAI Platform Architecture

  • Design and implement scalable ML and LLM infrastructure on AWS (SageMaker, EKS, S3, IAM, Lambda, Step Functions, CloudWatch).
  • Architect end-to-end ML and Generative AI lifecycle workflows:
    • Data ingestion & preprocessing o Feature engineering / embedding generation o Model training & fine-tuning (traditional ML + foundation models)
    • Model evaluation & validation
    • Deployment (real-time, batch, streaming)
    • Monitoring & retraining
  • Integrate LLM pipelines (prompt workflows, RAG architectures, fine-tuning flows) into the enterprise MLOps stack.
  • Define standards for CI/CD/CT pipelines across ML and GenAI workloads.

Generative AI & LLM Operationalization

  • Architect Retrieval-Augmented Generation (RAG) pipelines including:
    • Embedding generation workflows
    • Vector database integration
    • Document ingestion and chunking strategies
    • Retrieval evaluation and monitoring
  • Design and deploy LLM-based services using:
    • Managed services (e.g., SageMaker endpoints, Bedrock-style APIs)
    • Containerized custom inference services
  • Establish prompt versioning, evaluation frameworks, and experiment tracking for LLM systems.
  • Implement guardrails for hallucination control, safety monitoring, bias detection, and usage logging.
  • Define architecture for LLM fine-tuning workflows (including data curation, evaluation, and cost controls).
  • Implement scalable orchestration of LLM pipelines using workflow engines and event-driven patterns.

Deployment, Monitoring & Reliability

  • Architect scalable inference patterns for:
    • Traditional ML models
    • LLM APIs
    • RAG systems
  • Implement model monitoring frameworks for:
    • Performance degradation
    • Drift detection
    • LLM output quality
    • Latency and token usage metrics
  • Define SLAs/SLOs for ML and GenAI systems.
  • Design safe deployment strategies (blue/green, canary, shadow testing).
  • Establish logging, observability, and traceability standards for GenAI systems

FinOps & Cost Optimization

  • Implement cost tracking for:
    • Training workloads o GPU utilization
    • Inference endpoints o Token consumption (LLM APIs)
    • Vector database storage
  • Optimize LLM workloads for cost-performance tradeoffs (model size, batching, caching strategies).
  • Design autoscaling and compute optimization strategies for GPU and CPU-based inference.
  • Partner with finance and engineering teams to forecast ML/GenAI infrastructure spend.

Platform Enablement & Standards

  • Define enterprise standards for:
    • Experiment tracking
    • Model registry
    • Prompt registry
    • Artifact management
    • Embedding versioning
  • Provide architectural guidance to data science, AI, and engineering teams.
  • Evaluate and recommend tooling across the ML/GenAI stack (MLflow, feature stores, vector databases, orchestration tools).
  • Drive documentation and reusable patterns for ML and GenAI development.

What We’re Looking for

  • 6+ years of experience in ML engineering, data engineering, or MLOps roles.
  • Proven experience architecting ML platforms in AWS.
  • Strong hands‑on experience with SageMaker (training, pipelines, deployment).
  • Experience operationalizing LLM or Generative AI systems in production.
  • Experience building RAG pipelines and integrating vector databases.
  • Experience working with Databricks in production.
  • Experience implem
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

MLOps Architect - Gen Al
MLOps Architect - Gen Al

Kapitus • Arlington (VA)

On-site
USD 117,000 - 189,000
Health insurance
401(k) with company match
Tuition reimbursement
+2
Staff MLOps Engineer – ML Platform
Staff MLOps Engineer – ML Platform

BrightAI Corporation • Palo Alto (CA)

On-site
USD 150,000 - 190,000
AI ARCHITECT
AI ARCHITECT

Signature IT World Inc • Massachusetts

On-site
USD 180,000 - 240,000
Senior AI Engineer - GenAI + Data Platform - AWS
Senior AI Engineer - GenAI + Data Platform - AWS

Compunnel, Inc. • Los Angeles (CA)

On-site
USD 120,000 - 160,000
MLOPs Architect
MLOPs Architect

Quantum World Technologies Inc. • Dallas (TX)

On-site
USD 120,000 - 180,000
Senior AI Engineer: Scalable LLMs & MLOps Leader
Senior AI Engineer: Scalable LLMs & MLOps Leader

Compunnel, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Architect – Generative AI, LLMs & AWS
AI Architect – Generative AI, LLMs & AWS

Apexon • West Palm Beach (FL)

On-site
USD 150,000 - 190,000
ML Ops Engineer — Agentic AI Lab (Founding Team)
ML Ops Engineer — Agentic AI Lab (Founding Team)

Fabrion • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Meaningful equity
Data Scientist Specialist
Data Scientist Specialist

Compunnel, Inc. • McLean (VA)

On-site
USD 190,000 - 260,000
null
AI Architect
AI Architect

M2S Group • United States

On-site
USD 170,000 - 250,000