Technical Architect - ML

Quantiphi

United States

Remote

USD 180,000 - 260,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Quantiphi is seeking an experienced ML/AI engineering leader to architect and deliver enterprise-grade MLOps and LLMOps pipelines on AWS. You will shape model lifecycle, governance, and end-to-end deployment strategies across cloud-native platforms.

You will collaborate with data engineering, platform, and DevOps teams to build scalable, secure ML platforms (EKS-first) and ensure production-grade reliability with observability and compliance.

Qualifications

  • 8+ years working in ML/AI engineering or MLOps roles with strong architecture exposure.
  • Strong expertise in AWS cloud-native ML stack including SageMaker, EKS, Lambda, API Gateway, CI/CD.
  • Hands-on with major MLOps toolsets: MLflow, Kubeflow, SageMaker Pipelines, Airflow, BentoML, KServe, Seldon.
  • Deep understanding of model lifecycle management: feature engineering → training → registry → deployment → monitoring.
  • Experience implementing or supporting LLMOps pipelines, including prompt versioning and evaluation metrics.
  • Deep understanding of ML lifecycle: data ingestion, feature engineering, training, evaluation, packaging, CI/CD, drift monitoring.
  • Strong experience with AWS SageMaker (Pipelines, Feature Store, Model Registry, Model Monitor).
  • Experience with ML CI/CD pipelines including automated training, testing, validation, promotion, and endpoint deployment.
  • Experience with Infrastructure as Code (IaC) tools and CI/CD pipelines.
  • Experience with Kubernetes-based development and feature store management.
  • Understanding of lineage tracking: data snapshots, feature versions, code/versioning, reproducibility.
  • Hands-on with AWS Bedrock and Agentcore services, CloudWatch, SageMaker Monitor, Prometheus/Grafana.
  • Strong Python skills and cloud-native development patterns.
  • Solid understanding of security, IAM, secrets management, and artifact governance.

Responsibilities

  • Architect and implement the MLOps strategy aligned with proposal and delivery roadmap.
  • Design and own enterprise-grade ML/LLM pipelines for training, validation, deployment, versioning, monitoring, and CI/CD.
  • Build container-oriented ML platforms (EKS-first) while evaluating Kubeflow, SageMaker, MLflow, Airflow, etc.
  • Implement hybrid MLOps + LLMOps workflows, including prompt governance and evaluation frameworks.
  • Serve as a technical authority across internal and customer projects, contributing patterns and reusable frameworks.
  • Enable observability, monitoring, drift detection, lineage tracking, and auditability across ML/LLM systems.
  • Define standards for deployment, monitoring, governance, and automation to ensure production-grade reliability.
  • Collaborate with data engineering, platform, DevOps, and client stakeholders to deliver production-ready ML solutions.
  • Ensure security, governance, and compliance, especially around cloud services and Kubernetes workloads.
  • Conduct architecture reviews, troubleshoot complex ML system issues, and guide teams across platforms.
  • Mentor engineers and guide on modern MLOps tools and practices.

Skills

ML/AI engineering
MLOps
Python
Cloud-native development
Security best practices

Tools

SageMaker
EKS
Kubeflow
Airflow
Seldon
Terraform
CI/CD (CodeBuild/CodePipeline)
Kubernetes
Feature Store
Model Registry

Job description

While technology is the heart of our business, a global and diverse culture is the heart of our success. We love our people and we take pride in catering them to a culture built on transparency, diversity, integrity, learning and growth.

If working in an environment that encourages you to innovate and excel, not just in professional but personal life, interests you- you would enjoy your career with Quantiphi!

Must have skills & Qualifications:
  • 8+ years working in ML/AI engineering or MLOps roles with strong architecture exposure.
  • Strong expertise in AWS cloud-native ML stack, including: SageMaker(primary), EKS, Lambda, API Gateway, CI/CD (CodeBuild/CodePipeline or equivalent)
  • Hands-on experience with at least one major MLOps toolset and awareness of alternatives: MLflow, Kubeflow, SageMaker Pipelines, Airflow, BentoML, KServe, Seldon.
  • Deep understanding of model lifecycle management (feature engineering->training → registry → deployment → monitoring).
  • Experience implementing or supporting LLMOps pipelines, including: prompt versioning, evaluation metrics, automation frameworks
  • Deep understanding of ML lifecycle: data ingestion, feature engineering, training, evaluation, model packaging, CI/CD, drift detection, monitoring, and governance.
  • Strong experience with AWS SageMaker (Pipelines, Feature Store, Model Registry, Model Monitor).
  • Experience implementing ML CI/CD pipelines including automated training, testing, validation, model promotion, and endpoint deployment.
  • Experience working on Infrastructure as Code (IaC) tools and CI/CD pipelines
  • Experience with Kubernetes based development
  • Experience with feature engineering pipelines and Feature Store management.
  • Understanding of lineage tracking: training data snapshot, feature versions, code versioning, metadata tracking, reproducibility.
  • Hands-on experience with AWS Bedrock and Agentcore service
  • Experience with CloudWatch, SageMaker Model Monitor, Prometheus/Grafana.
  • Strong foundation in Python and cloud-native development patterns.
  • Solid understanding of security best practices, IAM, secrets management, and artifact governance.
Good to have skills:
  • Experience with vector databases, RAG pipelines, or multi-agent AI systems.
  • Exposure to DevOps and infrastructure-as-code (Terraform, Helm, CDK).
  • Hands-on understanding of model drift detection, A/B testing, canary rollouts, and blue-green deployments.
  • Familiarity with Observability stacks (Prometheus, Grafana, CloudWatch, OpenTelemetry).
  • SQL and data transformation experience using Snowflake, Databricks, Spark.
  • Ability to translate business goals into scalable AI/ML platform designs.
  • Strong communication and cross-team collaboration skills.
  • Ability to guide engineering teams through technical uncertainty and design choices.
Key Responsibilities:
  • Architect and implement the MLOps strategy for the programme, ensuring alignment with the project proposal and delivery roadmap.
  • Design and own enterprise-grade ML/LLM pipelines covering model training, validation, deployment, versioning, monitoring, and CI/CD automation.
  • Build container-oriented ML platforms (EKS-first) while evaluating alternative orchestration tools with similar capabilities (Kubeflow, SageMaker, MLflow, Airflow, etc.).
  • Implement hybrid MLOps + LLMOps workflows, including prompt/version governance, evaluation frameworks, and monitoring for LLM-based systems.
  • Serve as a technical authority across multiple internal and customer projects, contributing architectural patterns, best practices, and reusable frameworks.
  • Enable observability, monitoring, drift detection, lineage tracking, and auditability across ML/LLM systems.
  • Define and implement standards for model deployment, monitoring, governance, and automation to ensure production-grade reliability and scalability.
  • Collaborate with cross-functional teams — data engineering, platform, DevOps, and client stakeholders — to deliver production-ready ML solutions.
  • Ensure all solutions adhere to security, governance, and compliance expectations, particularly around handling cloud services, Kubernetes workloads, and MLOps tools.
  • Conduct architecture reviews, troubleshoot complex ML system issues, and guide teams through implementation across cloud-native ML platforms.
  • Mentor engineers and provide guidance on modern MLOps tools, platform capabilities, and best practices.

If you like wild growth and working with happy, enthusiastic over-achievers, you’ll enjoy your career with us_!_

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Associate Technical Architect - ML
Associate Technical Architect - ML

Quantiphi • Boston (MA)

On-site
CAD 120,000 - 170,000
Architect - Machine Learning
Architect - Machine Learning

Quantiphi • Boston (MA)

Hybrid
CAD 130,000 - 190,000
Remote Canada
Competitive salary
Health benefits
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
Senior ML Platform Architect - Cloud-Native MLOps
Senior ML Platform Architect - Cloud-Native MLOps

Quantiphi • United States

Remote
USD 180,000 - 260,000
Technical Architect - Machine Learning
Technical Architect - Machine Learning

Quantiphi • New Jersey

On-site
USD 170,000 - 250,000
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
Confidential
Consultant Machine Learning Engineer
Consultant Machine Learning Engineer

Careervitablr • United States

Remote
USD 140,000 - 210,000
Architect - Platform Engineering - USA
Architect - Platform Engineering - USA

Quantiphi, Inc. • Chicago (IL)

On-site
USD 130,000 - 180,000
Opportunity to work with Fortune 500 companies
Exposure to cutting-edge AI technologies
Dynamic team environment
Technical Architect - ML - GenAI
Technical Architect - ML - GenAI

Quantiphi, Inc. • United States

Remote
USD 150,000 - 230,000
Technical Architect - ML - GenAI
Technical Architect - ML - GenAI

Quantiphi • Northern (KY)

Hybrid
USD 170,000 - 210,000