AI/ML Engineer, Amazon Global Data Center Ops Central Insight and Analytics Team (AWS)

Amazon Inc.

Seattle (WA)

On-site

USD 144,000 - 194,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave
RSUs

Job summary

Amazon Data Services, Inc. is seeking an AI/ML Engineer to build, deploy, and operate ML/AI systems powering the agentic decision intelligence workflow. You will take models from notebooks to production, build the LLM integration layer, implement RAG pipelines, and establish evaluation frameworks for reliability.

This role is hands‑on with ML infrastructure, LLM applications, and production engineering, focusing on scalable, observable AI systems and end‑to‑end testing across the stack.

Qualifications

  • 3+ years of non‑internship professional software development experience
  • Bachelor's degree in Computer Science, Machine Learning, or related field (or equivalent experience)
  • 2+ years deploying ML models to production environments
  • Strong Python proficiency + experience with ML frameworks
  • Experience with LLM APIs and prompt engineering
  • Experience with cloud ML services
  • Experience building data pipelines for ML (feature engineering, preprocessing, training data management)
  • Solid software engineering fundamentals (testing, CI/CD, code review, production operations)

Responsibilities

  • Build and maintain LLM-powered components: structured reasoning chains, narrative generation, recommendation rationale
  • Implement and optimize prompt engineering pipelines with version control, A/B testing, and regression detection
  • Build RAG (Retrieval-Augmented Generation) systems that ground LLM outputs in operational data, historical playbooks, and domain knowledge
  • Build guardrails, validation layers, and output parsing for LLM responses. Optimize latency, cost, and quality trade-offs across LLM providers
  • Deploy ML models to production. Implement model monitoring: drift detection, performance degradation alerts, automated retraining triggers
  • Build A/B testing infrastructure for model experiments. Manage model versioning, rollback, and canary deployment. Ensure SLA compliance for inference latency and availability
  • Own the operational health of AI/ML services: monitoring, alarming, on‑call, incident response, observability across the AI stack (prompt traces, latency histograms, token usage, error rates)
  • Write comprehensive tests (unit, integration, end‑to‑end) for ML pipelines

Skills

Python
LLM APIs
Cloud ML services
Data pipelines
Production software

Education

Bachelor's degree in CS/ML or related

Tools

LangChain
LangGraph
CrewAI
Bedrock Agents
MLflow
SageMaker Pipelines
Step Functions
CDK
CloudFormation
Terraform
Feature stores

Job description

AI/ML Engineer, Amazon Global Data Center Ops Central Insight and Analytics Team

Job ID: 10513234 | Amazon Data Services, Inc.

We are looking for an AI/ML Engineer to build, deploy, and operate the ML/AI systems that power the agentic decision intelligence workflow we are building. You are the person who takes a model from a notebook to production, builds the LLM integration layer, implements RAG pipelines, creates evaluation frameworks, and ensures our AI systems are reliable, observable, and continuously improving.

This is a hands‑on engineering role with deep ML/AI focus — you write production code that runs AI systems, not research papers. If you love the intersection of ML infrastructure, LLM applications, and production engineering, this role is for you.

Key job responsibilities
  • Build and maintain LLM-powered components: structured reasoning chains, narrative generation, recommendation rationale
  • Implement and optimize prompt engineering pipelines with version control, A/B testing, and regression detection
  • Build RAG (Retrieval-Augmented Generation) systems that ground LLM outputs in operational data, historical playbooks, and domain knowledge
  • Build guardrails, validation layers, and output parsing for LLM responses. Optimize latency, cost, and quality trade-offs across LLM providers
  • Deploy ML models to production. Implement model monitoring: drift detection, performance degradation alerts, automated retraining triggers
  • Build A/B testing infrastructure for model experiments. Manage model versioning, rollback, and canary deployment. Ensure SLA compliance for inference latency and availability
  • Own the operational health of AI/ML services: monitoring, alarming, on‑call, incident response, observability across the AI stack (prompt traces, latency histograms, token usage, error rates)
  • Write comprehensive tests (unit, integration, end‑to‑end) for ML pipelines
Basic Qualifications
  • 3+ years of non‑internship professional software development experience
  • Bachelor's degree in Computer Science, Machine Learning, or related field (or equivalent experience)
  • 2+ years deploying ML models to production environments
  • Strong Python proficiency + experience with ML frameworks
  • Experience with LLM APIs and prompt engineering
  • Experience with cloud ML services
  • Experience building data pipelines for ML (feature engineering, preprocessing, training data management)
  • Solid software engineering fundamentals (testing, CI/CD, code review, production operations)
Preferred Qualifications
  • 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Experience building RAG systems (vector databases, embedding models, retrieval pipelines)
  • Experience with agent/orchestration frameworks (LangChain, LangGraph, CrewAI, Bedrock Agents, or custom)
  • Experience with ML evaluation frameworks (especially for generative AI / LLM outputs)
  • Experience with time-series ML (forecasting, anomaly detection)
  • Experience with MLOps tooling (MLflow, SageMaker Pipelines, Step Functions, feature stores)
  • Experience with infrastructure-as-code (CDK, CloudFormation, Terraform)
  • Background in operational/infrastructure environments

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign‑on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

USA, WA, Seattle - 143,700.00 - 194,400.00 USD annually

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI/ML Engineer, Amazon Global Data Center Ops Central Insight and Analytics Team
AI/ML Engineer, Amazon Global Data Center Ops Central Insight and Analytics Team

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+2
Software Development Engineer, AWS Marketing, Data Science & Engineering (D:SE)
Software Development Engineer, AWS Marketing, Data Science & Engineering (D:SE)

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
RSUs
401(k) matching
+2
Senior Delivery Consultant - AI/ML, AWS Professional Services
Senior Delivery Consultant - AI/ML, AWS Professional Services

Amazon • Chicago (IL)

On-site
USD 154,000 - 208,000
Software Development Engineer - Expert Consultant, AGI - Data Services
Software Development Engineer - Expert Consultant, AGI - Data Services

Amazon • Bellevue (WA)

On-site
USD 144,000 - 194,000
Software Development Engineer, AAIS
Software Development Engineer, AAIS

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
RSUs
401(k) matching
+1
Software Development Engineer, AAIS
Software Development Engineer, AAIS

Amazon • Seattle (WA)

On-site
USD 143,700 - 194,400
Health insurance
401(k) matching
Paid time off
+1
Delivery Consultant- AI/ML, Data & Machine Learning (DML)
Delivery Consultant- AI/ML, Data & Machine Learning (DML)

Amazon Inc. • Arlington (VA), Northern (KY)

Hybrid
USD 131,000 - 178,000
Senior Delivery Consultant - AI/ML, AWS Professional Services
Senior Delivery Consultant - AI/ML, AWS Professional Services

Amazon • Mountain View (CA)

On-site
USD 177,000 - 239,000
Health insurance
401(k) matching
Paid time off
Machine Learning Engineer, Advertising & Marketing Performance Intelligence
Machine Learning Engineer, Advertising & Marketing Performance Intelligence

Amazon • Seattle (WA), Northern (KY)

Hybrid
USD 144,000 - 194,000
Health insurance
RSUs
401(k) matching
+1
Sr. Machine Learning Engineer, AWS Applied AI Solution
Sr. Machine Learning Engineer, AWS Applied AI Solution

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000