AI/ML Infrastructure Engineer: RAG & Agentic (Hybrid)

Steneral Consulting

Scottsdale (AZ)

Hybrid

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Steneral Consulting is seeking an experienced AI/ML Engineer to design, build and operate AI/ML infrastructure and agentic systems. You will design MCP servers/agents, integrate LLMs and implement RAG pipelines for production use.

The role emphasizes prompt engineering, toolchain telemetry, and scalable deployments on Kubernetes/Docker with Google Cloud. Collaboration across teams is required for reliable, cost-efficient AI workloads.

Qualifications

  • 5+ years of strong software engineering (Python/NodeJS) and production services.
  • 2+ years of experience with LLMs, prompts and agent frameworks.
  • 2+ years implementing RAG with embeddings and vector databases.
  • 2+ years using LangChain patterns and toolchain telemetry for tracing.
  • 5+ years Kubernetes, Docker, CI/CD and infrastructure-as-code.
  • 2+ years Google Cloud Platform services experience.
  • 2+ years observability, testing, security for distributed systems.
  • Knowledge of vector stores and embedding providers.

Responsibilities

  • Design, build and operate MCP servers and MCP agents that host, orchestrate and monitor AI/agent workloads.
  • Develop agentic AI, prompt engineering patterns, LLM integrations and developer tooling for production use.
  • Own deployment, scaling, reliability and cost-efficiency on Kubernetes/Docker and Google Cloud with automated CI/CD
  • Design and implement RAG (Retrieval Augmented Generation) pipelines and integrations with vector stores and retrieval tooling; use LangChain and Langfuse for orchestration, chaining, and observability.
  • Implement and maintain MCP server and agent code, APIs, and SDKs for model access and agent orchestration.
  • Design agent behavior, workflows and safety guards for agentic AI systems.
  • Create, test and iterate prompt templates, evaluation harnesses and grounding/chain of thought strategies.
  • Integrate LLMs and model providers (self hosted and cloud APIs) with unified adapters and telemetry.
  • Build developer tooling: CLI, local runner, simulators, and debugging tools for agents and prompts.
  • Containerize services (Docker), manage orchestration (Kubernetes/GKE), and optimize nodes, autoscaling and resource requests.
  • Ensure observability: logging, metrics, traces, dashboards, alerting and SLOs for model infra and agents.
  • Create runbooks, playbooks and incident response procedures; reduce MTTR and perform postmortems.
  • Design and maintain RAG workflows: document chunking, embeddings, vector indexing, retrieval strategies, re ranking and context injection.
  • Integrate and instrument LangChain for composable chains, agents and tooling; use Langfuse (or equivalent tracing) to capture prompts, model calls, RAG traces and evaluation telemetry.

Skills

Python
NodeJS
System design
CI/CD
Kubernetes
Docker
GCP
Observability
LangChain
Prompt engineering

Tools

LangChain
Langfuse
Vector stores
CI/CD tooling

Job description

Steneral Consulting is seeking an experienced AI/ML Engineer to design, build and operate AI/ML infrastructure and agentic systems. You will design MCP servers/agents, integrate LLMs and implement RAG pipelines for production use.

The role emphasizes prompt engineering, toolchain telemetry, and scalable deployments on Kubernetes/Docker with Google Cloud. Collaboration across teams is required for reliable, cost-efficient AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI/ML Infrastructure Engineer
Senior AI/ML Infrastructure Engineer

TechDigital Group • Scottsdale (AZ)

On-site
USD 150,000 - 210,000
GenAI & RAG Engineer – Agentic AI Specialist (Hybrid)
GenAI & RAG Engineer – Agentic AI Specialist (Hybrid)

Capgemini • Charlotte (NC)

Hybrid
USD 150,000 - 210,000
Healthcare benefits
Paid time off
Parental leave
+1
AI/ML Engineer – RAG, LLMs & Kubernetes (Hybrid)
AI/ML Engineer – RAG, LLMs & Kubernetes (Hybrid)

Tech Mirrors • Phoenix (AZ)

Hybrid
USD 63,000 - 87,000
Senior AI/ML Full-Stack Engineer: RAG & LLM Orchestration
Senior AI/ML Full-Stack Engineer: RAG & LLM Orchestration

MAS Global Consulting • Plano (TX)

On-site
USD 140,000 - 190,000
Senior AI Engineer — Agentic AI & RAG Pipelines (Hybrid)
Senior AI Engineer — Agentic AI & RAG Pipelines (Hybrid)

Dormont Manufacturing Co • Atlanta (GA)

Hybrid
USD 144,000 - 230,000
Remote AI/ML Engineer — LLMs, RAG & Multi-Agent
Remote AI/ML Engineer — LLMs, RAG & Multi-Agent

YO AI Labs • Phoenix (AZ)

Remote
USD 120,000 - 180,000
Senior AI Engineer: Agentic & RAG Systems (Remote)
Senior AI Engineer: Agentic & RAG Systems (Remote)

EPAM Systems • Georgia

On-site
USD 120,000 - 180,000
Remote work within Georgia
Relocation opportunities
Certifications (GCP/AWS)
+3
Senior AI/ML Engineer — Agentic Systems & RAG Platforms
Senior AI/ML Engineer — Agentic Systems & RAG Platforms

Vanguard • Philadelphia

Hybrid
USD 150,000 - 210,000
Senior AI/ML Engineer — Production-Grade Agentic AI
Senior AI/ML Engineer — Production-Grade Agentic AI

Involved Solutions • Town of Texas (WI)

On-site
Remote AI/ML Engineer: LLMs, RAG & Secure Cloud
Remote AI/ML Engineer: LLMs, RAG & Secure Cloud

YO AI Labs • Illinois

Remote
USD 120,000 - 180,000