AI/ML Engineer-W2 only

Steneral Consulting

Scottsdale (AZ)

Hybrid

USD 140,000 - 190,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Steneral Consulting is seeking an experienced AI/ML Engineer to design, build and operate AI/ML infrastructure and agentic systems. You will design MCP servers/agents, integrate LLMs and implement RAG pipelines for production use.

The role emphasizes prompt engineering, toolchain telemetry, and scalable deployments on Kubernetes/Docker with Google Cloud. Collaboration across teams is required for reliable, cost-efficient AI workloads.

Qualifications

  • 5+ years of strong software engineering (Python/NodeJS) and production services.
  • 2+ years of experience with LLMs, prompts and agent frameworks.
  • 2+ years implementing RAG with embeddings and vector databases.
  • 2+ years using LangChain patterns and toolchain telemetry for tracing.
  • 5+ years Kubernetes, Docker, CI/CD and infrastructure-as-code.
  • 2+ years Google Cloud Platform services experience.
  • 2+ years observability, testing, security for distributed systems.
  • Knowledge of vector stores and embedding providers.

Responsibilities

  • Design, build and operate MCP servers and MCP agents that host, orchestrate and monitor AI/agent workloads.
  • Develop agentic AI, prompt engineering patterns, LLM integrations and developer tooling for production use.
  • Own deployment, scaling, reliability and cost-efficiency on Kubernetes/Docker and Google Cloud with automated CI/CD
  • Design and implement RAG (Retrieval Augmented Generation) pipelines and integrations with vector stores and retrieval tooling; use LangChain and Langfuse for orchestration, chaining, and observability.
  • Implement and maintain MCP server and agent code, APIs, and SDKs for model access and agent orchestration.
  • Design agent behavior, workflows and safety guards for agentic AI systems.
  • Create, test and iterate prompt templates, evaluation harnesses and grounding/chain of thought strategies.
  • Integrate LLMs and model providers (self hosted and cloud APIs) with unified adapters and telemetry.
  • Build developer tooling: CLI, local runner, simulators, and debugging tools for agents and prompts.
  • Containerize services (Docker), manage orchestration (Kubernetes/GKE), and optimize nodes, autoscaling and resource requests.
  • Ensure observability: logging, metrics, traces, dashboards, alerting and SLOs for model infra and agents.
  • Create runbooks, playbooks and incident response procedures; reduce MTTR and perform postmortems.
  • Design and maintain RAG workflows: document chunking, embeddings, vector indexing, retrieval strategies, re ranking and context injection.
  • Integrate and instrument LangChain for composable chains, agents and tooling; use Langfuse (or equivalent tracing) to capture prompts, model calls, RAG traces and evaluation telemetry.

Skills

Python
NodeJS
System design
CI/CD
Kubernetes
Docker
GCP
Observability
LangChain
Prompt engineering

Tools

LangChain
Langfuse
Vector stores
CI/CD tooling

Job description

AI/ML Engineer- ** 3 OPENINGS**

HYBRID **MULTIPLE LOCATIONS

Location: Scottsdale, AZ or Dallas, TX (HYBRID) - Locals only

ROPES AI Assessment - First Step

MUST INTEVRIEW ONSITE FOR 2nd ROUND INTERVIEW

Must have a valid LinkedIn with Picture and Location

Need Genuine Visa Copy and DL

Job Description
About the Role

We are seeking an experienced AIML Engineer to design, build, and operate AI/ML infrastructure and agentic systems. This role involves developing MCP servers and agents, integrating LLMs, and implementing RAG pipelines for production environments.

Key Responsibilities
  • Design, build and operate MCP servers and MCP agents that host, orchestrate and monitor AI/agent workloads.
  • Develop agentic AI, prompt engineering patterns, LLM integrations and developer tooling for production use.
  • Own deployment, scaling, reliability and cost-efficiency on Kubernetes/Docker and Google Cloud with automated CI/CD
  • Design and implement RAG (Retrieval Augmented Generation) pipelines and integrations with vector stores and retrieval tooling; use LangChain and Langfuse for orchestration, chaining, and observability.
Core Responsibilities
  • Implement and maintain MCP server and agent code, APIs, and SDKs for model access and agent orchestration.
  • Design agent behavior, workflows and safety guards for agentic AI systems.
  • Create, test and iterate prompt templates, evaluation harnesses and grounding/chain of thought strategies.
  • Integrate LLMs and model providers (self hosted and cloud APIs) with unified adapters and telemetry.
  • Build developer tooling: CLI, local runner, simulators, and debugging tools for agents and prompts.
  • Containerize services (Docker), manage orchestration (Kubernetes/GKE), and optimize nodes, autoscaling and resource requests.
  • Ensure observability: logging, metrics, traces, dashboards, alerting and SLOs for model infra and agents.
  • Create runbooks, playbooks and incident response procedures; reduce MTTR and perform postmortems.
  • Design and maintain RAG workflows: document chunking, embeddings, vector indexing, retrieval strategies, re ranking and context injection.
  • Integrate and instrument LangChain for composable chains, agents and tooling; use Langfuse (or equivalent tracing) to capture prompts, model calls, RAG traces and evaluation telemetry.
Required Skills & Experience
  • 5+ years of Strong Software Engineering (Python/NodeJS), system design and production service experience.
  • 2+ years of Experience with LLMs, prompt engineering, and agent frameworks.
  • 2+ years of Experience Practical experience implementing RAG: embeddings, vector DBs and retrieval tuning.
  • 2+ years of Experience with LangChain patterns and with toolchain telemetry (Langfuse or similar) for prompt/model traceability.
  • 5+ years of Experience with Kubernetes, Docker, CI/CD and infrastructure as code experience.
  • 2+ years of Experience with Practical experience with Google Cloud Platform services
  • 2+ years of Experience with Observability, testing, and security best practices for distributed systems.
  • 2+ years of Experience with evaluating and mitigating retrieval/augmentation failures, hallucinations, and leakage risks in RAG systems.
  • Familiarity with vendor and open source vector stores and embedding providers
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

C2C Job opportunity: AI/ML Engineer at Phoenix, AZ
C2C Job opportunity: AI/ML Engineer at Phoenix, AZ

Tech Mirrors • Phoenix (AZ)

Hybrid
USD 63,000 - 87,000
AI Engineer
AI Engineer

TechDigital Group • Scottsdale (AZ)

On-site
USD 150,000 - 210,000
AI Engineer
AI Engineer

Eightelevengroup • Miami (FL)

On-site
USD <65,000
AI / ML Engineer
AI / ML Engineer

Highbrow LLC • Dallas (TX)

On-site
USD 100,000 - 130,000
AI ML Engineer
AI ML Engineer

Innoventrics • Charlotte (NC)

Hybrid
USD 150,000 - 190,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Innovation-Technology-Services • Pennington (NJ)

Hybrid
USD 120,000 - 160,000
Lead Agentic AI engineer
Lead Agentic AI engineer

Envision Technology Solutions • Dallas (TX)

Hybrid
USD 140,000 - 200,000
AI Engineer
AI Engineer

Kaleidoscope Innovation • Fort Worth (TX)

On-site
USD 140,000 - 190,000
AI/ML Engineer
AI/ML Engineer

Synergis • Atlanta (GA)

Hybrid
Medical, dental, vision insurance
401k
Commuter benefits
Title: Principal AI Engineer – Agentic AI | Contract to Hire | Irvine, CA (Hybrid) | AS
Title: Principal AI Engineer – Agentic AI | Contract to Hire | Irvine, CA (Hybrid) | AS

Central Business Solutions, Inc • Irvine (CA)

Hybrid
USD 180,000 - 240,000