MLOps Technical Lead

TechDigital Group

Austin (TX)

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology company is seeking an AI Engineer to build intelligent, data-driven platforms for Generative AI systems in Austin, Texas. The ideal candidate will have over 3 years of backend experience and a proven track record of deploying AI features at scale. Key responsibilities include developing automated tools for evaluation and conducting rigorous data analyses. Strong proficiency in Python, along with hands-on experience with Docker and Kubernetes, is essential. Qualified candidates will have experience integrating machine learning models effectively.

Qualifications

  • 3+ years of backend or distributed systems experience.
  • Experience shipping AI/LLM features serving real users.
  • Proficient in Python for backend development.

Responsibilities

  • Build automated evaluation tools for AI systems.
  • Conduct statistical analyses to ensure reliability.
  • Benchmark and integrate AI/ML models into systems.

Skills

Backend or distributed systems experience
Experience shipping AI/LLM features
Python proficiency
Cloud infrastructure experience (AWS/GCP/Azure)
Hands-on with Docker and Kubernetes
Understanding of LLM integration

Tools

Docker
Kubernetes

Job description

Build intelligent, data-driven platform. The focus is to support the development of next-generation test analytics and test agents that enable faster insights, improved diagnostics, and scalable infrastructure for Generative AI systems connecting test stations, line level data, and pipelines. You will build automated evaluation tools, and conduct rigorous statistical analyses to ensure the reliability of both human and AI-based assessment systems.

Benchmark, adapt, and integrate AI/ML models into existing software systems. Independently run and analyze ML experiments for real improvements.

Must-Have Requirements
  • 3+ years of backend or distributed systems experience, with pre-AI production background
  • Experience shipping AI/LLM features serving real users at scale — not just prototypes or demos
  • Built AI agents, skills, tools, or MCP (Model Context Protocol) integrations
  • Python proficiency for backend development
  • Secondary language knowledge of Go, TypeScript, or Rust
  • Deep cloud infrastructure experience with AWS/GCP/Azure, including cost optimization and compute decisions
  • Hands‑on with Docker and Kubernetes — build, deploy, debug, and scale services
  • Understanding of LLM integration: token economics, context limits, rate limiting, structured outputs, API failure modes
  • Knowing how to evaluate LLM outputs and handle challenges such as non‑determinism, quality measurement, and regression detection
  • Practice as a hands‑on engineer who writes code, debugs production issues, and deploys their own work
Preferred / Differentiators
  • Built multi‑step agentic workflows with tool use and function calling
  • Experience with agent orchestration frameworks (LangGraph, CrewAI, or custom)
  • Built guardrails, fallbacks, or graceful degradation for AI systems
  • Streaming inference and async agent orchestration
  • Cost/latency optimization: caching, batching, prompt compression
  • ML observability tools: Langfuse, Arize, Braintrust, W&B
  • Retrieval systems (vector search, hybrid search) as a tool, not the focus
Screening Questions for Candidates
  1. Describe a production AI agent or skill system you built. What broke and how did you fix it?
  2. Have you built MCP servers/integrations or custom tool‑use systems for LLMs?
  3. How do you evaluate whether an LLM‑based feature is working well? What makes this hard?
  4. Walk me through how you’d deploy and scale an AI service on Kubernetes.
Not a Fit If
  • Primarily a model trainer/fine‑tuner (we're not training models)
  • AI experience is mainly academic, research, or tutorial‑based
  • No production systems experience (only notebooks/demos)
  • Looking for entry‑level role with heavy mentorship
  • Background is primarily data science/analytics rather than engineering
  • Architects who don't write or deploy code themselves
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied AI Engineer
Applied AI Engineer

Envision Technology Solutions • Austin (TX)

On-site
USD 150,000 - 190,000
Senior AI Machine Learning Engineer
Senior AI Machine Learning Engineer

Techsa • Egypt (PA)

Hybrid
USD 120,000 - 190,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
MLOps Technical Architect
MLOps Technical Architect

Veriipro • Atlanta (GA)

On-site
USD 150,000 - 190,000
AI/ML Engineer
AI/ML Engineer

Jobtailor • Colorado

On-site
USD 120,000 - 160,000
Applied AI Engineer
Applied AI Engineer

Tata Consultancy Services • Austin (TX)

On-site
USD 70,000 - 130,000
ML Ops Engineer — Agentic AI Lab (Founding Team)
ML Ops Engineer — Agentic AI Lab (Founding Team)

Fabrion • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Meaningful equity
Senior AI Engineer
Senior AI Engineer

7AI • Boston (MA)

On-site
AI Engineer, AIOps & Infrastructure
AI Engineer, AIOps & Infrastructure

eloquentai • San Francisco (CA)

On-site
USD 130,000 - 160,000
AI Engineer
AI Engineer

Teserac, Inc. • Sunnyvale (CA)

On-site
USD 100,000 - 130,000
Health Care Plan (Medical, Dental & Vision)
Paid Time Off (Vacation, Sick & Public Holidays)
Free Food & Snacks
+2