LLM / GenAI Engineer

Evlo AI

Miami (FL)

On-site

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Evlo AI in Miami is seeking a software engineer to build production-grade AI systems. You will design robust RAG pipelines, agentic workflows, and fine-tuning pipelines, collaborating with applied scientists and backend engineers.

You will implement scalable Python-based solutions, optimize embeddings, and contribute to evaluation frameworks for LLMs in production environments.

Qualifications

  • 3–6 years of software engineering experience, with at least 2 years building LLM apps in production.
  • Deep familiarity with LLM orchestration frameworks and prompt engineering best practices.
  • Strong proficiency in Python and REST API design, plus cloud infra integration.
  • Solid understanding of embeddings, vector spaces, tokenization, and model quantization.
  • Bonus: experience with CUDA kernels, vLLM or TGI deployments, and open-source contributions.

Responsibilities

  • Design and implement production RAG pipelines.
  • Build and optimize vector DB integrations for low-latency search.
  • Develop LLM evaluation frameworks including benchmarks and regression tests.
  • Execute parameter-efficient fine-tuning pipelines (LoRA/QLoRA).
  • Write observable, tested Python code and participate in reviews.

Skills

Python
async programming
REST API design
cloud infrastructure
LLM orchestration
prompt engineering
context window optimization
embedding models
vector spaces
tokenization
model quantization
CUDA kernels
vLLM
TGI
open-source contributions

Job description

About The Role

The role is for someone who has moved beyond basic prompting and understands what it takes to build production-grade AI systems: robust RAG pipelines, agentic workflows, fine-tuning pipelines, and systematic evaluation frameworks.

About The Role

The role is for someone who has moved beyond basic prompting and understands what it takes to build production-grade AI systems: robust RAG pipelines, agentic workflows, fine-tuning pipelines, and systematic evaluation frameworks.

The engineering team owns complex pieces of a high-scale AI platform, working directly with applied scientists and backend engineers to deploy performant LLM applications.

Key Responsibilities
  • Design and implement production RAG pipelines using LangChain, LlamaIndex, or custom Python architectures
  • Build and optimize vector database integrations such as Pinecone, Weaviate, or pgvector for low-latency semantic search
  • Develop systematic LLM evaluation frameworks including benchmark suites, LLM-as-judge pipelines, and automated regression testing
  • Execute parameter-efficient fine-tuning pipelines using LoRA and QLoRA on specialized domain datasets
  • Write observable, tested, and well-documented Python code; participate actively in architecture reviews and deployment pipelines
What We Are Looking For
  • 3 to 6 years of software engineering experience, with at least 2 years specifically focused on building LLM applications in production
  • Deep familiarity with LLM orchestration frameworks, prompt engineering best practices, and context window optimization
  • Strong proficiency in Python, async programming, REST API design, and cloud infrastructure integration
  • Solid understanding of embedding models, vector spaces, tokenization, and model quantization techniques
  • Bonus: Experience with custom CUDA kernels, open-source model deployment via vLLM or TGI, and published contributions to AI open-source projects
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Ai/Llm Engineer
Ai/Llm Engineer

2T Consulting • Maywood (NJ)

On-site
USD 130,000 - 190,000
Ai/Llm Engineer
Ai/Llm Engineer

2T Consulting • Garfield (NJ)

On-site
USD 120,000 - 180,000
Senior AI EngineerWoodland Hills
Senior AI EngineerWoodland Hills

TechDigital Group • Los Angeles (CA)

On-site
USD 120,000 - 160,000
LLM Application Engineer
LLM Application Engineer

Emonics LLC • Phoenix (AZ)

On-site
USD 110,000 - 180,000
Senior AI Engineer
Senior AI Engineer

TechDigital Group • Los Angeles (CA)

On-site
USD 130,000 - 160,000
LLM Applications Engineer
LLM Applications Engineer

SupportFinity™ • New York (NY)

Hybrid
USD 130,000 - 175,000
Senior LLM Systems Architect — AI for Security & Automation
Senior LLM Systems Architect — AI for Security & Automation

7AI • Boston (MA)

On-site
USD 140,000 - 200,000
AI Engineer - NC
AI Engineer - NC

LawPro.ai • North Carolina

On-site
USD 140,000 - 190,000
AI Engineer - TX
AI Engineer - TX

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000
AI Engineer - VA
AI Engineer - VA

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000