LLM / GenAI Engineer

Evlo AI

New York (NY)

On-site

USD 140,000 - 200,000

Full time

8 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Evlo AI is seeking a software engineer to build production-grade LLM systems, owning RAG architectures, agentive workflows, evaluation pipelines, and inference optimization for real users at scale.

You will work at the intersection of applied research and software engineering with ML and platform engineers to translate state-of-the-art techniques into reliable services.

Qualifications

  • 3–6 years of software engineering, with 1–2+ years shipping LLM/GenAI systems in production.
  • Strong Python skills, async code, API design, and typed codebases.
  • Hands-on production RAG experience: chunking, hybrid retrieval, reranking, hallucination mitigation.
  • Knowledge of transformer architectures, prompts, and when to fine-tune vs. prompt vs. retrieve.
  • Experience deploying models in cloud: AWS/GCP/Azure; Docker/Kubernetes.
  • BS/MS in CS or related field; practical industry experience acceptable.
  • Bonus: LangGraph, CrewAI, GPU profiling, multimodal models, or OSS LLM tooling.

Responsibilities

  • Design, build, and ship LLM-powered features - RAG pipelines, agents, structured extraction - using LangChain, LlamaIndex, or custom orchestration layers
  • Integrate and optimize vector stores (Pinecone, Weaviate, pgvector) and embedding pipelines for retrieval quality and latency targets
  • Build systematic evaluation harnesses: golden datasets, LLM-as-judge scoring, regression suites, and guardrail testing before every release
  • Fine-tune and adapt open-weight models (Llama, Mistral, Qwen) using LoRA/QLoRA and distillation on domain-specific data
  • Optimize inference for cost and latency - quantization, batching, caching strategies, and vLLM or TensorRT-LLM deployment
  • Instrument LLM applications with tracing, observability, and monitoring (LangSmith, OpenTelemetry, dashboards) to catch drift and degradation
  • Collaborate with product/design teams to translate ambiguous requirements into scoped, testable AI capabilities

Skills

Python
Async programming
LLM production
APIs
Distributed systems

Education

BS/MS in CS or related

Tools

LangChain
LlamaIndex
Weaviate
Pinecone
Docker
Kubernetes

Job description

About The Role

The role focuses on building production LLM systems - not demos. Expect to own RAG architectures, agentic workflows, evaluation pipelines, and inference optimization that serve real users at scale.


About The Role

The role focuses on building production LLM systems - not demos. Expect to own RAG architectures, agentic workflows, evaluation pipelines, and inference optimization that serve real users at scale.


The work sits at the intersection of applied research and software engineering: translating state-of-the-art techniques into reliable, observable, cost-efficient services alongside a team of ML engineers and platform engineers.


Key Responsibilities


  • Design, build, and ship LLM-powered features - RAG pipelines, agents, structured extraction - using LangChain, LlamaIndex, or custom orchestration layers

  • Integrate and optimize vector stores (Pinecone, Weaviate, pgvector) and embedding pipelines for retrieval quality and latency targets

  • Build systematic evaluation harnesses: golden datasets, LLM-as-judge scoring, regression suites, and guardrail testing before every release

  • Fine-tune and adapt open-weight models (Llama, Mistral, Qwen) using LoRA/QLoRA and distillation on domain-specific data

  • Optimize inference for cost and latency - quantization, batching, caching strategies, and vLLM or TensorRT-LLM deployment

  • Instrument LLM applications with tracing, observability, and monitoring (LangSmith, OpenTelemetry, custom dashboards) to catch drift and degradation

  • Collaborate with product and design teams to translate ambiguous requirements into scoped, testable AI capabilities


What We Are Looking For


  • 3–6 years of software engineering experience, with at least 1–2 years building and shipping LLM/GenAI systems in production

  • Strong Python engineering skills; comfort with async code, API design, and type-checked codebases

  • Hands-on experience with RAG in production: chunking strategies, hybrid retrieval, reranking, and hallucination mitigation

  • Practical knowledge of transformer architectures, prompt engineering limits, and when to fine-tune vs. prompt vs. retrieve

  • Experience deploying models on cloud infrastructure (AWS, GCP, or Azure) with Docker/Kubernetes

  • BS/MS in Computer Science, a related technical field, or equivalent practical experience

  • Bonus: experience with agent frameworks (LangGraph, CrewAI), GPU profiling, multimodal models, or contributions to open-source LLM tooling

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI/ML Engineer
AI/ML Engineer

RiskForce • Northern (KY)

On-site
USD 120,000 - 155,000
LLM Integration / LangChain Engineer
LLM Integration / LangChain Engineer

Zoho • United States

Remote
USD 120,000 - 180,000
Senior/Principal Local LLM & Generative AI Platform Engineer
Senior/Principal Local LLM & Generative AI Platform Engineer

Parallel Wireless • United States

On-site
USD 180,000 - 280,000
Senior LLMOps Engineer
Senior LLMOps Engineer

UNAVAILABLE • McLean (VA)

On-site
USD 180,000 - 240,000
AI Engineer / LLM Systems Engineer
AI Engineer / LLM Systems Engineer

Sphere Software • United States

Remote
USD 120,000 - 170,000
LLM Training & Model Development Engineer
LLM Training & Model Development Engineer

InOpTra Digital • United States

On-site
USD 90,000 - 120,000
Competitive salary
Opportunity for remote work
Health benefits
Application Engineer – LLM
Application Engineer – LLM

Salt Digital Recruitment • United States

On-site
USD 120,000 - 180,000
DevOps Engineer
DevOps Engineer

Rivago Infotech Inc • Charlotte (NC)

On-site
USD 120,000 - 180,000
Ai Software Architect Llm Agentic Systems Msys Tech India Pvt Ltd Chennai
Ai Software Architect Llm Agentic Systems Msys Tech India Pvt Ltd Chennai

Vibehackers • United States

On-site
USD 180,000 - 260,000
Senior LLM Engineer – GenAI / ML (Python, Langchain)
Senior LLM Engineer – GenAI / ML (Python, Langchain)

Executive Staff Recruiters / ESR Healthcare • United States

Remote
USD 41,000 - 77,000
Referral Bonus $150