AI Engineer (RAG & On Prem LLMs) RAG, LLM, Python, PyTorch, LangChain, Neo4j, Docker, Kubernete[...]

Diverse CG Sp. z o.o. Sp.k.

Warszawa

On-site

PLN 212,044 - 339,270

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Diverse CG Sp. z o.o. Sp.k. in Warsaw is seeking a skilled professional to architect and implement Retrieval Augmented Generation (RAG) pipelines tailored for enterprise applications. The ideal candidate will join a dynamic team and spearhead projects involving cutting-edge AI technologies and open-source large language models.

The role requires at least 3 years in ML/NLP, with a focus on deploying LLM-based solutions in on-premises or hybrid settings. A strong background in Python and experience with vector databases are essential. This position offers challenges in a vibrant work environment.

Qualifications

  • At least 3 years of professional experience in ML/NLP roles, including 2+ years working with RAG systems.
  • Proven experience deploying and operating LLM-based solutions in on-prem or hybrid environments.
  • Hands-on experience with vLLM, LiteLLM, and open-source LLMs such as LLAMA 3.2.

Responsibilities

  • Architect, implement, and optimize end-to-end Retrieval Augmented Generation (RAG) pipelines.
  • Design and integrate retrieval mechanisms with generative models.
  • Implement and customize inference servers for scalable LLM serving.

Skills

ML/NLP experience
Python
Problem-solving
Communication
Knowledge of vector databases
Linux systems

Education

Bachelor's, Master's, or PhD in Computer Science

Tools

PyTorch
Hugging Face Transformers
LangChain
Neo4j
Docker
Kubernetes

Job description

Job location: Warsaw

Responsibilities:

  • Architect, implement, and optimize end-to-end Retrieval Augmented Generation (RAG) pipelines for enterprise use cases in on-premises environments
  • Design and integrate retrieval mechanisms (e.g. vector databases such as Neo4j) with generative models (e.g. LLAMA 3.2, Mistral)
  • Fine-tune and optimize retrieval and generation components to achieve high accuracy and low latency
  • Implement and customize inference servers using vLLM and LiteLLM for efficient and scalable LLM serving
  • Integrate open-source large language models with proprietary data sources and enterprise APIs
  • Design GPU-optimized, scalable on-prem infrastructure for model training and inference, ensuring security and data governance compliance
  • Collaborate with DevOps teams to containerize workflows using Docker and Kubernetes and automate MLOps pipelines
  • Apply performance optimization techniques such as quantization, pruning, and dynamic batching
  • Monitor system performance, troubleshoot bottlenecks, and ensure high availability
  • Work closely with data engineers and business stakeholders to translate business requirements into technical AI solutions in telco environments

Requirements:

  • At least 3 years of professional experience in ML/NLP roles, including 2+ years working with RAG systems
  • Proven experience deploying and operating LLM‑based solutions in on‑prem or hybrid environments
  • Hands‑on experience with vLLM, LiteLLM, and open‑source LLMs such as LLAMA 3.2, DeepSeek, or Mistral
  • Strong Python skills and experience with frameworks such as PyTorch, Hugging Face Transformers, and LangChain
  • Experience with vector databases (e.g. Neo4j)
  • Familiarity with Linux‑based systems and Red Hat OpenShift
  • Strong problem‑solving and analytical skills
  • Ability to clearly communicate complex AI concepts to non‑technical stakeholders
  • Bachelor's, Master's, or PhD degree in Computer Science, Artificial Intelligence, or a related field
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer (RAG & On Prem LLMs) RAG, LLM, Python, PyTorch, LangChain, Neo4j, Docker, Kubernete[...]
AI Engineer (RAG & On Prem LLMs) RAG, LLM, Python, PyTorch, LangChain, Neo4j, Docker, Kubernete[...]

DCG Poland • Warszawa

On-site
PLN 40,000 - 60,000
Lead GenAI Engineer – LLM, RAG & Life Sciences Domain
Lead GenAI Engineer – LLM, RAG & Life Sciences Domain

Mogi I/O : OTT/Podcast/Short Video Apps for you • Poland

On-site
PLN 296,000 - 382,000
AI Platform Engineer
AI Platform Engineer

ITDS • Wrocław

On-site
PLN 180,000 - 300,000
RAG & LLM Architect — On-Prem AI Pipelines (Warsaw)
RAG & LLM Architect — On-Prem AI Pipelines (Warsaw)

Diverse CG Sp. z o.o. Sp.k. • Warszawa

On-site
PLN 212,000 - 340,000
AI Architect
AI Architect

Luxoft • Poland

On-site
PLN 240,000 - 360,000
Private Medical & Dental care
Life Insurance
Internal Mobility program
Senior AI/ML Engineer - Remote, MLOps & RAG
Senior AI/ML Engineer - Remote, MLOps & RAG

Formamind sp. z o.o • Warszawa

On-site
PLN 180,000 - 300,000
RAG & LLM Architect — On-Prem AI Pipelines (Warsaw)
RAG & LLM Architect — On-Prem AI Pipelines (Warsaw)

DCG Poland • Warszawa

On-site
PLN 40,000 - 60,000
Architect ML/GenAI IRC296971
Architect ML/GenAI IRC296971

GlobalLogic • Kraków

On-site
PLN 180,000 - 340,000
Applied Scientist (LLM)
Applied Scientist (LLM)

SQUAD Ukraine Limited • Wrocław

Hybrid
Competitive salary packages
Guaranteed paid vacation
Private medical insurance
Architect ML/GenAI IRC296972
Architect ML/GenAI IRC296972

GlobalLogic • Kraków

On-site
PLN 80,000 - 120,000