Applied Scientist (LLM)

SQUAD

United States

Remote

USD 140,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Remote work within Ukraine
Medical insurance from day one
21 paid vacation days per year
Educational leaves
Free English classes

Job summary

SQUAD is seeking an experienced Applied Scientist with a strong background in Large Language Models to develop high-performance Generative AI features across Cloud and Edge environments. You will drive the transition from research to production by optimizing local inference through model compression and quantization for private, real-time Edge performance.

Your role includes engineering scalable RAG architectures, multi-agent systems for Cloud deployment, and end-to-end research lifecycle tasks

Qualifications

  • 3+ years of commercial experience in Machine Learning, focusing on NLP/LLM.
  • Strong programming experience in Python and ML libraries.
  • Hands-on experience deploying LLMs to production (PEFT/LoRA, RAG, or agentic frameworks).

Responsibilities

  • Design and implement prompt orchestration, fine-tuning (SFT/RLHF/DPO), and autonomous agent workflows.
  • Curate large-scale training data from text, image, and audio sources.
  • Identify patterns in model hallucinations and visualize evaluation metrics.
  • Tune hyperparameters and improve inference speed/accuracy via PEFT and prompt engineering.
  • Collaborate with Product and Data Engineering teams to integrate LLM features.
  • Track progress using standard benchmarks and internal KPIs.
  • Stay at the forefront of transformer variants and production-ready techniques.
  • Mentor junior team members to elevate expertise.

Skills

NLP/LLM expertise
Python
PyTorch
Hugging Face Transformers
PEFT/LoRA
Reinforcement Learning

Tools

PyTorch
Hugging Face Transformers
PEFT
ONNX/Triton

Job description

Team Summary

Our distributed team is looking for an experienced Applied Scientist with a strong background in Large Language models to develop high-performance Generative AI features across Cloud and Edge environments.

Job Summary

In this role you will drive the transition from research to production by optimizing local inference through model compression and quantization for private, real-time Edge performance, while also engineering scalable RAG architectures and multi-agent systems for Cloud deployment. Your daily responsibilities encompass the full research lifecycle, including formulating hypotheses, generating synthetic datasets, fine-tuning LLMs, and validating safety and alignment, ultimately culminating in technical reports.

Responsibilities and Duties
  • Design and implement advanced methods in prompt orchestration, fine-tuning (SFT/RLHF/DPO), and autonomous agentic workflows
  • Curate high-quality training data from large-scale text and multi-modal sources
  • Identify patterns in model hallucinations and visualize evaluation metrics for clear interpretation
  • Tune hyperparameters and improve inference speed/accuracy through PEFT (LoRA/QLoRA) and advanced prompt engineering
  • Collaborate with Product and Data Engineering teams to seamlessly integrate LLM features into the broader ecosystem
  • Track and report progress using industry-standard benchmarks (MMLU, HumanEval, etc.) and custom internal KPIs
  • Stay at the forefront of the field (e.g., State Space Models, new Transformer variants) and evaluate cutting-edge techniques for production readiness
  • Engage in continuous technical growth and mentor junior colleagues to elevate the team’s expertise
Qualifications and Skills
  • 3+ years of commercial experience in Machine Learning, with a specific focus on the NLP or LLM domain
  • Strong knowledge of Python3, NumPy, pandas, and modern text-processing libraries, PyTorch and Hugging Face (Transformers, PEFT, Accelerate)
  • Proficiency in PEFT/LoRA and Reinforcement Learning techniques
  • Deep understanding of attention mechanisms, tokenization, context window management, and embedding spaces
  • Practical experience in at least one of the following: Retrieval-Augmented Generation (RAG), Fine-tuning, or Agentic frameworks
  • Proven ability to manage and analyze massive datasets (>100GB) across text, image, and audio formats
  • Hands-on experience crafting high-fidelity datasets and building robust data pipelines
  • Expertise in prompt engineering, agentic framework design, and LLM pipeline orchestration
  • Experience deploying LLMs to production environments using Triton Inference Server, vLLM, TGI, or ONNX
  • Good written and spoken English
Nice to have
  • Practical experience with Pinecone, Weaviate, Milvus, or Chroma
  • Advanced quantization (GGUF, AWQ, EXL2), pruning, and knowledge distillation
  • Experience with LangChain, LlamaIndex, or AutoGen
  • Basic understanding of web/client-server architecture and streaming API responses (Asyncio, aiohttp)
  • Familiarity with RAGAS, DeepEval, or G-Eval
  • Experience using Docker, Kubernetes, and cloud GPU orchestration (e.g., Run:ai, Lambda Labs)
  • Knowledge of C++, Triton, or CUDA for custom kernel development
We offer multiple benefits that include
  • The environment of equal opportunities, transparent and value-based corporate culture and an individual approach to each team member
  • Competitive compensation and perks
  • Gig-contract
  • 21 paid vacation days per year, paid public holidays according to the Ukrainian legislation
  • Development opportunities like corporate courses, knowledge hubs, and free English classes as well as educational leaves
  • Medical insurance is provided from day one. Sick leaves and medical leaves are available
  • Remote working mode is available within Ukraine only
  • Free meals, fruits, and snacks when working in the office.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Engineer
AI Engineer

Quadcode • Georgia

Hybrid
USD 120,000 - 180,000
Senior Machine Learning Engineer (LLMs)
Senior Machine Learning Engineer (LLMs)

Albiware Inc. • Chicago (IL)

On-site
USD 140,000 - 210,000
Competitive salary
Generous PTO
Medical, dental, and vision coverage
+2
LLM Training & Model Development Engineer
LLM Training & Model Development Engineer

InOpTra Digital • United States

On-site
USD 90,000 - 120,000
Competitive salary
Opportunity for remote work
Health benefits
Senior/Principal Local LLM & Generative AI Platform Engineer
Senior/Principal Local LLM & Generative AI Platform Engineer

Parallel Wireless • United States

On-site
USD 180,000 - 280,000
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)

Bitus Labs • Irvine (CA)

On-site
USD 140,000 - 190,000
Senior/Principal Local LLM & Generative AI Platform Engineer
Senior/Principal Local LLM & Generative AI Platform Engineer

Parallelwireless • United States

On-site
USD 140,000 - 210,000
Lead AI Engineer
Lead AI Engineer

Zs Associates • Bellevue (WA)

On-site
USD 120,000 - 150,000
Comprehensive total rewards package
Robust skills-building programs
Multiple career progression paths
AIML Engineer
AIML Engineer

Qubeaxis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Performance bonus (up to 20% of base)
Equity participation
Health, dental, and vision insurance
+3
Member of ML Technical Staff
Member of ML Technical Staff

Pragmatike • San Francisco (CA)

On-site
USD 200,000 - 350,000
Artificial Intelligence Developer
Artificial Intelligence Developer

Five Data Products And Solutions • United States

Remote
USD 150,000 - 230,000