Senior AI Engineer – LLM Systems

Evollabs Tech

Anupgarh

On-site

INR 4,000,000 - 6,000,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Evollabs Tech in India is seeking an experienced engineer to optimize LLM inference performance on custom AI hardware. You will enhance latency, throughput, and energy efficiency for datacenter-scale deployments.

The role covers profiling, transformer-level optimizations, quantization, KV caching, batching, and close collaboration with hardware and software teams to deploy cutting-edge AI solutions for demanding workloads.

Qualifications

  • Strong understanding of transformer architectures and LLM internals.
  • Experience with multiple modern LLMs (e.g., LLaMA, Mistral, Qwen, DeepSeek).
  • Proficient in inference optimizations: quantization, pruning, KV caching, batching, etc.
  • Strong Python skills and ML frameworks (PyTorch, JAX, or similar).
  • Experience with distributed systems and large-scale inference workloads.
  • Ability to profile and debug performance bottlenecks across hardware and software stacks.
  • Strong systems thinking across model, runtime, and hardware layers.

Responsibilities

  • Analyze, profile and optimize LLM inference performance across distributed, multi-chip systems.
  • Bring a deep understanding of transformer architectures, including dense and Mixture-of-Experts (MoE) models.
  • Evaluate and benchmark different LLMs on custom hardware.
  • Design and implement optimizations for attention mechanisms and other model-level improvements.
  • Collaborate with hardware and compiler teams to co-design efficient inference pipelines.
  • Build and maintain benchmarking frameworks to evaluate latency and throughput.
  • Analyze trade-offs between model architecture choices and system-level performance.
  • Contribute to model deployment strategies for large-scale datacenter environments.
  • Stay up to date with latest research in LLM architectures and inference optimization.

Skills

Transformer architectures
LLM internals
Python
PyTorch
JAX
Distributed systems
Performance profiling
System-level thinking

Tools

PyTorch
JAX

Job description

Description

We are a technology company focused on designing and developing advanced, customized server hardware solutions optimized for artificial intelligence workloads. Our mission is to accelerate AI innovation by delivering high-performance, scalable, and energy-efficient infrastructure for datacenter-scale inference.

Description

We are a technology company focused on designing and developing advanced, customized server hardware solutions optimized for artificial intelligence workloads. Our mission is to accelerate AI innovation by delivering high-performance, scalable, and energy-efficient infrastructure for datacenter-scale inference.

Our chips in development are purpose-built for large-scale AI inference and will be deployed in rack-level systems where multiple devices collaborate to deliver optimal latency, throughput, and efficiency. We are building the next generation of AI infrastructure and are looking for engineers who deeply understand how modern large language models behave at scale.

This role is focused on LLM systems, architecture, and performance. You will work at the intersection of model internals and hardware, ensuring that state-of-the-art models run efficiently on our platform.

Responsibilities
  • Analyze, profile and optimize large language model (LLM) inference performance across distributed, multi-chip systems
  • Bring a deep understanding of transformer architectures, including dense and Mixture-of-Experts (MoE) models
  • Evaluate and benchmark different LLMs (e.g., LLaMA, Mistral, Qwen, DeepSeek) on custom hardware
  • Design and implement optimizations for attention mechanisms (e.g., Flash Attention, grouped-query attention, sliding window attention)
  • Work on model-level optimizations such as quantization (INT8/FP8), KV-cache management, batching, and parallelism strategies
  • Collaborate with hardware and compiler teams to co-design efficient inference pipelines
  • Build and maintain benchmarking frameworks to evaluate latency, throughput, and scaling behavior
  • Analyze trade-offs between model architecture choices and system-level performance
  • Contribute to model deployment strategies for large-scale datacenter environments
  • Stay up to date with the latest research in LLM architectures and inference optimization
Requirements
  • Strong understanding of transformer architectures and LLM internals
  • Hands-on experience working with multiple modern LLMs (e.g., LLaMA, Mistral, Qwen, DeepSeek, etc.)
  • Deep knowledge of different LLM architectures such as dense models and Mixture-of-Experts (MoE) architectures
  • Familiarity with attention mechanisms and their optimizations
  • Experience with LLM inference optimization techniques (quantization, pruning, KV caching, batching, etc.)
  • Strong Python skills and experience with ML frameworks (PyTorch, JAX, or similar)
  • Experience with distributed systems and large-scale inference workloads
  • Ability to profile and debug performance bottlenecks across hardware and software stacks
  • Strong systems thinking and ability to work across model, runtime, and hardware layers
Preferred Qualifications
  • 8+ years of relevant experience in deep learning, AI systems, or performance engineering
  • Experience working close to hardware (GPU, TPU, or custom accelerators)
  • Experience with parallelism strategies (tensor parallelism, pipeline parallelism, expert parallelism)
  • Familiarity with datacenter-scale deployment, orchestration and inference servers (e.g., vLLM)
  • Background in performance engineering or systems optimization
What We’re Not Looking For

This role is not focused on prompt engineering, or application-layer GenAI development. Instead, it is centered on deep LLM internals, architecture, and inference performance at scale.

Why Join Us?
  • Work on cutting-edge AI hardware designed specifically for LLM inference
  • Solve challenging problems at the intersection of AI models and systems
  • Collaborate with a team pushing the boundaries of datacenter-scale AI performance
  • Make a direct impact on the future of AI infrastructure
Responsibilities
  • Analyze, profile and optimize large language model (LLM) inference performance across distributed, multi-chip systems
  • Bring a deep understanding of transformer architectures, including dense and Mixture-of-Experts (MoE) models
  • Evaluate and benchmark different LLMs (e.g., LLaMA, Mistral, Qwen, DeepSeek) on custom hardware
  • Design and implement optimizations for attention mechanisms (e.g., Flash Attention, grouped-query attention, sliding window attention)
  • Work on model-level optimizations such as quantization (INT8/FP8), KV-cache management, batching, and parallelism strategies
  • Collaborate with hardware and compiler teams to co-design efficient inference pipelines
  • Build and maintain benchmarking frameworks to evaluate latency, throughput, and scaling behavior
  • Analyze trade-offs between model architecture choices and system-level performance
  • Contribute to model deployment strategies for large-scale datacenter environments
  • Stay up to date with the latest research in LLM architectures and inference optimization
Requirements
  • Strong understanding of transformer architectures and LLM internals
  • Hands-on experience working with multiple modern LLMs (e.g., LLaMA, Mistral, Qwen, DeepSeek, etc.)
  • Deep knowledge of different LLM architectures such as dense models and Mixture-of-Experts (MoE) architectures
  • Familiarity with attention mechanisms and their optimizations
  • Experience with LLM inference optimization techniques (quantization, pruning, KV caching, batching, etc.)
  • Strong Python skills and experience with ML frameworks (PyTorch, JAX, or similar)
  • Experience with distributed systems and large-scale inference workloads
  • Ability to profile and debug performance bottlenecks across hardware and software stacks
  • Strong systems thinking and ability to work across model, runtime, and hardware layers
Preferred Qualifications
  • 8+ years of relevant experience in deep learning, AI systems, or performance engineering
  • Experience working close to hardware (GPU, TPU, or custom accelerators)
  • Experience with parallelism strategies (tensor parallelism, pipeline parallelism, expert parallelism)
  • Familiarity with datacenter-scale deployment, orchestration and inference servers (e.g., vLLM)
  • Background in performance engineering or systems optimization
What We’re Not Looking For

This role is not focused on prompt engineering, or application-layer GenAI development. Instead, it is centered on deep LLM internals, architecture, and inference performance at scale.

Why Join Us?
  • Work on cutting-edge AI hardware designed specifically for LLM inference
  • Solve challenging problems at the intersection of AI models and systems
  • Collaborate with a team pushing the boundaries of datacenter-scale AI performance
  • Make a direct impact on the future of AI infrastructure
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Server Engineer
Inference Server Engineer

Evollabs Tech • Anupgarh

On-site
INR 3,000,000 - 6,000,000
Member of Technical Staff
Member of Technical Staff

eBay • Bengaluru

On-site
INR 4,000,000 - 7,500,000
Senior/Principal Local Llm & Generative Ai Platform Engineer
Senior/Principal Local Llm & Generative Ai Platform Engineer

Parallelwireless • Maharashtra

On-site
INR 3,000,000 - 5,500,000
Interesting Job Opportunity: Carelon - Artificial Intelligence Engineer - Python/LLM
Interesting Job Opportunity: Carelon - Artificial Intelligence Engineer - Python/LLM

Carelon Global Solutions India • Bengaluru

On-site
INR 1,200,000 - 1,800,000
LLM Ops Engineer
LLM Ops Engineer

gnani.ai • Bengaluru

On-site
INR 2,800,000 - 4,800,000
Gen AI Engineer
Gen AI Engineer

HCLTech • Bengaluru

On-site
INR 2,500,000 - 4,500,000
AI Engineer (Python, GenAI/LLMs + ML Fundamentals)
AI Engineer (Python, GenAI/LLMs + ML Fundamentals)

Solutions By Text • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Prismforce Pvt Ltd - AI Engineer - LLM/RAG
Prismforce Pvt Ltd - AI Engineer - LLM/RAG

Prismforce • Maharashtra

On-site
INR 1,800,000 - 3,200,000
AI/ML Engineer
AI/ML Engineer

Jash Data Sciences Pvt. Ltd. • Pune District

On-site
INR 800,000 - 1,200,000
Competitive salary
Learning opportunities
Exposure to latest AI technologies
LLM Engineer (Large Language Models)
LLM Engineer (Large Language Models)

Fospe UK Ltd • Bengaluru

Hybrid
INR 2,500,000 - 5,200,000
Competitive compensation with bonuses
Hybrid work at Bangalore Innovation Cn
Health, dental, wellness insurance
+3