Job description:
We are looking for an AI Engineer with strong AI/ML and GenAI systems expertise to build and produce next-generation AI systems. This role is for someone who can go beyond using frameworks someone who understands how models and AI systems work internally, can design the architecture around them, evaluate them rigorously, and take a prototype all the way to a scalable, production-grade system.
AI/ML & Multimodal AI
- Strong first-principles understanding of Machine Learning, Deep Learning and Computer Vision
- Deep understanding of LLMs and Vision-Language Models (VLMs) tokenization/encoding, vision & language representations, multimodal alignment and how semantic information is fused
- SFT for LLMs/VLMs, PEFT/LoRA and fine-tuning strategies
- Model optimization: quantization (AWQ, INT4, HQQ), distillation, batching and inference optimization
- vLLM and high-performance model serving
- Reward modelling, custom/verifiable rewards, DPO, RLHF and GRPO
- Hands-on experience with PyTorch, Hugging Face Transformers and TRL
RAG, Retrieval & AI Agents
- Build advanced RAG / GraphRAG / VisionRAG systems
- Retrieval & ranking: BM25, semantic retrieval, hybrid search, vector databases
- Understand and implement RRF, MRR, Recall, Precision, F1, RAGAS and other evaluation approaches
- Query optimization using HyDE, query expansion and rewriting
- Agentic systems using LangGraph, LangChain, CrewAI, Agno or equivalent
- Understanding of MCP, its communication/transport mechanisms, tool calling and context engineering
- Structured output generation, schema validation and reliable tool execution
Evaluation & ML Engineering
- Build evaluation harnesses, automated testing and benchmarking frameworks
- Design golden datasets, taxonomies and evaluation datasets
- Perform model benchmarking, error analysis and root-cause analysis of model failures
- Design data preprocessing, transformation and ML pipelines
- Understand model quality vs. latency vs. memory vs. cost trade-offs
- Experience with distributed ML systems using Ray or equivalent.
Production AI Systems & Architecture
- Design end-to-end AI/ML system architectures
- Taking rapid prototypes robust production systems
- Build Python/FastAPI microservices and REST APIs
- Docker/containerization, webhooks and SSE
- CI/CD, Git/GitHub and production engineering practices
- Design scalable pipelines involving queues, asynchronous processing and distributed workloads
- Hands-on exposure to multi-GPU training/inference is highly desirable
An ideal candidate will have the following Background and Skills:
Background:
- Undergraduate Degree in any quantitative discipline such as engineering or science from a Top-Tier institution. MBA is a plus.
- Minimum 3 years of relevant experience in GEN AI.
- Strong hands-on experience in AI/ML infrastructure, cloud platforms, and production-grade ML systems.
- Prior experience working with AWS or GCP cloud services, particularly services such as AWS Bedrock, SageMaker, EC2, S3, SQS, Lambda, ECR, EKS, CloudWatch, or GCP Vertex AI.
- Experience with GPU infrastructure, CUDA, multi-GPU environments, and distributed training is required.
- Hands-on experience with containerized and Kubernetes-based environments, including Docker and Kubernetes.
- Experience working with LLMs, model training, fine-tuning, inference, and deployment using modern ML frameworks and tools.
- Experience building and deploying scalable APIs and ML services in production environments.
- Strong understanding of Linux, networking/API fundamentals, REST APIs, and CI/CD practices.
Skills:
- Strong hands-on experience with Python, PyTorch, Hugging Face Transformers, TRL, FastAPI, Docker, vLLM, and Ray.
- Very good understanding of SQL/PostgreSQL, Vector Databases, and Redis
- Hands-on experience with Kubernetes and distributed computing environments.
- Strong understanding of GPU/CUDA fundamentals, multi-GPU systems, and distributed training.
- Experience with observability and monitoring tools, particularly Prometheus, Grafana, Loki, Promtail, and CloudWatch.
- Experience with AWS/GCP cloud infrastructure and relevant AI/ML services.
- Strong understanding of Git/GitHub, REST APIs, and CI/CD pipelines.
- Ability to design, build, deploy, and monitor scalable AI/ML systems and production inference services.
Location:
Chennai, Hybrid work mode (4 days Work from Office and 1 day Work from Home)
Rewards:
- Attractive compensation
- Rapid and clear growth path
- Interaction with industry experts
- On job training and skill development
Eucloid offers an expedited growth path along with a compensation package which is among the best in the industry.
Additional Information:
If you require reasonable accommodation during the application or selection process, please do not hesitate to reach out to Eucloid HR at careers@eucloid.com