Overview
We are seeking a versatile Data Engineer with 2+ years of experience to build and scale the data infrastructure powering our organization. You will develop robust pipelines and optimize architectures that bridge the gap between traditional analytics and next‑generation AI. In this role, you will work at the intersection of large‑scale data processing and modern AI, building the critical foundations for high-performance applications and agentic workflows. This position is based in Hyderabad, India.
Who you are
- Experienced Engineer: You have 2+ years of professional experience in data engineering or a backend‑heavy software engineering role.
- Python Expert: You possess expert‑level Python coding skills (Must Have), with a focus on writing clean, scalable, and asynchronous code.
- LLM Orchestrator: You have deep, hands‑on experience (Must Have) with LangChain or LangGraph to build sophisticated multi‑step chains and agentic systems.
- Vector DB Specialist: You have proven experience (Must Have) implementing and tuning Vector Databases for high‑volume RAG pipelines.
- Data Foundationist: You have a strong understanding of traditional data modelling, ETL/ELT processes, and working with SQL/NoSQL databases.
- AI‑Literate: You have a solid grasp of embedding models, tokenization, and modern information retrieval techniques.
- Agile & Innovative: You thrive in fast‑paced environments and enjoy staying updated with the rapidly evolving landscape of GenAI and search technologies.
What you’ll be doing
- Architect AI Data Pipelines: Design and maintain robust data ingestion and transformation pipelines tailored for LLM training, fine‑tuning, and Retrieval‑Augmented Generation (RAG).
- Build Agentic Workflows: Utilize LangGraph to develop complex, state‑managed AI agents and cyclical workflows that enhance automated user interactions.
- Optimize RAG Systems: Architect the retrieval layer of our AI applications, implementing efficient document embedding strategies and semantic search.
- Manage Vector Infrastructure: Implement and optimize Vector Databases (e.g., Pinecone, Weaviate, or Milvus) to ensure high‑performance data retrieval and storage.
- Scale Data Models: Create scalable data schemas that support both structured and unstructured data, ensuring seamless integration with our AI services.
- Performance Engineering: Identify and resolve latency bottlenecks in data retrieval and embedding generation to ensure real‑time AI responsiveness.
- Collaborate Cross‑Functionally: Partner with AI Researchers and Product Managers to transition experimental AI prototypes into production‑ready data products.
Seismic is an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to gender, age, race, religion, or any other classification which is protected by applicable law.