Job Title- Data engineer Senior associate
Years of experience -4-9yrs
Location - Bangalore/Hyderabad/Kolkata
Shoft timings-2-11PM
Skills requirement - Database-Sql, Mango DB, AI , Gen AI , AI engineer
Role Overview
We are looking for a Senior Associate Data Engineering & AI to design, develop, and manage scalable data platforms supporting analytics, reporting, and AI-driven applications. The ideal candidate will have strong experience with Oracle, SQL databases, MongoDB, Databricks, and database administration, along with an understanding of vectorization, embeddings, vector search, and Generative AI data patterns.
This role combines hands-on data engineering, database management, performance optimization, and the preparation of enterprise data for AI and machine-learning use cases.
Key Responsibilities
- Design and develop reliable ETL/ELT pipelines using SQL, Python, PySpark, and Databricks.
- Build and maintain data models across Oracle, relational databases, MongoDB, and lakehouse platforms.
- Develop and optimize complex SQL queries, stored procedures, views, indexes, and database objects.
- Administer databases, including user access, security, backup and recovery, monitoring, patching, and performance tuning.
- Build Databricks workflows using Spark SQL, Delta Lake, and medallion data layers.
- Prepare structured and unstructured data for analytics, machine learning, and Generative AI applications.
- Support AI use cases involving embeddings, vectorization, semantic search, vector databases, and Retrieval-Augmented Generation (RAG).
- Implement data-quality checks, metadata management, security controls, and data-governance standards.
- Troubleshoot production database and data-pipeline issues and perform root-cause analysis.
- Collaborate with data architects, application teams, analysts, data scientists, and AI engineers.
- Mentor junior team members and contribute to data engineering and database best practices.
Required Skills
- Strong hands-on experience with Oracle Database, SQL, and PL/SQL.
- Experience with relational database design, query optimization, indexing, partitioning, and performance tuning.
- Working knowledge of MongoDB, including document modeling, aggregation, indexing, and administration.
- Experience with Databricks, Apache Spark, Spark SQL, PySpark, and Delta Lake concepts.
- Experience building batch or near-real-time data pipelines.
- Knowledge of database administration, including backup, recovery, security, monitoring, migration, and high availability.
- Understanding of AI-ready data pipelines, embeddings, vector search, similarity matching, and vector databases.
- Proficiency in Python or another data-engineering scripting language.
- Strong analytical, troubleshooting, communication, and documentation skills.
Preferred Skills
- Exposure to Generative AI, large language models, RAG, prompt-based applications, or AI/ML platforms.
- Experience with Databricks Vector Search, MongoDB Atlas Vector Search
- Familiarity with Azure, AWS, or Google Cloud.
- Experience with orchestration and streaming tools such as Airflow, Azure Data Factory, Kafka, or Databricks Workflows.
- Knowledge of Git, CI/CD, automated testing, and DevOps practices.