Responsibilities
- Work on machine learning workflows involving embeddings, similarity search, and vector retrieval
- Assist in designing and optimizing semantic search and retrieval augmented generation pipelines
- Analyze performance, accuracy, and scalability of vector based systems
- Work with both structured and unstructured data for AI applications
- Support integration of vector databases into AI services and applications
- Write clean, well documented, and maintainable code
- Collaborate with engineering and research teams on AI system design
Requirements
Strong fundamentals in machine learning concepts
- Proficiency in Python
- Understanding of embeddings, similarity search, or NLP concepts
- Familiarity with Git and version control
- Ability to learn and work with new AI tools and systems
- Hands on experience with production grade AI infrastructure
- Practical exposure to vector databases and large scale retrieval systems
- Opportunity to contribute to an open source AI project
- Mentorship from experienced ML and systems engineers
- A strong portfolio project demonstrating real world AI skills
- Paid Internship
What We're Looking For
Key evaluation criteria for this role
1 This internship follows a project based evaluation approach. Candidates will be evaluated primarily based on a project created using the Endee vector database.
- Star the official Endee GitHub repository (https://github.com/endee-io/endee)
- Develop a well defined AI or ML project using Endee as the vector database
- Demonstrate a practical use case such as semantic search, RAG, recommendations, or similar AI workflows
- Host the project on GitHub
- Provide a clean and comprehensive README including:
- Project overview and problem statement
- System design or technical approach
- Explanation of how Endee is used
- Clear setup and execution instructions
Submission Requirements
Share the GitHub repository link as part of the application or evaluation process