Gen AI Data Engineer

Tiger Analytics, LLC

United States

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading analytics consulting firm in the United States is seeking an experienced Machine Learning Engineer with Gen AI expertise. The role involves designing scalable data solutions, collaborating with global teams, and ensuring system performance. Implement your skills in a dynamic environment focused on advanced analytics and significant career development opportunities.

Qualifications

  • 8+ years of experience in data engineering or platform engineering.
  • Experience building scalable distributed data systems.
  • Familiarity with generative AI tools and techniques.

Responsibilities

  • Designing and implementing scalable data engineering and ML solutions.
  • Collaborating to deliver client-ready analytics solutions.
  • Building and deploying data platforms for real-time processing.

Skills

Proficiency in Python
Proficiency in SQL
Proficiency in PySpark
Strong communication skills
Experience working with onshore/offshore teams

Tools

Snowflake
AWS
Apache Airflow
GitHub
VS Code

Job description

Overview

Tiger Analytics is looking for experienced Machine Learning Engineers with Gen AI experience to join our fast-growing advanced analytics consulting firm. Our employees bring deep expertise in Machine Learning, Data Science, and AI. We are the trusted analytics partner for multiple Fortune 500 companies, enabling them to generate business value from data. Our business value and leadership has been recognized by various market research firms, including Forrester and Gartner.

We are looking for top-notch talent as we continue to build the best global analytics consulting team in the world.

Responsibilities

You will be responsible for:

  • Designing and implementing scalable data engineering and ML solutions for Gen AI and RAG use cases.
  • Collaborating with onshore and offshore teams to deliver client-ready analytics solutions.
  • Building and deploying pipelines and data platforms to support real-time and batch processing.
  • Ensuring system performance, reliability, security, and governance of data assets.
Technical Skills Required
  • Programming Languages: Proficiency in Python, SQL, and PySpark.
  • Data Warehousing: Experience with Snowflake, NOSQL and Neo4j.
  • Data Pipelines: Proficiency with Apache Airflow.
  • Cloud Platforms: Familiarity with AWS (S3, RDS, Lambda, AWS Batch, SageMaker processing Job, CloudFormation, etc.) or GCP (Vertex AI RAG, Data Pipeline, BigQuery, GKE).
  • Operating Systems: Experience with Linux.
  • Batch/Realtime Pipelines: Experience in building and deploying various pipelines.
  • Version Control: Experience with GitHub.
  • Development Tools: Proficiency with VS Code.
  • Engineering Practices: Skills in testing, deployment automation, DevOps/SysOps.
  • Communication: Strong presentation and communication skills.
  • Collaboration: Experience working with onshore/offshore teams.
Desired Skills
  • Big Data Technologies: Experience with Hadoop and Spark.
  • Data Visualization: Proficiency with Streamlit and dashboards.
  • APIs: Experience in building and maintaining internal APIs.
  • Machine Learning: Basic understanding of ML concepts.
  • Generative AI: Familiarity with generative AI tools and techniques.
Additional Expertise
  • Knowledge Graphs: Experience with creation and retrieval.
  • Vector Databases: Proficiency in managing vector databases.
  • Data Persistence: Ability to develop and maintain multiple forms of data persistence and retrieval methods (RDBMS, Vector Databases, buckets, graph databases, knowledge graphs, etc.).
  • Cloud Technologies: Experience with AWS, especially SageMaker, Lambda, OpenSearch.
  • Automation Tools: Experience with Airflow DAGs, AutoSys, and CronJobs.
  • Unstructured Data Management: Experience in managing data in unstructured forms (audio, video, image, text, etc.).
  • CI/CD: Expertise in continuous integration and deployment using Jenkins and GitHub Actions.
  • Infrastructure as Code: Advanced skills in Terraform and CloudFormation.
  • Containerization: Knowledge of Docker and Kubernetes.
  • Monitoring and Optimization: Proven ability to monitor system performance, reliability, and security, and optimize them as needed.
  • Security Best Practices: In-depth understanding of security best practices in cloud environments.
  • Scalability: Experience in designing and managing scalable infrastructure.
  • Disaster Recovery: Knowledge of disaster recovery and business continuity planning.
  • Problem-Solving: Excellent analytical and problem-solving abilities.
  • Adaptability: Ability to stay up-to-date with the latest industry trends and adapt to new technologies and methodologies.
  • Team Collaboration: Proven ability to work well in a team environment and contribute to a positive, collaborative culture.
GenAI Engineer Specific Skills
  • Industry Experience: 8+ years of experience in data engineering, platform engineering, or related fields, with deep expertise in designing and building distributed data systems and large-scale data warehouses.
  • Data Platforms: Proven track record of architecting data platforms capable of processing petabytes of data and supporting real-time and batch ingestion processes.
  • Data Pipelines: Strong experience in building robust data pipelines for document ingestion, indexing, and retrieval to support scalable RAG solutions. Proficiency in information retrieval systems and vector search technologies (e.g., FAISS, Pinecone, Elasticsearch, Milvus).
  • Graph Algorithms: Experience with graphs/graph algorithms, LLMs, optimization algorithms, relational databases, and diverse data formats.
  • Data Infrastructure: Proficient in infrastructure and architecture for optimal extraction, transformation, and loading of data from various data sources.
  • Data Curation: Hands-on experience in curating and collecting data from a variety of traditional and non-traditional sources.
  • Ontologies: Experience in building ontologies in the knowledge retrieval space, schema-level constructs (including higher-level classes, punning, property inheritance), and Open Cypher.
  • Integration: Experience in integrating external databases, APIs, and knowledge graphs into RAG systems to improve contextualization and response generation.
  • Experimentation: Conduct experiments to evaluate the effectiveness of RAG workflows, analyze results, and iterate to achieve optimal performance.

This position offers an excellent opportunity for significant career development in a fast-growing and challenging entrepreneurial environment with a high degree of individual responsibility.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Gen AI Data Engineer
Gen AI Data Engineer

Tiger Analytics • United States

Hybrid
USD 120,000 - 150,000
Significant career development opportunities
Work in a fast-growing entrepreneurial environment
Gen AI Engineer
Gen AI Engineer

Tiger Analytics Inc. • United States

On-site
USD 150,000 - 210,000
Gen AI Engineer
Gen AI Engineer

Tiger Analytics • United States

On-site
USD 150,000 - 190,000
Lead Data Scientist - GenAI
Lead Data Scientist - GenAI

Tiger Analytics, LLC • United States

On-site
USD 130,000 - 160,000
Senior Data Engineer
Senior Data Engineer

Tiger Analytics • Jersey City (NJ)

On-site
USD 120,000 - 150,000
Senior Data Scientist GenAI / RAG
Senior Data Scientist GenAI / RAG

Compunnel, Inc. • Houston (TX)

On-site
USD 120,000 - 150,000
Senior Data Scientist - Agentic AI
Senior Data Scientist - Agentic AI

Tiger Analytics, LLC • New York (NY)

On-site
USD 130,000 - 170,000
Lead Data Scientist - GenAI
Lead Data Scientist - GenAI

Tiger Analytics • New Jersey

On-site
USD 130,000 - 160,000
Data Engineer - Jersey City
Data Engineer - Jersey City

Tiger Analytics, LLC • Jersey City (NJ)

On-site
USD 120,000 - 160,000
Full Stack Gen AI Engineer
Full Stack Gen AI Engineer

Tiger Analytics • Town of Texas (WI)

Remote
USD 120,000 - 150,000