Contact with us through our representative or submit a business inquiry online.
Data Scientist with Bachelor’s degree in Computer Science, Computer Information Systems, Information Technology, or a combination of education and experience equating to the U.S. equivalent of a Bachelor’s degree in one of the aforementioned subjects.
Job Duties and Responsibilities:
- Collaborate with product owners, data scientists, data engineers, analysts, architects, and business stakeholders to understand data, artificial intelligence, and analytics requirements.
- Contribute to translating business requirements into analytical tasks, technical specifications, success metrics, acceptance criteria, and implementation steps.
- Collect, clean, profile, analyze, and interpret structured, semi-structured, and unstructured datasets to identify trends, patterns, anomalies, and business opportunities.
- Perform exploratory data analysis, statistical analysis, hypothesis testing, correlation analysis, segmentation, and feature engineering using established analytical methods.
- Develop, train, tune, test, and evaluate machine-learning models for regression, classification, clustering, forecasting, recommendation, and anomaly-detection use cases.
- Compare model performance using metrics such as precision, recall, F1 score, AUC, RMSE, and MAE, and prepare dashboards, reports, and visualizations to communicate the results.
- Develop and support Retrieval-Augmented Generation pipelines using enterprise data, embeddings, semantic search, hybrid search, reranking, vector databases, and large language models.
- Apply prompt-engineering, grounding, citation-validation, structured-output, and response-evaluation techniques to improve the accuracy and reliability of Generative AI applications.
- Develop, test, and support single-agent and multi-agent workflows using LangChain, LangGraph, LlamaIndex, or comparable Agentic AI frameworks.
- Support human-in-the-loop approvals, confidence thresholds, fallback mechanisms, retry policies, and testing for hallucinations, bias, prompt injection, sensitive-data exposure, and unauthorized tool usage.
- Develop Python-based components and reusable APIs for data processing, machine learning, Generative AI, and Agentic AI applications.
- Contribute to the development and maintenance of data models and analytical schemas using MongoDB, relational databases, cloud data warehouses, and data lakes.
- Develop and support ETL and ELT pipelines for data ingestion, cleansing, transformation, validation, normalization, enrichment, and analytical preparation.
- Support batch, streaming, Change Data Capture, and event-driven data pipelines, including data-quality checks, schema validation, logging, metadata management, and monitoring.
- Participate in testing, deployment, versioning, monitoring, troubleshooting, and documentation of machine-learning models, AI applications, and data pipelines while coordinating with cross-functional, onshore, and offshore teams.
Technologies Involved / Skills required for the position:
- Effective communication, collaboration, analytical, problem-solving, and documentation skills for working with cross-functional technical and business teams.
- Ability to understand business requirements and translate them into analytical tasks, technical specifications, success metrics, acceptance criteria, and implementation steps.
- Proficiency in Python and SQL, with experience using Pandas, NumPy, PySpark, or comparable technologies to collect, clean, profile, transform, and analyze structured, semi-structured, and unstructured data.
- Working knowledge of exploratory data analysis, statistical analysis, hypothesis testing, correlation analysis, segmentation, and feature-engineering techniques.
- Experience developing and evaluating regression, classification, clustering, forecasting, recommendation, and anomaly-detection models using Scikit-learn, XGBoost, LightGBM, TensorFlow, PyTorch, or comparable frameworks.
- Knowledge of model-evaluation metrics, including precision, recall, F1 score, AUC, RMSE, and MAE, with experience creating dashboards, reports, and visualizations using Power BI, Tableau, Looker, Omni, or comparable tools.
- Experience developing Retrieval-Augmented Generation solutions using embeddings, semantic search, hybrid search, reranking, vector databases, large language models, and enterprise data.
- Working knowledge of prompt engineering, grounding, citation validation, structured outputs, response evaluation, and guardrails for Generative AI applications.
- Experience developing and testing single-agent and multi-agent workflows using LangChain, LangGraph, LlamaIndex, or comparable Agentic AI frameworks.
- Knowledge of human-in-the-loop processes, confidence thresholds, fallback mechanisms, retry policies, hallucination evaluation, bias detection, prompt-injection prevention, and sensitive-data protection.
- Proficiency in developing Python-based data-processing, machine-learning, Generative AI, and Agentic AI components, with knowledge of reusable APIs using FastAPI, Flask, or comparable frameworks.
- Experience working with data models and analytical schemas using MongoDB, PostgreSQL, MySQL, SQL Server, BigQuery, Snowflake, Databricks, or comparable database and cloud data technologies.
- Experience developing and supporting ETL and ELT pipelines for data ingestion, cleansing, transformation, validation, normalization, enrichment, and analytical preparation.
- Working knowledge of batch, streaming, Change Data Capture, and event-driven pipelines using technologies such as Apache Airflow, Kafka, Google Cloud Pub/Sub, AWS SQS, SNS, or EventBridge.
- Understanding of testing, deployment, versioning, monitoring, troubleshooting, and documentation practices for machine-learning models, AI applications, and data pipelines, including familiarity with Docker, CI/CD, MLOps, and LLMOps.
Work location is Portland, ME with required travel to client locations throughout USA.
Rite Pros is an equal opportunity employer (EOE).