Data Scientist

Vytwo

Dallas (TX)

On-site

USD 90,000 - 150,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible work from home options

Job summary

Vytwo is looking for a Data Scientist to build end-to-end ML and NLP solutions across structured and unstructured data. The role is based in Dallas, TX onsite, with flexibility for India-based consultants.

You’ll work on forecasting, predictive modeling, and cost/risk modeling, deploying models in production with Databricks and Azure. You’ll handle large-scale feature engineering, model lifecycle work, and real-time inference for text understanding and semantic retrieval, leveraging PyTorch,

Qualifications

  • 3+ years of data science experience with ML/NLP focus.
  • Hands-on expertise with Databricks, Delta Lake and MLflow.
  • Strong PySpark and Azure services proficiency (ADF, Azure ML, AKS).
  • Experience deploying models to production and building NLP pipelines.

Responsibilities

  • Build, deploy, and optimize ML models for forecasting and NLP tasks.
  • Develop large-scale feature engineering with PySpark and big data tools.
  • Create batch pipelines with model versioning and experiment tracking.
  • Design prompts, context graphs, and agentic workflows for LLM systems.
  • Implement real-time, low-latency inference for text understanding and retrieval.
  • Use Databricks and Azure for ETL, training, and deployment orchestration.
  • Collaborate with Git-based workflows and AI coding tools like Copilot.

Skills

Databricks
Delta Lake
MLflow
PySpark
Azure
ADF
Azure ML
AKS
Transformers
HuggingFace
PyTorch
TensorFlow
LangChain
Semantic Kernel
CrewAI
AutoGen
Prompt engineering
RAG
Vector databases
Kubernetes
REST APIs
Git
GitHub Copilot

Tools

Docker
Kubernetes
REST APIs
Git
GitHub Copilot

Job description

Vytwo is hiring a Data Scientist to build end-to-end machine learning and NLP solutions across both structured and unstructured data. The role is based in Dallas, TX (onsite), with the India team also open to local consultants eligible for INDIA. You will work on forecasting and predictive modeling as well as low-latency, production-ready NLP and LLM systems.

This position sits at the intersection of traditional analytics and modern generative AI, with an emphasis on deploying models and maintaining them in production using Databricks and Azure. Expect a blend of large-scale feature engineering, model lifecycle work, and real-time inference for text understanding and semantic retrieval.

What you’ll do
  • Build, deploy, and optimize ML models for predictive analytics, forecasting, classification, and regression.
  • Perform large-scale feature engineering using PySpark and big data tools.
  • Develop batch pipelines, including model versioning and experiment tracking.
  • Create cost estimation and risk/likelihood models using statistical and machine learning techniques.
  • Build NLP pipelines using deep learning frameworks such as PyTorch and TensorFlow (or similar).
  • Develop real-time, low-latency inference for use cases including text classification, embeddings, semantic search, summarization, and retrieval.
  • Design prompts, context graphs, and agentic workflows for LLM-based systems.
  • Apply prompt engineering, context engineering, and autonomous agent frameworks in production systems.
  • Use Databricks for ETL, feature engineering, model training, and orchestration.
  • Use Azure services for model deployment, data pipelines, and infrastructure.
  • Collaborate using Git-based workflows and leverage AI coding tools such as GitHub Copilot and Claude Code.
  • Implement model monitoring, observability, drift detection, and performance tracking.
Key focus areas
  • Structured Data (80–90%): predictive analytics, forecasting, cost estimation, likelihood modeling, and batch-oriented machine learning pipelines.
  • Text / Unstructured Data (NLP & GenAI): low‑latency real‑time systems using deep learning, LLMs, prompt engineering, and agentic AI frameworks.
What you bring
  • 3+ years of data science experience.
  • Strong hands‑on experience with Databricks (including Delta Lake, MLflow, and Job Orchestration).
  • Excellent PySpark skills for large-scale distributed data processing.
  • Proficiency in Azure services, including ADF, Azure ML, AKS, and Databricks on Azure.
  • Strong understanding of ML algorithms, statistical methods, and data analysis.
  • Deep learning experience with PyTorch and TensorFlow.
  • Experience with Transformers via HuggingFace.
  • Experience with model monitoring and ML observability.
  • Ability to write clean, optimized code and leverage AI code assistants.
  • Prompt engineering experience involving task prompts, chain of thought, tool calling, and retrieval prompts.
  • Context engineering experience including retrieval pipelines, RAG, memory management, and context structuring.
  • Knowledge of LLM-based agentic frameworks such as LangChain, Semantic Kernel, CrewAI, and AutoGen.
  • Vector databases and embedding models are plus.
Technologies you may use
  • PySpark, PyTorch, TensorFlow, Transformers (HuggingFace), Databricks
  • Delta Lake, MLflow, Job Orchestration, Azure
  • ADF, Azure ML, AKS, Databricks on Azure
  • Git, GitHub Copilot, Claude Code, LangChain, Semantic Kernel, CrewAI, AutoGen
  • Vector databases, embedding models, Docker, Kubernetes
  • Azure DevOps, GitHub Actions, Kafka, EventHub, Spark Streaming
  • REST APIs
Flexible work option
  • Flexible work from home options available.
Good to have
  • Experience with containerization (Docker, Kubernetes, AKS).
  • Experience deploying models to production (real‑time endpoints and REST APIs).
  • Knowledge of streaming technologies (Kafka, EventHub, Spark Streaming).
  • Understanding of CI/CD for ML using Azure DevOps and/or GitHub Actions.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

Jobtailor • Utah

On-site
USD 120,000 - 180,000
Lead AI Engineer - Bengaluru, INDIA
Lead AI Engineer - Bengaluru, INDIA

Vytwo • Dallas (TX)

Hybrid
USD 120,000 - 160,000
Lead AI Engineer - Bengaluru, INDIA
Lead AI Engineer - Bengaluru, INDIA

Vytwo Technologies Inc. • Prosper (TX)

Hybrid
USD 120,000 - 150,000
Flexible work from home options
Senior/Lead Data Software Engineer (Python, Spark, Azure)
Senior/Lead Data Software Engineer (Python, Spark, Azure)

EPAM Systems • Town of Poland (NY)

On-site
USD 120,000 - 150,000
Data Scientist (AI/ML)
Data Scientist (AI/ML)

Insight Global • Chicago (IL)

On-site
USD 120,000 - 180,000
Sr Data Engineer, Data Analytics & Intelligence, NA
Sr Data Engineer, Data Analytics & Intelligence, NA

Vantage Data Centers Management Company LLC • Denver (CO)

On-site
USD 130,000 - 155,000
Medical, dental, and vision coverage
Life and AD&D
401k with company match
+2
AI/ML Engineer
AI/ML Engineer

Winaxis LLC • Dallas (TX)

On-site
USD 120,000 - 180,000
Data Scientist — Real-Time NLP & LLM Systems on Azure
Data Scientist — Real-Time NLP & LLM Systems on Azure

Vytwo • Dallas (TX)

On-site
USD 90,000 - 150,000
Flexible work from home options
Senior AI Engineer
Senior AI Engineer

Flex Employee Services • Irvine (CA)

On-site
USD 64,747,000 - 146,136,000
Dental insurance
Health insurance
Referral program
+1
Data Scientist
Data Scientist

AI Squared • Washington

On-site
USD 110,000 - 140,000