Research Data Scientist (Remote)

Innodata Inc.

Philippines

On-site

PHP 1,200,000 - 2,400,000

Full time

20 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Innodata Inc. seeks a Research Data Scientist to advance AI/natural language processing and generative AI initiatives. You will design experiments, develop evaluation methods and translate research findings into practical solutions for delivery teams.

The role emphasizes strong statistical and analytical capabilities, with collaboration across researchers, data scientists, AI/ML engineers, domain experts and client-facing staff. Leadership is advantageous but not required.

Qualifications

  • Master's or PhD in computer science, AI, ML, data science or related fields.
  • Proven research capability with publications, patents or conference presentations preferred.

Responsibilities

  • AI/ML and Generative AI Research: conduct independent and collaborative research in generative AI, LLMs, NLP, multimodal AI, ML, evaluation and data.
  • LLM and Model Evaluation: develop LLM evaluation frameworks, benchmarks, datasets and evaluation criteria.
  • Data Science and Statistical Research: collect, analyze, interpret large datasets; perform exploratory and statistical analyses.
  • AI Data and Dataset Development: develop datasets, sampling, annotation frameworks and data quality evaluation.
  • Research, Innovation and Collaboration: contribute to papers, reports and internal publications; present findings to stakeholders.

Skills

Python programming
Statistical methods
Experiment design
Research communication
Independent research

Education

Master's or PhD in related field
Assess research capability (publications/patents)

Tools

PyTorch
TensorFlow
Git

Job description

The Research Data Scientist works on research-driven AI and machine learning initiatives across generative AI, large language models, NLP, multimodal AI, model evaluation and AI data. The role designs experiments, develops evaluation methodologies, analyzes complex datasets, builds research prototypes and turns research findings into practical solutions the delivery organization can use.

This is a depth role. It calls for a strong research orientation, real statistical and analytical capability and the ability to work alongside researchers, data scientists, AI/ML engineers, domain experts and client-facing teams. Leadership and delivery experience are read as advantages here rather than requirements; the primary bet is on technical and research capability.

KEY RESPONSIBILITIES
AI/ML and Generative AI Research
  • Conduct independent and collaborative research in generative AI, LLMs, NLP, multimodal AI, machine learning, model evaluation and AI data
  • Formulate research questions and translate complex AI/ML problems into structured research methodologies and experiments
  • Design, execute and analyze experiments to evaluate and improve AI/ML models and solutions
  • Build analytical models, prototypes and research pipelines using Python and relevant ML frameworks
  • Stay current with emerging research, methodologies and developments in generative AI, LLMs, NLP, multimodal models and AI evaluation
LLM and Model Evaluation
  • Develop and implement LLM evaluation frameworks, benchmarks, datasets and evaluation criteria
  • Evaluate models for accuracy, robustness, bias, hallucination, reasoning, relevance and response quality
  • Conduct model benchmarking, error analysis, comparative analysis and performance evaluation
  • Work on RAG, SFT, RLHF and DPO, prompt engineering, fine-tuning, embeddings and LLM optimization as applicable
  • Identify model and data gaps and recommend improvements to model performance and reliability
Data Science and Statistical Research
  • Collect, clean, analyze and interpret large and complex structured and unstructured datasets
  • Perform exploratory data analysis, statistical analysis, hypothesis and significance testing, correlation analysis, sampling and error analysis
  • Develop data-driven insights and identify the patterns, trends and relationships relevant to AI/ML research
  • Apply appropriate statistical and quantitative methodologies to validate research findings
AI Data and Dataset Development
  • Develop and evaluate datasets, sampling methodologies, taxonomies, annotation frameworks, data quality frameworks and evaluation criteria
  • Analyze data quality and identify the issues affecting model performance
  • Collaborate with annotation, data engineering and AI/ML teams to improve training and evaluation data
  • Translate data and research findings into actionable recommendations for improving AI systems
Research, Innovation and Collaboration
  • Contribute to research papers, technical reports, whitepapers, patents, benchmarks and internal publications where applicable
  • Explore new methodologies, models, datasets and evaluation approaches, and identify where emerging research applies to real AI and data challenges
  • Work closely with researchers, data scientists, AI/ML engineers, annotation teams, domain experts and delivery teams
  • Present research findings and technical recommendations to senior technical stakeholders, and take part in client-facing technical discussions where required
CANDIDATE PROFILE
Education
  • Master's or PhD required in computer science, artificial intelligence, machine learning, data science, statistics, mathematics, computational science or a related discipline.
  • Assessed on research capability. A visible research record through publications, patents, conference presentations, open-source contributions or substantial applied AI/ML research is strongly preferred
Experience
  • 4 to 7 years of hands-on research experience in AI/ML, data science, NLP, generative AI or LLMs
  • Independent research capability. Able to formulate research questions, design experiments, analyze results and communicate findings without supervision
  • Demonstrated research track record. Publications, patents, conference presentations, open-source contributions or significant applied AI/ML research projects. Candidates with publications in reputed conferences or journals are preferred
  • Large-scale data experience. Working with large structured and unstructured datasets in a research setting
  • People or project leadership is an advantage. Research leadership, mentoring or program management is welcome but is not a requirement for this role
Certifications (Advantage, Not Required)
  • Technical: cloud or ML certification from AWS, Azure or Google Cloud
  • Research: peer-reviewed publication record carries more weight here than any certification
DOMAIN SPECIALIZATION

Candidates should demonstrate research or applied expertise in at least one of the following AI/data domains. Multiple specializations are an advantage.

  • Computational Linguistics & NLP: Computational Linguistics, Natural Language Processing (NLP), language and LLM evaluation
  • Law: Legal research, legal NLP, legal reasoning or AI applications in law
  • Medicine: Internal Medicine & Subspecialties, Surgical Specialties, Pediatrics & Geriatrics, clinical AI and medical NLP
  • Finance: Financial research, financial NLP, quantitative analysis or AI applications in finance
  • Advanced STEM: Engineering, Computer & Data Science, Agricultural & Food Science, Environmental Science & Technology, Applied Mathematics & Physics, Biology, Physics, Chemistry, Earth Science, Astronomy, Mathematics, Statistics, Computer Science, Systems & Game Theory
  • Multilingual & Dialects: Multilingual NLP, dialect variation, linguistic nuance, speech and language data
  • Process & Operations: Operations and Project Management, workflow analysis, process optimization and AI-enabled process design
  • Social Sciences: Psychology, Sociology, Anthropology, Political Science, Economics and related computational or research applications
Niche and highly relevant specializations

include multimodal and visual evaluation, speech/audio and live speech, prompting and model evaluation, computational linguistics and NLP, medical subspecialties, multilingual and dialect research, advanced STEM research, robotics/AI applications, and domain-specific AI evaluation.

Candidates do not need to cover all domains. Depth of expertise in one or more relevant specialization areas, combined with demonstrated research capability, is the primary consideration.

Technical Skills
  • Core programming. Strong proficiency in Python and SQL, with hands-on NumPy, Pandas and Scikit-learn
  • Deep learning frameworks. PyTorch or TensorFlow preferred
  • Statistical foundation. Machine learning algorithms, statistics and experimentation, data analysis and feature engineering, model evaluation and performance metrics, hypothesis testing and statistical inference
  • LLM and generative AI. Hands-on exposure to LLMs, NLP, generative AI and multimodal AI, with direct experience in one or more of RAG, LLM evaluation, prompt engineering, fine-tuning, SFT, RLHF or DPO, embeddings or model benchmarking
  • Engineering hygiene. Familiarity with Git and with cloud platforms such as AWS, Azure or Google Cloud is desirable
COMPETENCIES AND WORKING STYLE
  • Owns the methodological choice and can name the limitations of a finding without being asked
  • Has been misled by an evaluation metric at least once and learned from it
  • Still ships code personally and recently
  • Translates research for non-technical stakeholders without overclaiming certainty
  • Curious about the delivery problem, not only the research problem; connects findings to what the client actually needs
  • Works comfortably alongside annotation and data engineering teams rather than treating data as an input received from elsewhere
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Scientist (Gen/Agentic AI solutions)
Data Scientist (Gen/Agentic AI solutions)

Lafarge Africa Plc • Hinoba-an

On-site
PHP 7,524,000 - 11,285,000
Lead Data Scientist
Lead Data Scientist

V2 Solutions • Hinoba-an

On-site
PHP 1,200,000 - 1,800,000
AI Architect
AI Architect

Career Connect • Philippines

Hybrid
PHP 1,800,000 - 3,000,000
LLM Data Scientist
LLM Data Scientist

Greytip Software Private Limited • Hinoba-an

On-site
PHP 360,000 - 720,000
PH - GenAI Engineer
PH - GenAI Engineer

Thinking Machine • Philippines

Hybrid
PHP 900,000 - 1,500,000
Health benefits
Hybrid setup
Professional development budget
+1
Data Scientist
Data Scientist

Smart Communications, Inc. • Makati

On-site
PHP 900,000 - 1,500,000
Machine Learning Specialist (Research & Engineering)
Machine Learning Specialist (Research & Engineering)

Indra Philippines, Inc. • Philippines

On-site
PHP 900,000 - 1,500,000
Special Projects - AI & Automation
Special Projects - AI & Automation

Shopee • Manila

On-site
PHP 1,000,000 - 1,600,000
AI Developer – Backend & LLM Systems
AI Developer – Backend & LLM Systems

Salvo Software LLC • Mexico

On-site
PHP 1,228,000 - 2,150,000
AI Research and Innovation - Gen AI (Lead, Manager, Sr. Manager, Associate Director)
AI Research and Innovation - Gen AI (Lead, Manager, Sr. Manager, Associate Director)

Accenture in the Philippines • Philippines

On-site
PHP 1,500,000 - 2,500,000