Senior AI Research Engineer

Phrase

United States

Remote

USD 140,000 - 205,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Phrase is seeking a Senior AI Researcher to lead a major research workstream focused on fine-tuning LLMs for high-quality translation and language tasks. You will design evaluation frameworks, build agentic workflows, and collaborate with Product and Engineering to translate research into production capabilities.

You will mentor junior researchers, contribute to hiring, and help shape the future of our language intelligence platform, working across a global, distributed team.

Qualifications

  • PhD in Computer Science or Computational Linguistics (or MSc with equivalent hands-on research).
  • Strong hands-on experience fine-tuning LLMs or transformer models.
  • Proven NLP expertise in translation, MT evaluation, or related fields.

Responsibilities

  • Design, train, and evaluate LLMs for translation and language-quality tasks.
  • Build and evaluate agentic workflows for multi-step reasoning and tool use.
  • Own a significant research workstream from data collection to production integration.
  • Mentor junior researchers and contribute to hiring discussions.

Skills

LLM fine-tuning
NLP/ML expertise
Python
Evaluation frameworks
Research leadership

Education

PhD in CS/Computational Linguistics

Tools

Hugging Face
PyTorch
Weights & Biases
vLLM
AWS SageMaker
GCP Vertex AI

Job description

Senior AI Researcher

At Phrase, we help open the door to global business by providing the world’s leading Language Intelligence Platform.

The Phrase Platform combines AI, agentic orchestration, and a headless, API-first architecture in one composable system. Beyond translation, it orchestrates and adapts content to culture, audience, channel, brand voice, and intended outcome in any language, for every audience. It applies the context that makes content perform in every market: quality standards, glossaries, prior translations, and cultural nuance. Every team, in every region, can ship content that is on-brand, on-point, and ready for any audience.

Phrase gives enterprises the intelligence to automate workflows, the freedom to connect their own tools and engines, and the control to govern global content at scale.

With a global team based in our offices, and remote colleagues across Europe, the UK, the US, and the APAC region, Phrase offers an international environment built on collaboration, innovation, and shared purpose.

Phrase’s AI Research team builds, trains, and evaluates the proprietary translation models and agentic workflows that power the platform’s language quality — live systems serving production traffic, not proof-of-concept work. This role takes ownership of a significant research workstream within that portfolio, extending Phrase’s live systems and shaping the direction of the team’s wider research. You’ll work across fine‑tuning and instruction‑tuning LLMs for domain‑specific machine translation, designing evaluation frameworks that combine automatic metrics and LLM‑as‑judge approaches, and building agentic workflows for multi‑step reasoning and tool use. It’s built for someone with strong hands‑on research experience who wants real scope to lead a body of work, mentor others, and grow their technical influence over time.

What you’ll be responsible for:
  • Design, train, and evaluate LLMs for translation and language-quality tasks, using fine‑tuning techniques such as LoRA and DPO and instruction‑tuned models.
  • Design and implement evaluation frameworks for translation quality, combining automatic metrics (BLEU, TER, ChrF, MQM, COMET), LLM‑as‑judge approaches, and hybrid pipelines.
  • Build and evaluate agentic workflows for NLP problems, spanning multi‑step reasoning, tool use, and structured output generation.
  • Own a significant workstream within the team’s model‑development programme, from data and training through evaluation and production integration.
  • Design and run reproducible experiments, documenting results and decisions clearly.
  • Work with Product, Engineering, and Solutions to turn research findings into shippable capabilities.
  • Benchmark frontier models — OpenAI, Gemini, Claude, and open‑source — against internal baselines.
  • Evolve the team’s quality‑evaluation systems, including style‑guide integration and customer‑led, outcome‑oriented metrics.
  • Shape decisions on active learning, feedback‑loop design, and data‑pipeline strategy for continuous model improvement.
  • Mentor junior researchers, review PRs and research write‑ups, and contribute to reading groups and deep dives.
  • Contribute to hiring, including take‑home review and technical interviews.
  • Represent your work in cross‑functional forums with Product and Engineering.
What you need:
  • Strong hands‑on experience fine‑tuning LLMs or large transformer‑based models; familiarity with pre‑training and preference‑tuning (RLHF/DPO) is desirable.
  • Solid NLP background, with demonstrable expertise in at least one of: machine translation, MT evaluation, information extraction, text generation, search optimization, or reinforcement learning.
  • Ability to design, build, and evaluate agentic systems, including tool‑calling agents, multi‑agent pipelines, or LLM‑orchestrated workflows.
  • Experience designing and running evaluation frameworks, including automatic metrics and human‑in‑the‑loop evaluation design.
  • Strong Python skills and experience with ML/NLP libraries such as Hugging Face or PyTorch.
  • Familiarity with workflow orchestration tools such as Flyte, Prefect, or Airflow (desirable).
  • PhD in Computer Science, Computational Linguistics, or a related field, or an MSc with equivalent hands‑on research or industry experience.
  • Industry experience with production NLP systems, including shipping models, handling data pipelines, and monitoring quality in live traffic (desirable).
  • Familiarity with the localization and MT domain, including MQM, COMET, post‑editing workflows, segment‑level quality signals, or TMS integrations (desirable).
  • Experience with experiment tracking tools such as Weights & Biases, and model‑serving or cloud ML infrastructure such as vLLM, TGI, AWS SageMaker, or GCP Vertex AI (desirable).
  • Experience with frameworks for structured agentic pipelines such as PydanticAI or LangGraph (desirable).
  • Strong publication record or open‑source contributions in NLP/ML (desirable).
  • Product‑minded, thinking about research in terms of what it enables and who it serves, not just whether the numbers went up.
  • A natural driver who identifies open problems and moves work forward without waiting to be pushed.
  • Growth‑oriented, motivated to expand technical scope and mentor others, with a path toward greater research leadership over time.
Our current tech stack in this role:
  • Python
  • Hugging Face
  • PyTorch
  • Weights & Biases
  • PydanticAI
  • LangGraph
  • vLLM / TGI
  • AWS SageMaker / GCP Vertex AI
  • Flyte / Prefect / Airflow
What you’ll get:
  • Work experience in a successful and growing global SaaS company
  • Be part of an international team in Europe, APAC, and the Americas
  • Expert colleagues in their field who are determined to build the best localization platform on the market, creating a world where language never limits opportunity
  • An agile work environment, where it is encouraged to take smart risks
  • Take part in a culture full of trust, support and loyalty, where respectful and open feedback is valued, and diversity is fully embraced
  • A positive, open‑minded, and innovative atmosphere
  • Support in your professional development and personal career goals
What’s on top:
  • 4 Company holidays additional to your regular holidays (1 day per quarter where the entire company is off to celebrate our achievements).
  • In addition, the company also provides employees with a Christmas break to allow you to spend time with family and friends without use of your vacation allocation.
  • Your birthday is off because it is important to celebrate you as well.
  • 2 Giveback days where you can support the local community, volunteer, and/or participate in charity events and activities.
  • Professional and extensive onboarding.
  • Enterprise Claude Licence
  • Additional local benefits depending on the entity you’re hired at, just ask your Talent Acquisition Partner

Phrase is committed to ensuring equal pay for equal work and work of equal value between women and men. Our aim is that our workforce will be truly representative of all sections of society and that each worker feels respected and able to give their best. The job requirements and salary range for this job have been established through a gender‑neutral job evaluation and classification process. This process objectively assesses the skills, responsibility, effort and working conditions required for each job. This systematic approach ensures fair pay regardless of who performs the work. We welcome applications from all qualified candidates regardless of their sex, race or ethnicity, disability, religion/belief, sexual orientation, gender identity or age. We value and welcome different perspectives, experiences and backgrounds as we believe that these differences make our team even stronger on our mission of opening the door to global business by giving everybody access to the content they need in the language they speak.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Web Developer, Marketing
Senior Web Developer, Marketing

Phrase • United States

Remote
USD 120,000 - 180,000
4 Company holidays (additional)
Christmas break
Your birthday off
+4
Lead AI Translation Research Engineer
Lead AI Translation Research Engineer

Phrase • United States

Remote
USD 140,000 - 205,000
Senior Staff Research Scientist | Voice
Senior Staff Research Scientist | Voice

DeepL • London (KY)

On-site
USD 180,000 - 240,000
ML Data & Platform Engineer
ML Data & Platform Engineer

Speechmatics • Cambridge (MA)

Hybrid
USD 150,000 - 190,000
Private Medical
Dental for you and family
Pension/401K matching
+4
ML Data & Platform Engineer
ML Data & Platform Engineer

Speechmatics • London (KY)

Hybrid
USD 110,000 - 170,000
Private Medical and Dental for you and
Global working opportunities
Generous holiday allowance
+3
Research Engineer/ Applied Scientist
Research Engineer/ Applied Scientist

Madrona Venture Labs • San Mateo (CA)

On-site
USD 252,000 - 308,000
Research Engineer/Research Scientist, Pre-training
Research Engineer/Research Scientist, Pre-training

Menlo Ventures • San Francisco (CA)

On-site
USD 340,000 - 425,000
Competitive compensation and benefits
Generous vacation and parental leave
Flexible working hours
Research Engineer, Machine Learning Systems
Research Engineer, Machine Learning Systems

Deepgram • United States

On-site
USD 100,000 - 150,000
Holistic health
Unlimited PTO
Learning stipend
+1
Software Engineer, Agent (Spanish speaking)
Software Engineer, Agent (Spanish speaking)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 170,000 - 250,000
Flexible PTO
Medical, dental, and vision benefits
Life insurance and disability benefits
+6
Solutions Architect
Solutions Architect

LILT (Production) • United States

Hybrid
USD 150,000 - 230,000