Machine Learning Software Engineer

Ema

Vancouver

On-site

CAD 120,000 - 170,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Ema in Vancouver is building AI employees that carry out complex workflows across enterprise applications. This role focuses on designing context, tools, and orchestration for multi-step agents, spanning documents, slides, images, audio and video, with careful test decisions on latency and cost.

You’ll join researchers and engineers across the Bay Area, Vancouver and India, contributing to post-training methods, evaluation, and production deployment.

Job description

  • Ema builds AI employees that carry out complex workflows across enterprise applications. Our ML team works on the loop that makes them better: production traces become data, data becomes training and evaluation, and better agents produce better traces. The hard part is deciding which intervention will improve behavior in the next real workflow
  • Harnesses and inference-time compute. Design context, tools, skills and orchestration for multi-step agents, including work across documents, slides, images, audio and video. Test where extra reasoning, search or verification earns its latency and cost. Build self-improvement loops with explicit permissions, evaluation gates and rollback
  • Agent post-training. Curate trajectories for SFT, optimize preferences, or run RL on real agent tasks. Investigate methods such as DPO, GRPO or DAPO where they fit; compare process and outcome supervision, shape rewards, and distill useful frontier behavior into smaller models. Measure whether gains transfer beyond the training environment
  • Environments and rewards. Turn enterprise workflows into reproducible training and evaluation environments: fixture tenants, simulated users who may get impatient and leave, and rewards grounded in verifiable outcomes. Find the shortcuts an agent can exploit before a training run optimizes for them
  • Data engines and evaluation. Mine production agent-steps for failures; build curated corpora and useful synthetic augmentation. Calibrate judges against human labels, construct behavior-level benchmarks from real workflows, and quantify data quality, performance uplift and reliability across stochastic runs
  • Retrieval, memory and context graphs. Connect enterprise information with user- and tenant-level learnings. Separate failures of retrieval from failures to use retrieved context; test what to retain, update and retrieve so that past experience improves the next decision
  • Quality per dollar. Build and evaluate routing, ensembles, caching and small-model specialization. Measure downstream task success alongside latency and cost; a cheaper model is useful only if the complete agent still succeeds
  • You’ll go deep in a subset of these areas. Projects combine applied research with the engineering needed to make the result work in production
  • Most projects develop over roughly four to six months, with useful improvements shipping along the way. You’ll define the problem and baseline, build the data or environment needed to test it, run experiments, and own serving and integration
  • Follow the system through deployment, monitoring and failure analysis until it is ready for a clear engineering handoff
  • You’ll work with researchers and engineers across the Bay Area, Vancouver and India. We’re hiring from junior through senior levels, with project scope matched to your experience

A master’s or PhD in a relevant field, or equivalent work or research experience. Papers, substantial open-source contributions, trained models and well-documented experiments can demonstrate that depthStatistical judgment. You can size an experiment, choose meaningful baselines and held-out tests, and account for variation across tasks, seeds and repeated runs. You can distinguish a real improvement from judge bias, data leakage or a benchmark shortcutEvidence of zero-to-one ownership. A system, model or research project you took from an ambiguous problem to a working result. We want to understand your contribution, the tradeoffs you made, and what changed when the work met real users or realistic tasksThese are project-specific strengths; post-training experience is optional, and no candidate needs the entire list. Bring a repository, paper, model or technical write-up that lets us examine how you think and what you builtDepth you can defend. Substantial work in at least one of agent/tool-use systems, post-training, reward modeling or RL environments, retrieval and memory, or evaluation design. Be ready to explain the mechanism, the alternatives you rejected and the failure modes you found. One area you can teach us beats five you’ve touchedProduction engineering judgment. You can debug across the model and system boundary, isolate a failure, and turn the result into maintainable production codeHonest measurement. You would rather retire your own approach after a clean negative result than ship an improvement that disappears under a stronger evaluationInteractive agent environments/harnesses for software engineering, web or tool use; large-scale trace analysis, data curation or synthetic generationPractical security work on prompt injection, data governance or permission boundaries for agents that can act and improve themselvesOpen-model post-training with TRL, veRL, OpenRLHF, or similar or a custom loop, especially debugging reward hacking or unstable optimizationDesigning systems to support complex, long-horizon agent work across a multitude of modalities and platformsServing with vLLM or SGLang, distillation, quantization, or multi-node GPU training

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal AI Engineer (Context, Agents and Context)
Principal AI Engineer (Context, Agents and Context)

Elastic • Ottawa

Hybrid
CAD 140,000 - 210,000
Health coverage
Flexible location
Vacation days
+3
Member of Technical Staff - ML Operations
Member of Technical Staff - ML Operations

Veeda AI • Toronto

On-site
CAD 120,000 - 150,000
RESEARCH SCIENTIST
RESEARCH SCIENTIST

Good Start Labs • Ottawa

Hybrid
CAD 90,000 - 130,000
Member of Technical Staff - World Models
Member of Technical Staff - World Models

Veeda AI • Toronto

On-site
CAD 120,000 - 180,000
Principal Scientist, Physical AI
Principal Scientist, Physical AI

Sanctuary Corp • Vancouver

On-site
CAD 140,000 - 220,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Aimlroles • Toronto

On-site
CAD 120,000 - 170,000
Senior Product Manager
Senior Product Manager

Ampliwork • Montreal (administrative region)

On-site
CAD 120,000 - 180,000
Research Intern, Small Language Models (Mitacs)
Research Intern, Small Language Models (Mitacs)

Onix (doing business as “Onix”; legal entity listed as 16445039 Canada Inc.) • Montreal (administrative region)

Hybrid
CAD 17,000 - 23,000
Agentic AI Optimization Developer
Agentic AI Optimization Developer

Equifax, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Member of Technical Staff - Robotics
Member of Technical Staff - Robotics

Veeda AI • Toronto

On-site
CAD 110,000 - 150,000