Get more replies from employers
Send a job-specific resume in minutes.
Synphony is the deployment layer for physical AI, bringing frontier policies to real-world environments like agriculture, manufacturing, mining, and oil & gas. You will own the learning problem end‑to‑end on the machine, defining tasks, collecting data, training policies, evaluating results, diagnosing failures, and iterating until a plant manager signs off.
The role emphasizes on-site collaboration with customers, handling data collection, model calibration, and integration into live automation
Synphony · On-site with customers, heavy travel · Agriculture · Manufacturing · Mining · Oil & Gas
Robot Learning Engineer
train a policy that does a task nobody has automated, put it on a robot in a customer's plant, and stand next to it while it runs a shift.
Synphony is the deployment layer for physical AI. We take frontier policies — VLAs, foundation models, learned controllers — and make them work inside industries that run on 40-year-old equipment. Not a lab, not a demo video. Wash-down environments, 110°F greenhouses, vibration, variable lighting, and operators who have done the task by hand for thirty years and will tell you exactly why your robot is wrong.
Robot learning is the whole job, and the emphasis is on robot. Every model you train runs on a physical arm in front of a customer who is measuring cycle time. If your ML experience has only ever touched benchmarks and leaderboards, or your background is web services, cloud platforms, or LLM applications, you will not be competitive here regardless of how strong that work is. One question filters most of it: have you trained something that moved real hardware, and did it work?
You own the learning problem end to end, on the machine. Define the task formally, decide what data would be sufficient, collect it, train the policy, evaluate it honestly, diagnose why it fails, and iterate until the success rate is high enough that a plant manager signs off. Then deploy it and watch it run.
The hard part is not the training run. The environment is not i.i.d., demonstrations are inconsistent, the reward is unspecified, the eval set is whatever came off the line that week, and the distribution shifts when the customer changes a supplier. You will spend more time deciding what to measure than optimizing what you measure.
Turn "the robot needs to trim this" into an action space, an observation space, a horizon, a success criterion, and a data budget. Choose imitation vs. RL vs. residual vs. classical control on the merits. Know when the honest answer is a fixture and a threshold, not a model.
Teleop rig design, synchronized multi-camera capture, timestamp discipline, calibration, action-representation consistency between recording and execution, normalization, curation. Most policy failures are data or calibration failures wearing a costume, and you will be the one who finds that out.
Fine-tune VLAs and imitation-learning policies — π₀/π₀.₅-class, ACT, diffusion and flow‑matching policies, and whatever supersedes them. LoRA, action tokenization, chunk horizons. Get them onto Jetson‑class edge hardware inside the latency budget.
A first‑class deliverable, not overhead. Offline metrics that actually predict online success, sim harnesses, held‑out contexts, confidence intervals on small samples, and the discipline to say "this regression is noise" with numbers behind it.
Perception, calibration, policy, controller, or fixture — attribute it correctly before you fix it. Ablate. Instrument. Run the experiment that discriminates between two hypotheses rather than the one that confirms your favorite.
Get the policy onto the customer's hardware, wire it into the control stack, respect the real‑time boundary — a 10 Hz policy does not sit inside a 4 ms loop, and you should know what goes between them — and commission it against their takt.
You don't need every line. The first one is not optional.
Robot learning moves monthly and someone who stops reading is obsolete in a year. Track the frontier — new VLA architectures, tactile and force representations, world models, long‑horizon manipulation — reproduce what matters, and hold a defensible view on which results are a genuine step change and which are a well‑shot demo. We fund the compute and the hardware to find out. Some of what you build here should be publishable, and some of it will be.
You'll travel, and you'll be on‑site somewhere hot, loud, dusty, and far from good coffee. Your beautiful eval curve will meet a conveyor running 15% faster than anyone documented. Data will arrive mislabeled, unsynchronized, and short. Research instincts that reward chasing an interesting tangent have to coexist with a customer expecting a working cell in eleven weeks, and you will have to choose. If that tension sounds miserable, this isn't your job. If it sounds like the only place the interesting problems actually are, we should talk.
You'll work on manipulation problems with no published solution, on data nobody else has, with a deadline that forces the question of whether the method actually works. Full ownership from formulation to production, not one slice of it. And every deployment feeds real operational data back into policies that improve every other site we run — in the industries that grow the food, dig the materials, and make the things. That's the point.