An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Hume AI is seeking a systems-oriented engineer to own the path from trained model to production inference in a high-performance environment in New York City. You will handle graph export, engine compilation, runtime integration, and production verification, collaborating with researchers and backend engineers to scale inference across products.
The role emphasizes performance, numerical correctness, reliability, and accelerator efficiency, with autonomy over lifecycle from design to deployment
Hume AI is looking for a systems-oriented engineer to own the path from trained model to production inference. Join us in the heart of New York City and contribute to our endeavor to ensure that AI is guided by human values, the most pivotal challenge—and opportunity—of the 21st century.
Hume AI is a Series B startup dedicated to building artificial intelligence that is directly optimized for human well-being. As the first company to release speech language models, we’re focused on expanding our research to encompass audio understanding models and evaluation platforms for enterprises.
Our goal is to enable a future in which technology draws on an understanding of human emotional expression to better serve human goals. As part of our mission, we also conduct groundbreaking scientific research, publish in leading scientific journals like Nature, and support a non-profit, The Hume Initiative, that has released the first concrete ethical guidelines for empathic AI www.thehumeinitiative.org. You can learn more about us on our website https://hume.ai/ and read about us in WIRED, Forbes, and Venturebeat.
As the first engineer dedicated full-time to inference backends at Hume, you will own the systems that take trained models from checkpoints to production inference: graph export, engine compilation, runtime integration, serving contracts, client libraries, artifact verification, and the performance and correctness of what runs in production.
You will work closely with research scientists, machine learning engineers, backend engineers, and the Data Plane team to bring new models and inference capabilities into production.
This is a systems-oriented role focused on performance, numerical correctness, reliability, and efficient use of accelerators. You will operate with a high degree of autonomy, make sound architectural decisions, and own systems throughout their lifecycle—from initial design and implementation through deployment, observability, optimization, and production support.
As the first engineer dedicated full-time to inference backends at Hume, you will own the systems that take trained models from checkpoints to production inference. This includes graph export, engine compilation, runtime integration, serving contracts, artifact verification, and production performance and correctness.
You will also own Hume’s internal inference platform, which serves many of our in-house models across products. This includes deployment, routing, load balancing, health checking, observability, capacity management, and safe model rollout.
You will work closely with research scientists, machine learning engineers, backend engineers, and product teams to turn new model capabilities into reliable production systems.
This is a systems-oriented role focused on performance, numerical correctness, reliability, and efficient use of accelerators. You will operate with a high degree of autonomy and own systems from design through production.