Turn this role into an interview — a resume and cover letter built around what this employer wants.
Hume AI in New York City seeks a Senior Software Engineer for the Data Plane team to build and optimize the on-premises inference and workflow platform powering large-scale AI workloads. You will own end-to-end system design, drive performance, reliability, and resource efficiency, mentor teammates, collaborate with research and product, and document architectures for production readiness.
This role offers strong ownership, opportunities to shape high-impact systems, and close collaboration with
Hume AI is looking for a systems-oriented engineer to help build Hume’s on-premises inference and workflow platform for large-scale AI workloads. Join us in the heart of New York City and contribute to our endeavor to ensure that AI is guided by human values, the most pivotal challenge (and opportunity) of the 21st century.
Hume AI is a Series B startup dedicated to building artificial intelligence that is directly optimized for human well-being. As the first company to release speech language models, we’re focused on expanding our research to encompass audio understanding models and evaluation platforms for enterprises.
Our goal is to enable a future in which technology draws on an understanding of human emotional expression to better serve human goals. As part of our mission, we also conduct groundbreaking scientific research, publish in leading scientific journals like Nature, and support a non-profit, The Hume Initiative, that has released the first concrete ethical guidelines for empathic AI (www.thehumeinitiative.org). You can learn more about us on our website (https://hume.ai/) and read about us in WIRED, Forbes, and Venturebeat.
As a Senior Software Engineer on the Data Plane engineering team, you will help build Hume’s on-premises inference and workflow platform for running large-scale AI workloads. You will work closely with backend engineers, research scientists, machine learning engineers, product managers, and infrastructure engineers to bring new models and inference capabilities into production.
This is a systems-oriented role focused on reliability, performance, and efficient resource utilization. You will operate with a high degree of autonomy, make sound architectural decisions, and own systems throughout their lifecycle—from initial design and implementation to deployment, observability, and production support.
Design, build, and maintain the data plane software that executes AI models and voice data workflows.
Develop capabilities such as declarative workflow execution, parallel and fan-out processing, inference-engine integrations, tracing, and execution-status reporting.
Integrate new models and runtimes while maintaining consistent performance, reliability, and operational behavior.
Diagnose complex production issues across application, operating-system, container, networking, and hardware boundaries.
Profile and optimize latency, throughput, memory usage, and resource utilization.
Improve the observability, resilience, testability, and operational safety of the platform.
Collaborate with research and product teams to turn emerging model capabilities into reliable production systems.
Write clear technical documentation for the systems and features you build.
Significant professional experience building server-side or systems software.
Strong Linux fundamentals and hands-on experience troubleshooting and profiling production systems.
Experience in one or more relevant systems areas, such as distributed systems, networking, filesystems, concurrency, resource management, or computer architecture.
Professional experience with at least one systems-oriented programming language, such as Rust, C, or C++.
A strong sense of ownership and the ability to make pragmatic engineering decisions independently.
Excellent written and verbal communication skills. Engineers at Hume document the systems and products they build.
The ability to use AI-assisted coding tools effectively to accelerate delivery while retaining ownership of the result—including being able to explain, validate, debug, and modify the resulting code independently when necessary.
Demonstrated experience designing or building systems with a control-plane/data-plane architecture.
Experience with PyTorch or other machine-learning runtimes.
Experience with GPU inference, CUDA, ONNX, TensorRT, or performance profiling for accelerated workloads.
Experience with Python, or a willingness to learn it. Model execution and integration often require a small amount of supporting Python code.
Contributions to backend, systems, or machine-learning infrastructure open-source projects.
Familiarity with how neural networks are executed, including model graphs, tensor operations, and forward propagation.
Experience building container images and maintaining CI/CD pipelines.
Experience deploying software into on-premises or customer-managed environments.