Senior Manager, AI Foundation Model

Merlinlabs

Boston (MA)

Hybrid

USD 180,000 - 260,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) with match
Unlimited vacation
Hybrid work in Boston HQ

Job summary

Merlin Labs in Boston, MA is hiring a senior AI systems leader to drive world-model and post-training efforts for autonomous flight. You will design deterministic interfaces to the autonomy stack, lead a small engineering team, and define evaluation criteria and safety requirements for model releases.

This role supports hybrid work with a preference for Boston-based relocation. The position requires extensive experience in AI systems, PyTorch, and production-grade robotics or aerospace

Qualifications

  • Degree in Computer Science, Artificial Intelligence, Data Science, Computer Engineering, Applied Math, or related field.
  • 8+ years building AI systems.
  • 3+ years leading technical teams or owning a major model program.
  • Proven track record shipping AI-powered models into production in regulated or real-time environments.
  • Strong PyTorch experience; able to read and reason about C++ real-time systems.
  • Clear technical communication for engineers, safety, and regulators.

Responsibilities

  • Own Merlin's foundation and world-model work — architecture selection and roadmap.
  • Lead and mentor a small team of world-model engineers; set technical bar and review culture.
  • Design model interfaces to the autonomy stack with deterministic verifier-compatible outputs.
  • Define success criteria and build evaluation harnesses, taxonomy, and regression suites.
  • Establish uncertainty quantification and out-of-distribution detection as first-class outputs for safety.
  • Benchmark learned planning against current rule-based systems across mission profiles.
  • Collaborate with Certification and Systems Engineering to keep designs defendable to regulators.
  • Track external research frontier and make disciplined choices on adoption.

Job description

Merlin (NASDAQ: MRLN) is a publicly traded aerospace and defense company building a non-human pilot to deliver full-stack autonomy for any aircraft from takeoff to touchdown. The Merlin Pilot autonomy system powers a growing range of aircraft and mission profiles and has been proven through hundreds of autonomous flights from Merlin's global flight test facilities, including Kerikeri, New Zealand; Quonset Point, Rhode Island; and soon, Bedford, Massachusetts. Headquartered in Boston, Merlin is expanding its organization to accelerate the development and deployment of its autonomy platform, helping customers solve some of aviation's most pressing challenges, from pilot shortages to improving flight safety. Backed by some of the world's leading investors prior to its public listing, Merlin continues to advance the certification and commercialization of autonomous flight across commercial and defense aviation.

  • You have built learned decision-making systems that left the lab and ran on real hardware with real consequences.
  • You are fluent in modern model architecture and post-training, but you are not a benchmark chaser — you have been in the room when a learned system was asked to justify itself to people who sign off on safety, and you know the difference between a model that performs well and a model whose behavior you can characterize.
  • You want to work on a problem where “it works most of the time” is not a result.
  • Technical strategy: own Merlin's foundation and world-model work — architecture selection, build-vs-adapt decisions, post-training approach, and the capability roadmap that supports it.
  • Team leadership: lead and mentor a small team of world-model and post-training engineers; set the technical bar and the review culture for model work across AI Core.
  • System interface: design the model interface to the rest of the autonomy stack — structured, schema-constrained plan outputs that a deterministic verifier can accept or reject, never free-form actuator authority.
  • Evaluation: define what “good” means before training begins — build the evaluation harness, capability taxonomy, and regression suite that gate every model release, in partnership with the Data/Sim/Release pillar.
  • Safety-relevant outputs: establish uncertainty quantification and out-of-distribution detection as first-class model outputs, not afterthoughts — downstream safety monitoring depends on them.
  • Benchmarking: deliver an honest, reproducible comparison between learned planning and Merlin's current rule-based behavior planning across representative mission profiles, including the cases where the learned approach loses.
  • Certification partnership: work with Systems Engineering, Certification, and the Chief Architect to keep model design inside what is defensible to a regulator, and to shape what “defensible” will mean for learned components.
  • Research judgment: track the external research frontier and make disciplined calls about what Merlin adopts, builds, or ignores.
  • Degree in Computer Science, Artificial Intelligence, Data Science, Computer Engineering, Applied Math, or a related subject.
  • 8+ years building AI systems, with 3+ years leading technical teams or owning a major model program.
  • Proven team management experience shipping high-tech, AI-powered models into production — hiring and developing AI engineers, setting technical direction and priorities, and owning delivery from research through deployment.
  • Demonstrated ownership of a learned system that shipped into a physical, real-time product — robotics, autonomous vehicles, aerospace, or industrial autonomy.
  • Depth in at least two of: world models and learned dynamics; sequence models applied to planning or control; post-training (SFT, preference optimization, RL fine-tuning); structured or constrained generation.
  • Rigorous evaluation practice: you have built eval harnesses that caught regressions before customers did, and you can explain why a model's aggregate metric improved while a specific behavior got worse.
  • Strong PyTorch; comfortable reading and reasoning about the C++ real-time systems your models feed.
  • You write clearly. Architecture decisions here get read by systems engineers, safety engineers, and regulators — not only by other AI engineers.
  • Experience with learned components in a certified or regulated product (DO-178C, ISO 26262, IEC 62304).
  • Background in classical planning, behavior trees, MCTS, or hierarchical task networks — you'll be replacing and interoperating with exactly these.
  • Familiarity with aviation domain structure: flight phases, ARINC 424 procedures, ATC phraseology.
  • Publications or open-source contributions in embodied AI, world models, or robot learning.
  • We welcome remote applicants for this role, with a preference for candidates based in or willing to relocate to Boston, MA for a hybrid work schedule at our HQ.

Merlin Labs offers an innovative, entrepreneurial, and team-focused startup environment. We also offer a top-notch benefits package (health, dental, life, unlimited vacation, and 401k with match) and work/life integration. Being part of the Merlin team allows you to become part of a small team that supports professional development while working together to achieve our mission.

Merlin Labs is an equal opportunity employer and values diversity. We do not discriminate on the basis of race, religion, color, national origin, genetic information, sex (including pregnancy), gender, gender identity and expression, sexual orientation, age, marital status, military service or obligation or disability status, or any other characteristic protected by law. All job offers are contingent upon the candidate passing background and reference checks.

At this time, we are unable to provide visa sponsorship or consider candidates who require visa transfers. Applicants must be authorized to work in the United States without the need for visa sponsorship now or in the future.

In compliance with federal law, all persons hired will be required to verify identity and eligibility to work in the United States and to complete the required employment eligibility verification form upon hire.

If you require reasonable accommodation in completing an application, interviewing, completing any pre-employment testing, or otherwise participating in the employee selection process, please direct your inquiries to: people@merlinlabs.com

Merlin Labs does not accept unsolicited resumes from any source other than directly from candidates.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Manager, AI Foundation Model
Senior Manager, AI Foundation Model

Lever, Inc. • Boston (MA)

Hybrid
USD 230,000 - 330,000
Health insurance
Dental
Life insurance
+2
Senior Manager, AI Foundation Model
Senior Manager, AI Foundation Model

Merlin • Boston (MA)

Hybrid
USD 180,000 - 280,000
Remote work possible
Hybrid schedule in Boston HQ
Senior Manager, AI Foundation Model
Senior Manager, AI Foundation Model

Merlin Labs • Boston (MA)

On-site
USD 180,000 - 250,000
Health insurance
Dental insurance
Life insurance
+4
Senior Manager, AI Foundation Model
Senior Manager, AI Foundation Model

Quiet Capital • San Francisco (CA)

On-site
USD 230,000 - 330,000
Health insurance
Dental
Life insurance
+2
Staff Software Engineer, AI Foundation Model
Staff Software Engineer, AI Foundation Model

Merlinlabs • Boston (MA), San Francisco (CA)

On-site
USD 180,000 - 240,000
Catered lunches
Snacks and beverages on-site
401k with match
+1
Staff Engineer, World Model Development
Staff Engineer, World Model Development

Merlinlabs • Boston (MA)

On-site
USD 160,000 - 230,000
Catered lunches
Snacks
Beverages
+3
Staff Software Engineer, AI Foundation Model
Staff Software Engineer, AI Foundation Model

Merlin • Boston (MA)

On-site
USD 180,000 - 260,000
Catered lunches
Snacks and beverages
Health, dental, life insurance
+2
Staff Engineer, World Model Development
Staff Engineer, World Model Development

Merlin • Boston (MA)

On-site
USD 150,000 - 230,000
Catered lunches
Snacks and beverages
On-site perks
Director, AI Core Software
Director, AI Core Software

Quiet Capital • San Francisco (CA)

On-site
USD 252,000 - 360,000
Catered lunches
Snacks and beverages
401k with company match
+2
Lead AI Dataloop and Release Engineer
Lead AI Dataloop and Release Engineer

Merlin Labs, Inc. • Boston (MA)

Hybrid
USD 220,000 - 290,000
Catered lunches
Snacks & drinks
Benefits package