Senior Engineering Manager - Cloud Platform & SRE | Mission-Critical AI Software Platform

Techfellow Limited

Boston (MA)

On-site

USD 220,000 - 260,000

Full time

19 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Relocation funded

Job summary

Techfellow Limited in Boston is seeking a senior leader to own the cloud platform and SRE organisation for a mission-critical civil aviation software programme. You will shape architecture, drive production operations, and scale teams across the US and Europe while maintaining resilience and availability.

The role requires deep AWS and Kubernetes expertise, multi-region design, and a rigorous engineering culture.

Qualifications

  • 9+ years across software engineering, platform engineering or SRE with at least 3 years of engineering-management responsibility.
  • Own architecture, delivery and production operation of the AWS platform for a large-scale civil aviation programme.
  • Deep practical AWS and Kubernetes expertise, multi-region design and high-availability production architecture.
  • Establish an engineering model supporting ~4-nines availability with strong observability and incident response.
  • Lead the US platform and SRE organisation from a small team, with metrics, planning and procedures.

Responsibilities

  • Own the cloud platform and SRE organisation across the programme, including evolution of existing services and new infrastructure.
  • Drive CI/CD, progressive delivery, automated validation, controlled releases and safe production changes.
  • Develop 24/7 production support model balancing readiness with delivery pace.
  • Lead roadmap toward FedRAMP and multi-region deployment, translating regulatory needs into infra.
  • Collaborate with engineering leadership to implement tooling and AI-assisted development without reducing rigour.

Skills

Software leadership
SRE management
AWS platform
Kubernetes
Observability

Tools

AWS
Kubernetes
IaC

Job description

[Up to c. $260k Base Salary + Significant Equity | On-Site Working]
Role Overview

We’re representing a venture-backed technology company delivering software used within mission-critical US civil aviation infrastructure. As a major new platform moves towards production, the organisation is strengthening the engineering leadership responsible for the cloud foundation, reliability model and operational capability supporting it.

This person will take ownership of the cloud platform and SRE organisation supporting that programme, covering both the evolution of an existing production platform and the build-out of new infrastructure for a highly regulated environment. The immediate priority is delivery - establishing the architecture, engineering systems and operating model needed to move quickly without compromising resilience or availability. It is a senior player-coach position with substantial organisational scope. You will inherit teams in the US and Europe, grow the US organisation significantly, develop the 24/7 reliability function and remain technically credible enough to make architectural decisions and work directly with engineers while the team scales.

*The position is based in Boston five days per week, with funded relocation available.

Role Snapshot
  • Typically 9+ years across software engineering, platform engineering or SRE, including at least 3 years of meaningful engineering-management responsibility
  • Own the architecture, delivery and production operation of the AWS platform supporting a large-scale civil aviation programme, including both existing services and new infrastructure
  • Bring deep practical AWS and Kubernetes expertise, including multi-region design, infrastructure-as-code, container orchestration and high-availability production architecture
  • Establish an engineering model capable of sustaining approximately four-nines availability, with strong observability, failure handling, incident response, postmortems and continuous reliability improvement
  • Build and scale the US platform and SRE organisation from an initial small team, while leading through existing managers and establishing effective metrics, planning, feedback loops and operational processes
  • Create the SRE and on-call model required for continuous 24/7 production support, balancing operational readiness with the pace demanded by a major delivery programme
  • Strengthen CI/CD and progressive-delivery practices across the platform, including automated validation, controlled releases, rollback mechanisms and safe production change
  • Lead the roadmap towards a FedRAMP environment and multi-region production deployment, partnering across engineering and leadership to translate regulatory requirements into workable infrastructure
  • Develop shared engineering tooling, including practical use of modern AI-assisted development for code generation, review and testing, without losing engineering rigour or accountability
  • (Preferred) A strong software-engineering or computer-science foundation before moving into SRE/platform engineering, plus experience in hyperscale, regulated, defence, critical-infrastructure or similarly demanding production environments
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Engineer (AWS / Azure / GCP) - VP
Senior Cloud Engineer (AWS / Azure / GCP) - VP

Morgan Stanley • New York (NY)

On-site
USD 150,000 - 210,000
Senior DevOps Engineer | Dynamic Asset Management Leader
Senior DevOps Engineer | Dynamic Asset Management Leader

Techfellow Limited • New York (NY)

Hybrid
USD 350,000 - 425,000
DevOps Engineer / SRE - Platform Engineer / AI Spacetech
DevOps Engineer / SRE - Platform Engineer / AI Spacetech

Attis • Houston (TX)

On-site
USD 130,000 - 150,000
Flexible time-off policy
Cost-effective healthcare
401k matching plan
+1
Director of Infrastructure Engineering
Director of Infrastructure Engineering

Appsierra Group • United States

On-site
USD 350,000 - 500,000
Equity compensation eligibility
Performance-based bonuses
Health insurance reimbursement up to 1
+3
Remote | Director of Infrastructure Engineering — $350,000–$500,000/year
Remote | Director of Infrastructure Engineering — $350,000–$500,000/year

24-Mag Llc • Northern (KY), New York (NY)

Hybrid
USD 350,000 - 500,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Platform Engineering Manager
Platform Engineering Manager

Huxley • Boston (MA)

Remote
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

JLA Resourcing Ltd • New York (NY)

Hybrid
USD 150,000 - 200,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Chicago (IL)

Hybrid
USD 140,000 - 170,000
Comprehensive healthcare benefits
401(k) match
Paid Time Off (PTO)
+2
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Wand AI • Palo Alto (CA)

On-site
USD 180,000 - 250,000