Platform Delivery & Reliability Engineer (Remote) - 29337

Huntington Ingalls Industries

Colorado Springs (CO)

Hybrid

USD 119,574 - 235,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Enlighten is seeking a Senior Platform Delivery & Reliability Engineer to own the rollout of our data lakehouse platform across a large government enterprise. This role blends platform engineering, SRE and delivery leadership to deploy, debug, and improve the system while building processes that make every deployment faster and more reliable.

You will collaborate with the Ingest, Query, Application and Testing teams, government customers and site personnel to ensure predictable rollouts, quick

Qualifications

  • Must obtain and maintain a U.S. Government Security Clearance.
  • Extensive experience deploying, operating production Kubernetes clusters.
  • Proven track record delivering complex distributed systems to production.

Responsibilities

  • Plan, coordinate, and execute deployments across 50 production Kubernetes clusters.
  • Debug and troubleshoot issues to root cause.
  • Contribute patches and automation improvements.
  • Catalog and triage field issues; maintain a living knowledge base.
  • Enable and improve site reliability: monitoring, alerting, incident response.
  • Coordinate with stakeholders and mentor engineers.
  • Other duties as assigned.

Skills

Excellent verbal and written comms
Cross-team collaboration
Problem solving
Leadership

Education

Bachelor's degree in related field
Master's degree in related field

Tools

Kubernetes
Terraform
AWS
Azure
GCP
GitLab CI
Go
Python
Bash

Job description

Platform Delivery & Reliability Engineer (Remote) - 29337

Enlighten, honored as a Top Workplace from USA Today, is a leader in big data solution development and deployment, with expertise in cloud-based services, software and systems engineering, cyber capabilities, and data science. Enlighten provides continued innovation and proactivity in meeting our customers’ greatest challenges.

We recognize that the most effective environment for your projects doesn’t always look the same. Our hybrid work approach ensures that you can make lasting relationships with your team and collaborate in-person to get the job done—while having the flexibility to work from home when needed to achieve focused results.

Why Enlighten? At Enlighten, our team’s unwavering work ethic, top talent and celebration of innovative ideas have helped us thrive. We know that our employees are essential to our company’s success, so we seek to take care of you as much as you take care of us. Here are a few highlights of our benefits package:

  • 100% paid employee premium for healthcare, vision and dental plans.
  • 10% 401k benefit.
  • Generous PTO + 10 paid holidays.
  • Education/training allowances.

Anticipated Salary Range: $119,574.00-$235,000.00. The salary range for this role is intended as a good faith estimate based on the role's location, expectations, and responsibilities. When extending an offer, Enlighten takes a variety of factors into consideration which include, but are not limited to, the role's function, internal equity and a candidate's education or training, work experience, certifications and key skills. Occasionally positions/roles may include additional non-recurrent compensation and will be addressed by the recruiter during the interview process.

Job Description

Enlighten is looking for a Senior Platform Delivery & Reliability Engineer to own the rollout of our data lakehouse platform across a large, multi-site government enterprise, currently ~50 production Kubernetes clusters and growing. This role is a rare hybrid of platform engineer, SRE, and delivery lead. You can deploy the platform, debug anything you encounter in the field, feed what you learn back to the engineering teams, and help fix underlying issues in the code base. Just as importantly, you can step back from any individual issue and fix the system that produced it by building the processes, tooling, and communication channels inside Enlighten that make every deployment faster and less painful than the one before it.

You will be a full member of the Infrastructure team, working daily with our Ingest, Query, Application and Testing teams as well as government customers and site personnel. Success in this role looks like: rollouts across the enterprise happen predictably and efficiently, issues found in the field are cataloged, communicated, and resolved quickly, and the friction that slows deployments steadily disappears.

  • Plan, coordinate, and execute deployments and upgrades of the data lakehouse platform across 50 production Kubernetes clusters in customer environments.
  • Debug and troubleshoot critical issues anywhere in the stack (infrastructure, Kubernetes, platform services, data services, and applications) and drive them to root cause.
  • Contribute patches, configuration changes, and automation improvements directly back to the platform.
  • Catalog and triage issues discovered in the field, communicate them clearly to the Infrastructure, Ingest, Query, and Application teams, and maintain a living knowledge base of failure modes, fixes, and runbooks.
  • Enable and improve site reliability for fielded environments: monitoring, alerting, incident response, and continuous reliability improvement.
  • Identify friction and dysfunction in how deployments happen, such as unclear handoffs, communication gaps, and repeated manual work; then design, implement, and institutionalize the processes that eliminate them (release checklists, readiness reviews, escalation paths, cross-team communication cadences).
  • Continuously improve deployment tooling and automation so rollouts become faster, safer, and more repeatable.
  • Coordinate with a large set of stakeholders (the engineering teams, government programs, security, and site personnel) and keep them informed.
  • Mentor engineers, both junior and senior, on debugging, deployment, and operational excellence.
  • Other duties as assigned.
Minimum Qualifications
  • Clearance Requirement: Must obtain and maintain a U.S. Government Security Clearance, but not required on day one; U.S. Citizenship required.
  • 9 years relevant experience with Bachelors in related field; 7 years relevant experience with Masters in related field; or High School Diploma or equivalent and 13 years relevant experience.
  • Deep, hands-on experience deploying, operating, and debugging production Kubernetes clusters and their ecosystem (networking, storage, service meshes, observability, volume management).
  • Proven record of delivering complex distributed systems into production across many environments or sites: not just building platforms, but landing them with customers.
  • Elite troubleshooting and analytical skills across the full stack: Linux systems, hosts, networks, security, containers, and application services.
  • Experience with infrastructure as code (e.g., Terraform) and modern cloud environments (e.g., AWS, Azure, GCP).
  • Experience with CI/CD pipelines (e.g., GitLab CI) and proficiency in scripting or programming (e.g., Go, Python, Bash).
  • Working knowledge of SRE practices: monitoring and alerting, incident management, blameless postmortems, and runbook development.
  • Demonstrated experience creating or improving engineering and delivery processes that other teams actually adopted; you can point to a workflow that exists because you built it.
  • Excellent verbal and written communication skills; able to translate deep technical issues for engineers, leadership, and customers, and comfortable coordinating a large number of people across organizational boundaries.
  • Work Location: *Remote or Hybrid. This role is fully remote unless you are located near one of our offices in Columbia, MD; San Antonio, TX; Boise, ID; Greenville, SC; or Augusta, GA, where a hybrid schedule applies. Note: Work models are subject to change based on business needs.
Preferred Requirements
  • Experience deploying or operating large-scale data platforms and lakehouse technologies (e.g., Spark, Trino/Presto, Kafka, NiFi, object storage, Iceberg/Delta/Hudi)
  • Experience delivering into DoD, IC, or other federal environments, including STIG-hardened, disconnected, or air-gapped deployments and familiarity with the RMF/ATO process
  • Experience with Kubernetes Operators/Controllers development
  • Prior release management, delivery lead, field engineering, or deployment engineering experience on a multi-team program
  • Understanding of agile software development methodologies and use of standard software development tool suites (e.g., YouTrack, GitLab, Nexus)
  • DoD 8140 / 8570 compliance certifications may be required in this position as directed by the customer

We have many more additional great benefits/perks that you can find on our website at www.enlighten.com.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps Engineer - 30149
DevOps Engineer - 30149

Enlighten • San Antonio (TX)

On-site
USD 100,000 - 130,000
Healthcare covered
401k 10%
PTO + 10 holidays
+1
Platform Engineer (Hybrid) - 27674
Platform Engineer (Hybrid) - 27674

Mission Technologies, a division of HII • Columbia (MD)

Hybrid
USD 115,000 - 145,000
100% paid employee premium for healthcare
10% 401k benefit
Generous PTO + 10 paid holidays
+1
DevOps Engineer (Hybrid) - 29376
DevOps Engineer (Hybrid) - 29376

Mission Technologies, a division of HII • Columbia (MD)

Hybrid
USD 115,000 - 150,000
Health insurance
401k
PTO + holidays
+1
Platform Engineer (Hybrid) - 26835
Platform Engineer (Hybrid) - 26835

Mission Technologies, a division of HII • San Antonio (TX)

Hybrid
USD 97,000 - 135,000
100% paid employee premium for healthcare
10% 401k benefit
Generous PTO + 10 paid holidays
+1
Platform Engineer (Hybrid) - 30077
Platform Engineer (Hybrid) - 30077

Huntington Ingalls Industries • San Antonio (TX)

Hybrid
USD 70,000 - 100,000
Healthcare premium
401k benefit
PTO + holidays
+1
DevOps Engineer (Hybrid) - 29376
DevOps Engineer (Hybrid) - 29376

Huntington Ingalls Industries • Columbia (MD)

On-site
USD 115,000 - 150,000
Healthcare coverage
401k 10%
PTO and holidays
+1
Platform Engineer (Hybrid) - 28480
Platform Engineer (Hybrid) - 28480

Huntington Ingalls Industries • Columbia (MD)

Hybrid
USD 90,166 - 115,000
100% paid employee premium for healthcare
10% 401k benefit
Generous PTO + 10 paid holidays
+1
Platform Engineer (Hybrid) - 27786
Platform Engineer (Hybrid) - 27786

Mission Technologies, a division of HII • Columbia (MD)

Hybrid
USD 115,000 - 145,000
100% paid employee premium for healthcare, vision, and dental plans
10% 401k benefit
Generous PTO + 10 paid holidays
+1
Platform Engineer (Hybrid) - 27029
Platform Engineer (Hybrid) - 27029

Huntington Ingalls Industries • Columbia (MD)

Hybrid
USD 111,662 - 145,000
100% paid employee premium for healthcare, vision, and dental
10% 401k benefit
Generous PTO + 10 paid holidays
+1
DevOps Engineer - 30149
DevOps Engineer - 30149

HII • San Antonio (TX)

On-site
USD 100,000 - 130,000
Healthcare + vision & dental coverage
10% 401k benefit
Paid time off + 10 holidays
+1