Staff Platform Engineer (MANTL)

Alkami Technology, Inc.

United States

Remote

USD 140,000 - 175,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Remote-first environment
Unlimited PTO
401(k) with employer match

Job summary

Alkami Technology is seeking a Staff Platform Engineer to advance reliability across the MANTL platform. You will own monitoring, tracing, and performance improvements while shaping a roadmap independent of feature delivery timelines.

You will collaborate with Cloud Infra and application teams, setting technical direction and standards for platform reliability, and mentoring engineers on fault-injection and defensive design patterns.

Qualifications

  • Deep proficiency in TypeScript with production experience in a Node.js/TypeScript service environment.
  • Strong hands-on experience deploying, operating, and troubleshooting containerized workloads in Kubernetes.
  • Proven experience designing, building, and maintaining CI/CD pipelines, GitHub Actions preferred.
  • Proven experience designing monitoring, dashboard, and alerting strategy in an APM/observability tool, Datadog preferred.
  • Deep experience troubleshooting distributed system and microservice communication failures across a variety of integration points and protocols.
  • Strong experience with distributed tracing tools and practices at scale.
  • Working familiarity with relational databases, sufficient to diagnose and resolve complex query and schema-level performance issues.
  • Proven track record resolving significant application performance issues (caching, inefficient code paths, slow queries).
  • Experience designing and leading failure-mode testing (fault injection, resilience/chaos-style testing) and driving adoption of defensive patterns such as idempotency, retries, and circuit breakers.
  • Demonstrated ability to work independently on ambiguous, high-impact reliability problems and to set technical direction for others.
  • Excellent communication skills, with the ability to present technical root cause, risk, and remediation plans to engineering and non-technical leadership.
  • Experience mentoring other engineers.

Responsibilities

  • Lead investigation, troubleshooting, and resolution of the most complex reliability issues within the MANTL platform application code.
  • Set direction for how the platform identifies and addresses failure modes across integrations.
  • Own the design and evolution of monitoring, dashboards, and alerting strategy (Datadog preferred) across the platform.
  • Establish standards and lead implementation of distributed tracing across microservices.
  • Lead diagnosis and remediation of significant application performance issues (caching, slow queries).
  • Set direction for platform and application hardening practices, including fault-injection and resilience testing.
  • Own and improve CI/CD build pipeline architecture (GitHub Actions).
  • Guide deployment and troubleshooting of container-native (Kubernetes) workloads.
  • Define, prioritize, and drive execution of a roadmap of known reliability risks.
  • Establish and maintain documentation and runbook standards covering platform reliability issues.
  • Serve as the senior escalation point for complex platform reliability issues.
  • Define reliability targets (SLOs/SLIs) for key platform services and advise leadership on risk.
  • Provide technical mentorship and guidance to other Platform Engineers.

Skills

TypeScript proficiency
Node.js/TypeScript
Kubernetes
GitHub Actions
Datadog monitoring
Distributed tracing
Relational databases
Performance tuning
Chaos testing
Fault injection
System reliability
Mentoring

Education

Bachelor's degree in CS/Engineering

Tools

GitHub Actions
Datadog
OpenTelemetry

Job description

Alkami is the digital sales and service platform provider for U.S. banks and credit unions. Our unified Platform integrates onboarding, digital banking, and data and marketing—each solution can stand alone, but together they deliver more—to help institutions onboard, engage, and grow relationships. As the future shifts toward Anticipatory Banking, we help data-informed bankers meet the moment with technology that drives action.

Founded in 2009, we continue to be recognized for our intentional culture and tremendous growth (Best Place to Work in Fintech; Best & Brightest to Work For Nationally; and Comparably's Best Company Culture, Best Career Growth, Best Engineering Team, and Best Places to Work in Dallas, among others). We're building a culture where each Alkamist can perform to their highest potential, and we're always on the lookout for the best and brightest minds. If you're ready to experience the power of alchemy - transforming the ordinary into the extraordinary - come join one of the fastest growing SaaS companies in the U.S.

As a remote-first company, most of our positions can be remote in the US, except for key roles, which will be indicated in the Job Title.

Follow us on Glassdoor and LinkedIn!

The Staff Platform Engineer leads the effort to locate, make visible, and remediate sources of unreliability in the MANTL platform, including correctness problems that surface under failure conditions. This role works directly in the platform’s application codebase (TypeScript) and its container-native deployment environment, combining application-engineering skill with reliability-engineering practice at a level of scope and independence beyond the Sr Platform Engineer. In addition to resolving reactive issues as they surface, this role owns and prioritizes a standing roadmap of known reliability risks, driving that work forward on its own timeline. This role partners with, but is organizationally and functionally distinct from, both Cloud Infrastructure Engineering and product application engineering teams, and is expected to set technical direction and standards for platform reliability work.

Essential Duties & Responsibilities
  • Lead investigation, troubleshooting, and resolution of the most complex reliability issues within MANTL platform application code, including microservice communication failures and correctness issues that emerge under failure conditions

  • Set direction for how the platform identifies and addresses failure modes across its third-party and internal system integrations, establishing resilience strategies suited to each integration’s specific behavior

  • Own the design and evolution of monitoring, dashboards, and alerting strategy (Datadog preferred) across the platform, ensuring proactive visibility into emerging risks

  • Establish standards and lead implementation of distributed tracing across microservices to accelerate root-cause identification organization-wide

  • Lead diagnosis and remediation of significant application performance issues, including caching strategy, inefficient code paths, and query performance

  • Set direction for platform and application hardening practices, including fault-injection and resilience testing, and drive adoption of defensive design patterns to prevent recurrence

  • Own and improve CI/CD build pipeline architecture (GitHub Actions) supporting deployment of the platform

  • Guide deployment and troubleshooting of container-native (Kubernetes) workloads as part of resolving complex platform reliability issues

  • Define, prioritize, and drive execution of a roadmap of known reliability risks, independent of feature-delivery timelines

  • Establish and maintain documentation and runbook standards covering platform reliability issues, root causes, and remediations

  • Serve as the senior escalation point for complex platform reliability issues, partnering with Cloud Infrastructure Engineering and application engineering leadership on issues that cross domain boundaries

  • Define reliability targets (SLOs/SLIs) for key platform services and advise leadership on reliability risk and tradeoffs

  • Provide technical mentorship and guidance to other Platform Engineers

Recommended Experience & Education

Minimum Years of Experience

7 to 10 years of experience in software engineering, platform engineering, or a hybrid development/reliability engineering role

Education Level

Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent work experience

Required

  • Deep proficiency in TypeScript, with significant production experience in a Node.js/TypeScript service environment

  • Strong hands-on experience deploying, operating, and troubleshooting containerized workloads in Kubernetes

  • Proven experience designing, building, and maintaining CI/CD pipelines, GitHub Actions preferred

  • Proven experience designing monitoring, dashboard, and alerting strategy in an APM/observability tool, Datadog preferred

  • Deep experience troubleshooting distributed system and microservice communication failures across a variety of integration points and protocols

  • Strong experience with distributed tracing tools and practices at scale

  • Working familiarity with relational databases, sufficient to diagnose and resolve complex query and schema-level performance issues

  • Proven track record resolving significant application performance issues (caching, inefficient code paths, slow queries)

  • Experience designing and leading failure-mode testing (fault injection, resilience/chaos-style testing) and driving adoption of defensive patterns such as idempotency, retries, and circuit breakers

  • Demonstrated ability to work independently on ambiguous, high-impact reliability problems and to set technical direction for others

  • Excellent communication skills, with the ability to present technical root cause, risk, and remediation plans to engineering and non-technical leadership

  • Experience mentoring other engineers

Preferred

  • Experience with message brokers or event-streaming platforms such as Kafka, particularly at scale

  • Experience with OpenTelemetry or comparable distributed tracing frameworks at scale

  • Experience in a regulated or compliance-driven environment (fintech, banking, or similar)

  • Familiarity with infrastructure-as-code tooling (Terraform or similar)

  • Experience partnering with infrastructure/SRE teams on issues that cross application and infrastructure boundaries

  • Experience influencing engineering roadmap prioritization to secure time for platform stability work

The salary range for this position is: $140,000 - $175,000

Cool Things to Know

Not Just Any Company : Alkami has an awesome diverse and inclusive environment. We have a FUN culture and offer great benefits, including remote-first environment, unlimited paid time off, 401(k) with employer match, and more.

Work Authorization : We cannot offer employment sponsorship at this time. Candidates must be eligible to work in the US for full-time employment.

Recruiters : We are not looking for outside recruiting firms to help us in this search. Thank you for understanding.

Pay Transparency: As of January 1, 2023, new states and locales have enacted pay equity laws that require more pay transparency by employers in the following states: California, Colorado (effective January 1, 2021), Connecticut, Maryland, Nevada, New Jersey, New York, Ohio, Rhode Island and Washington.

The Important Stuff

Alkami Technology is an Equal Opportunity Employer and Prohibits Discrimination and Harassment of Any Kind: Alkami is committed to the principle of equal employment opportunity for all employees and to providing employees with a work environment free of discrimination and harassment. All employment decisions at Alkami are based on business needs, job requirements and individual qualifications, without regard to race, color, religion or belief, national, social or ethnic origin, sex (including pregnancy), age, physical, mental or sensory disability, HIV Status, sexual orientation, gender identity and/or expression, marital, civil union or domestic partnership status, past or present military service, family medical history or genetic information, family or parental status, or any other status protected by the laws and regulations in the locations where we operate. Alkami will not tolerate discrimination or harassment based on any of these characteristics. Alkami encourages applicants of all ages.

#LI-REMOTE

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer (MANTL)
Staff Software Engineer (MANTL)

Alkami Technology, Inc. • United States

Remote
USD 144,000 - 180,000
Remote-first environment
Unlimited paid time off
401(k) with employer match
Sr. Software Engineer, Data Platform (MANTL)
Sr. Software Engineer, Data Platform (MANTL)

Alkami • United States

Remote
USD 132,000 - 165,000
Remote-first environment
Unlimited PTO
401(k) with employer match
Staff Software Engineer (MANTL)
Staff Software Engineer (MANTL)

Alkami Technology • United States

Remote
USD 144,000 - 180,000
Remote-first culture
Unlimited PTO
401(k) with employer match
Sr. Software Engineer, Data Platform (MANTL)
Sr. Software Engineer, Data Platform (MANTL)

Alkami Technology, Inc. • United States

Remote
USD 132,000 - 165,000
Remote-first environment
Unlimited PTO
401(k) with employer match
Software Engineer (MANTL)
Software Engineer (MANTL)

Alkami • United States

Remote
USD 125,000 - 140,000
remote-first environment
unlimited paid time off
401(k) with employer match
Sr. Software Engineer (MANTL)
Sr. Software Engineer (MANTL)

Alkami Technology • United States

On-site
USD 132,000 - 165,000
Unlimited paid time off
401(k) with employer match
Remote-first environment
Sr. Solution Architect, Integrations (MANTL)
Sr. Solution Architect, Integrations (MANTL)

Alkami • United States

Remote
USD 96,000 - 120,000
remote-first environment
unlimited paid time off
401(k) with employer match
Software Engineer (MANTL)
Software Engineer (MANTL)

Alkami Technology, Inc. • Northern (KY)

Hybrid
USD 125,000 - 140,000
Remote-friendly culture
Unlimited PTO
401(k) with employer match
+1
Sr. Software Engineer
Sr. Software Engineer

Alkami • Northern (KY)

On-site
USD 120,000 - 140,000
Remote-first environment
Unlimited paid time off
401(k) with employer match
Sr. Software Engineer (MANTL)
Sr. Software Engineer (MANTL)

Alkami Technology • Northern (KY)

On-site
USD 130,000 - 165,000
Remote-first environment
Unlimited paid time off
401(k) with employer match