Sequen - Staff Software Engineer – Infrastructure

fabric

New York (NY)

On-site

USD 250,000 - 350,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health insurance
Unlimited time off
Equity

Job summary

Sequen AI is seeking an experienced Infrastructure Engineer to join our team to support the development, scaling, and maintenance of cutting-edge AI systems. You will lead the systems group, ensuring reliable training and serving of frontier ranking models.

Responsibilities include setting technical strategy, overseeing high-scale infrastructure, and building proactive monitoring with Prometheus and Datadog. Strong leadership, cloud expertise, and distributed systems passion are essential.

Qualifications

  • 10+ years of relevant industry experience.
  • Experience leading large-scale infrastructure projects or teams.
  • Strong expertise in cloud infra and distributed systems.

Responsibilities

  • Consult with stakeholders to understand infra and compute needs.
  • Set technical strategy and oversee development of high-scale infrastructure.
  • Design processes to improve reliability and prevent repeats.

Skills

Distributed systems
Leadership
Programming: Python/Go/Java
Problem solving
Independent worker

Tools

Kubernetes
Prometheus
Datadog
Terraform
AWS
GCP

Job description

About Us

Sequen AI is leading the charge for building frontier ranking models for search and recommendations. Sequen AI's technology specializes in designing end-user behavior for large consumer enterprises.

About the Role

We are currently looking for talented and experienced Infrastructure Engineers to join our team and support the development, scaling, and maintenance of our cutting-edge AI systems.

  • Core Infrastructure: The systems team is responsible for supporting clusters used to train, research, and ultimately serve AI models. Your work will be crucial in ensuring Sequen is able to continue to reliably train and serve frontier ranking models.
  • Observability: We build and maintain the infrastructure that monitors the health, performance, and efficiency of our AI systems. You'll work across teams to implement monitoring solutions using tools like Prometheus,, and Datadog, while developing automated approaches for dashboards and alerts. Your work will create reliable, low-maintenance systems that enable proactive monitoring and operational excellence.
Responsibilities
  • Consult with different stakeholders to deeply understand infrastructure, data and compute needs, identifying potential solutions to support frontier research and product development
  • Set technical strategy and oversee development of high scale, reliable infrastructure systems.
  • Design processes (e.g. postmortem review, incident response, on-call rotations) that help the team operate effectively and never fail the same way twice
About You
  • Have 10+ years of relevant industry experience, 3+ years leading large scale, complex projects or teams as an engineer or tech lead
  • Possess deep knowledge of modern cloud infrastructure including Kubernetes, Infrastructure as Code, AWS, and GCP
  • Are obsessed with distributed systems at scale, infrastructure reliability, scalability, security, and continuous improvement
  • Strong proficiency in at least one programming language (e.g., Python, Go, Java)
  • Strong problem-solving skills and ability to work independently
  • Have a passion for supporting internal partners like research to understand their needs
  • Have excellent communication skills to build consensus with stakeholders, both internally and externally
Strong Candidates May Have
  • Security and privacy best practice expertise
  • Hands-on experience with data pipelines and processing large-scale datasets
  • Experience with machine learning infrastructure like GPUs,
  • Technical expertise: Quickly understanding systems design tradeoffs, keeping track of rapidly evolving software systems
What We Offer
  • An senior role with massive impact on infrastructure and engineering velocity
  • Ownership of mission-critical systems that power ML and data workflows
  • Annual Salary : $250,000—$350,000 USD + Equity
  • Health insurance, unlimited time off, awesome team
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff, MLOps Engineer
Staff, MLOps Engineer

Sequen AI • United States

Remote
USD 220,000 - 280,000
Senior Infrastructure Engineer for ML Systems & Scale
Senior Infrastructure Engineer for ML Systems & Scale

fabric • New York (NY)

On-site
USD 250,000 - 350,000
Health insurance
Unlimited time off
Equity
Software Engineer - Infrastructure
Software Engineer - Infrastructure

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Software Engineer - Infrastructure
Software Engineer - Infrastructure

Baseten • San Francisco (CA)

On-site
USD 90,000 - 130,000
100% coverage of medical, dental, and vision insurance
Generous PTO policy including Winter Break
Company-facilitated 401(k)
Senior Software Engineer - Infrastructure
Senior Software Engineer - Infrastructure

AfterQuery • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 280,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Strativ Group • Menlo Park (CA)

On-site
USD 250,000 - 320,000
Software Engineer - Infrastructure
Software Engineer - Infrastructure

The Consensus • New York (NY)

On-site
USD 120,000 - 150,000
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
Paid parental leave
+2
Member of Technical Staff - Sandbox Platform
Member of Technical Staff - Sandbox Platform

Prime Intellect • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Visa sponsorship
Relocation support
Professional development budget
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • Arlington (VA)

On-site
USD 175,000 - 250,000