Senior Site Reliability Engineer

Navan

Dallas (TX)

On-site

USD 110,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Navan is seeking a passionate Site Reliability Engineer in Dallas, TX or Austin, TX. The candidate will design tooling, automation, and infrastructure services for thousands of travelers. Responsibilities include building reliable cloud infrastructure, identifying reliability issues, and automating processes. Ideal candidates should have 5+ years of SRE or DevOps experience, excellent communication skills, and hands-on expertise with Java and AWS. The role offers an exciting opportunity to work in a fast-paced environment with a focus on automation and innovation.

Qualifications

  • 5+ years of experience as a Senior SRE or DevOps Lead.
  • 2+ years in a production, 24x7 environment.
  • Excellent communication skills with stakeholders.

Responsibilities

  • Designing and operating cloud infrastructure.
  • Identifying and solving reliability issues systemically.
  • Automating operational processes to improve efficiency.

Skills

Cloud infrastructure design
Automation
Java applications
AWS
Infrastructure as Code
Distributed systems

Tools

Terraform
Jenkins
DataDog
NewRelic

Job description

At Navan, “It’s all about the user. All of them.” We’re passionate about providing a seamless one-stop experience for business travelers, no matter how they travel, where they stay, or where they’re going. We are constantly striving to make the most reliable and scalable systems possible to ensure that our services are available to our travelers when they need it most. With our exponential growth, we have many exciting challenges ahead and we’re looking for a passionate Site Reliability Engineer to join our team in Dallas, TX or Austin, TX. As an SRE you will design and develop tooling, automation and infrastructure services that power the Navan services, used by thousands of travelers on a daily basis. You will work closely with development teams, release and productivity teams and security teams to identify customer needs and build innovative solutions to solve them. You will work across a vast array of systems and technologies, aiming to build an autonomous, monitored, fault‑tolerant infrastructure that is optimized for both simplicity and uptime. You will collaborate with the backend and frontend engineering teams to ensure that product solutions are scalable, efficient, and reliable. You will design infrastructure to support our massive growth and work with the team to maintain the highest level of service.

What You’ll Do
  • Building a fast moving, high growth service. Navan is revolutionizing travel and expense services for the enterprise, and the product is evolving quickly. You are comfortable in a startup environment, enjoy seeing the product take shape, and have strong ownership of the success of your services.
  • Designing, implementing and operating cloud infrastructure. You’re a fit for us if you think in terms of infrastructure as code, deployment pipelines, and building the guardrails to make going fast also going safely.
  • Identifying reliability anti‑patterns and solving them systemically. You dive deep into the data to evaluate the health of your systems, and you use it to improve visibility and reliability across the fleet of services.
  • Finding and automating the toil out of our processes. You’d prefer to automate it entirely, or build a tool to empower your users rather than be the gatekeeper to the tool.
  • Leveraging AI tools and platforms in your daily work to achieve autonomous operations, reduce toil, and improve system observability.
  • Defining and driving the adoption of system reliability standards, including formalizing SLO/SLI frameworks, observability standards, and blameless post‑mortem practices across multiple engineering teams.
  • Driving the adoption of AI‑assisted developer tools and platforms to increase engineering productivity, enforce code quality standards, and enable real‑time architectural validation.
What We’re Looking For
  • 5+ years of progressive experience as a Senior SRE or DevOps Lead (or equivalent role)
  • 2+ years of experience in working on a production, 24x7 product environment
  • Passionate about solving problems and learning new tools and technologies
  • Excellent communication skills working with stakeholders and domain experts across the company to design solutions to user problems
  • Thrive in a fast‑paced environment
  • Demonstrated experience mentoring and leading junior and mid‑level engineers, and acting as a technical owner for cross‑functional infrastructure projects.
  • Operate with a strong sense of ownership demonstrated through shipping production‑quality code and infrastructure equipped with testing, monitoring and documentation
  • Hands‑on operational experience with Java based applications and services including JVM profiling and performance tuning (python, Node.js and Go are a plus)
  • Hands‑on experience building and operating distributed systems in a public cloud environment (preferably AWS), using CI/CD to deploy, manage and operate production systems, focusing on tooling and automation using tools such as maven and Jenkins.
  • Hands‑on experience with microservice architecture and related reliability and resiliency patterns such as throttling, queueing, and retries
  • Hands‑on experience with writing Infrastructure as Code in Terraform or Cloudformation or similar tools
  • A passion for automating away everything, using scripting languages such as python, bash groovy (we prefer lazy engineers)
  • Built, using, and automating monitoring systems such as NewRelic, DataDog, SignalFX, Kibana
  • Hands‑on experience deploying, operating, and monitoring production‑grade AI/ML microservices (e.g., RAG pipelines, agentic systems) on cloud platforms like AWS Fargate/ECS.
  • Experience leveraging AI/LLM platforms (e.g., Gemini, Braintrust) and managing their secrets and infrastructure using Infrastructure as Code (Terraform) and AWS SSM.
  • Demonstrated ability to integrate AI‑specific telemetry and advanced observability practices to enable predictive insights and systemic root‑cause analysis.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: AI‑Driven, Autonomous Cloud Uptime
Senior SRE: AI‑Driven, Autonomous Cloud Uptime

Navan • Dallas (TX)

On-site
USD 110,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Cvent, Inc. • Tysons (VA)

Hybrid
USD 100,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Namely • United States

Hybrid
USD 120,000 - 150,000
Senior Manager, Strategy & Innovation
Senior Manager, Strategy & Innovation

Traveltechessentialist • Dallas (TX)

On-site
USD 150,000 - 210,000
Healthcare coverage
Equity plans
Flexible time off
Senior Software Engineer, Flights
Senior Software Engineer, Flights

Traveltechessentialist • New York (NY)

On-site
USD 113,000 - 252,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Cvent • Tysons (VA)

Hybrid
USD 110,000 - 140,000
Sr. Manager, Analytics
Sr. Manager, Analytics

Traveltechessentialist • Palo Alto (CA)

On-site
USD 122,000 - 272,000
Fullstack Software Engineer II, Flights
Fullstack Software Engineer II, Flights

Navan • New York (NY)

On-site
USD 92,000 - 206,000
Healthcare
Insurance
Wellness resources
+1
Associate Engineer, Site Reliability
Associate Engineer, Site Reliability

Calabrio • United States

On-site
USD 90,000 - 130,000
Senior Director of Event Travel, Navan Events
Senior Director of Event Travel, Navan Events

Navan • San Francisco (CA)

On-site
USD 180,000 - 240,000