Software Engineer, Infrastructure

Gray Swan AI

United States

On-site

USD 180,000 - 290,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

401k with up to 4% matching
28 days annual leave
Health, dental, and vision coverage
Catered lunches (Pittsburgh office)
Flexible work arrangements
Visa sponsorship available for high-pq

Job summary

Gray Swan AI is seeking an Infrastructure Engineer to design, build, and scale the backend systems powering our AI security platform. You’ll own cloud infrastructure across Kubernetes and cloud providers, build scalable APIs, and work with ML, security, and product teams to deliver reliable AI workloads.

This role emphasizes ownership, scalable architectures, and fast-moving startup dynamics. You’ll collaborate with engineers to keep systems fast, resilient, and secure as Gray Swan grows in the

Qualifications

  • 5+ years of experience building backend infrastructure or distributed systems in production environments.
  • Strong programming skills in C/C++, Go, Python, Rust, or Java.
  • Experience operating services on Kubernetes and cloud platforms such as AWS, GCP, or Azure.
  • Deep understanding of networking, distributed systems, containers, service orchestration, and scalable architectures.
  • Experience designing APIs, microservices, asynchronous systems, and event-driven architectures.
  • Comfortable debugging complex production issues and improving reliability through automation and operational excellence.

Responsibilities

  • Design, build, and maintain highly available backend services and distributed systems that power Gray Swan's AI security platform.
  • Own cloud infrastructure across Kubernetes, AWS, networking, storage, and compute to ensure reliable production environments.
  • Build scalable APIs, internal platform services, and infrastructure tooling that improve developer productivity and system reliability.
  • Improve system observability through logging, metrics, tracing, dashboards, and automated alerting.
  • Optimize performance, latency, and infrastructure costs while maintaining reliability and security.
  • Partner closely with machine learning, security, and product engineering teams to deliver production-ready infrastructure for AI workloads.

Skills

Backend infra
Kubernetes
Cloud platforms
APIs & microservices
Networking
Programming languages
Observability

Tools

Kafka
Redis
PostgreSQL
ClickHouse

Job description

About Gray Swan

Gray Swan is on a mission to empower the world to use AI safely and securely. We evaluate AI models for the leading frontier labs along with building real-time threat detection and adaptive adversarial red teaming agents for teams deploying AI.

We're a team of approximately 50 people, well-funded, growing quickly. Our work directly influences how the world deploys AI agents and systems at scale..

Learn more about how we work.

The Role

Gray Swan is looking for an Infrastructure Engineer to build and scale the systems that power our AI security platform. You'll design the backend services, distributed infrastructure, and cloud architecture that enable Gray Swan to build AI systems and for customers to safely deploy frontier AI models at scale.

This role is ideal for an engineer who enjoys solving infrastructure challenges across reliability, scalability, observability, and performance. You'll work closely with machine learning engineers, product engineers, and security researchers to ensure our platform remains fast, resilient, and secure as we grow.

You'll have significant ownership over foundational systems and the opportunity to influence technical direction in a rapidly evolving AI startup.

What You’ll Do:
  • Design, build, and maintain highly available backend services and distributed systems that power Gray Swan's AI security platform.
  • Own cloud infrastructure across Kubernetes, AWS, networking, storage, and compute to ensure reliable production environments.
  • Build scalable APIs, internal platform services, and infrastructure tooling that improve developer productivity and system reliability.
  • Improve system observability through logging, metrics, tracing, dashboards, and automated alerting.
  • Optimize performance, latency, and infrastructure costs while maintaining reliability and security.
  • Partner closely with machine learning, security, and product engineering teams to deliver production-ready infrastructure for AI workloads.
Who You Are:
  • 5+ years of experience building backend infrastructure or distributed systems in production environments.
  • Strong programming skills in C/C++, Go, Python, Rust, or Java.
  • Experience operating services on Kubernetes and modern cloud platforms such as AWS, GCP, or Azure.
  • Deep understanding of networking, distributed systems, containers, service orchestration, and scalable architectures.
  • Experience designing APIs, microservices, asynchronous systems, and event-driven architectures.
  • Comfortable debugging complex production issues and improving reliability through automation and operational excellence.
  • Passionate about writing clean, maintainable code and building infrastructure that other engineers love using.
  • Excited to work in a fast-moving startup with significant ownership and ambiguity.
Bonus Points If You Have:
  • Experience supporting machine learning or LLM infrastructure.
  • Familiarity with infrastructure-as-code tools.
  • Experience with Kafka, Redis, PostgreSQL, ClickHouse, or similar distributed data systems.
  • Experience building internal developer platforms or platform engineering tooling.
  • Knowledge of cloud security, infrastructure hardening, or zero-trust architectures.
  • Previous experience at a high-growth startup or building products from zero to one.
  • Interest in AI safety, cybersecurity, or adversarial machine learning.
You’ll Thrive Here If You:
  • You thrive on ownership and solving hard problems. You're energized by ambiguity, enjoy building systems from the ground up, and take pride in delivering reliable solutions from design through production.
  • You think at scale. You enjoy designing resilient infrastructure, optimizing performance, and building systems that are secure, observable, and built to grow.
  • You collaborate across disciplines. You work effectively with machine learning engineers, security researchers, and product teams, knowing that the best infrastructure enables everyone else to move faster.
  • You're excited by our mission and startup environment. You enjoy moving quickly, adapting to change, and helping build the foundation for the future of secure AI.
What We Offer:

We offer a competitive compensation package designed to reward impact and incentivize growth. Our compensation philosophy is informed by our current valuation and recent industry data.

Compensation: $180,000 - $290,000 (depending on level)plus performance based bonus and meaningful equity package.

Benefits:

  • 401k with up to 4% matching
  • 28 days annual leave (vacation + holidays)
  • Health, dental, and vision coverage
  • Catered lunches (Pittsburgh office)
  • Flexible work arrangements
  • Visa sponsorship available for exceptional candidates
Interview Process

Application review. We read everything; we’ll respond within 10 days.

Online technical screen (15 min). Complete a simple, job-relevant exercise.

Intro call (30 min). We learn about you; you learn about us.

Technical interview (90 min). Live coding with some tasks requiring AIand others not.

Experience & culture interview (60 min). Conversational exploration of the skills fit.

Reference checks. We’ll reach out to 3-5 references that you provide.

Offer. If it’s mutual, we move fast.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Infrastructure
Software Engineer, Infrastructure

Gray Swan • Pittsburgh

Hybrid
USD 180,000 - 290,000
401k with up to 4% matching
28 days annual leave
Health, dental, and vision coverage
+3
Software Engineer, Infrastructure
Software Engineer, Infrastructure

Gray Swan • United States

On-site
USD 180,000 - 290,000
401k with up to 4% matching
28 days annual leave
Health, dental, and vision coverage
+3
Software Engineer, Infrastructure
Software Engineer, Infrastructure

Madrona Venture Labs • Pittsburgh

On-site
USD 180,000 - 290,000
401k with matching
28 days annual leave
Health, dental, and vision
+3
Cyber Security Engineer
Cyber Security Engineer

Gray Swan AI • Pittsburgh

On-site
USD 95,000 - 145,000
401k matching
Vacation + holidays
Health, dental, and vision
+3
Senior Software Engineer (Pittsburgh)
Senior Software Engineer (Pittsburgh)

Madrona Venture Labs • Pittsburgh

On-site
USD 180,000 - 220,000
401k matching
Paid vacation
Health insurance
+3
Staff Software Engineer (Pittsburgh)
Staff Software Engineer (Pittsburgh)

Madrona Venture Labs • Pittsburgh

On-site
USD 210,000 - 280,000
401k with up to 4% matching
28 days annual leave
Health, dental, and vision coverage
+3
Senior Software Engineer (Pittsburgh)
Senior Software Engineer (Pittsburgh)

Gray Swan • Pittsburgh

On-site
USD 180,000 - 260,000
401k with matching
28 days annual leave
Health, dental, and vision coverage
+3
Software Engineer
Software Engineer

Gray Swan • Pittsburgh

On-site
USD 115,000 - 165,000
28 days off between holidays & PTO
401(k) with up to 4% matching
Flexible work arrangements
+4
Machine Learning Engineer
Machine Learning Engineer

Gray Swan AI • Pittsburgh

On-site
USD 140,000 - 225,000
401k with up to 4% matching
28 days annual leave (vacation +bahold
Health, dental, and vision coverage
+3
Senior Forward Deployed Engineer – Enterprise
Senior Forward Deployed Engineer – Enterprise

Gray Swan • Seattle (WA), San Francisco (CA)

Hybrid
USD 175,000 - 220,000
401k match
Paid time off
Health insurance
+3