Senior AI Infrastructure Engineer, Inference at Scale

Bigbear.ai

Columbia (MD)

On-site

USD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

BigBear.ai is seeking a Software Engineer to support our AI infrastructure team, focusing on inference services and scalable AI platform components. You will help design, implement, and operate critical infrastructure and production systems across the AI stack.

8+ years of experience or a technical degree with 4+ years, plus TS/SCI w/Poly clearance, are valued. You will collaborate with multiple teams, improve observability, and guide junior engineers in a fast-evolving environment.

Qualifications

  • 8+ years of relevant experience or a technical Bachelor's with 4+ years experience.
  • Clearance: TS/SCI w/Poly
  • Proven experience building and maintaining production systems at scale.
  • Experience with high-volume web application architecture and performance optimization.
  • Strong background in systems integration across diverse technologies and platforms.
  • Hands-on experience with cloud engineering in AWS.
  • Proficiency with Kubernetes administration and deployment patterns.
  • Strong Python programming skills.
  • Experience implementing observability solutions (APM, OpenTelemetry, Grafana, Prometheus).
  • Familiarity with CI/CD pipelines and DevOps practices.
  • Strong change management and organizational influence skills.
  • Ability to thrive in ambiguous environments and create structure where needed.
  • Excellent communication and collaboration skills.

Responsibilities

  • Design, implement, and optimize infrastructure for AI model inference at scale.
  • Support the development and maintenance of production AI services and applications, including retrieval augmented generation (RAG), autonomous agents, and emerging technologies.
  • Navigate ambiguity and define solutions for underspecified systems and requirements.
  • Drive adoption of new technologies and practices across engineering teams.
  • Implement monitoring, logging, and observability solutions for AI services.
  • Automate infrastructure provisioning and configuration using Infrastructure-as-Code (IaC) principles.
  • Ensure high availability, reliability, and performance of AI platform components.
  • Contribute to security best practices for AI systems and data.
  • Provide technical guidance and informal mentorship to junior engineers.

Skills

Python
Kubernetes
AWS
Observability
CI/CD
IaC
Security best practices
Systems integration
Production systems
Mentorship

Education

Bachelor's degree in a technical discipline

Tools

LangChain
OpenTelemetry
Grafana
Prometheus
vLLM/LiteLLM

Job description

BigBear.ai is seeking a Software Engineer to support our AI infrastructure team, focusing on inference services and scalable AI platform components. You will help design, implement, and operate critical infrastructure and production systems across the AI stack.

8+ years of experience or a technical degree with 4+ years, plus TS/SCI w/Poly clearance, are valued. You will collaborate with multiple teams, improve observability, and guide junior engineers in a fast-evolving environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer (AI Infrastructure)
Software Engineer (AI Infrastructure)

BigBear.ai • Columbia (MD)

On-site
USD 150,000 - 190,000
Senior Software Engineer - ML Analytics (TS/SCI)
Senior Software Engineer - ML Analytics (TS/SCI)

Bigbear.ai • Columbia (MD)

On-site
USD 140,000 - 200,000
Senior Software Engineer (ML)
Senior Software Engineer (ML)

BigBear.ai • Columbia (MD)

On-site
USD 140,000 - 200,000
Senior AI Inference Infrastructure Engineer
Senior AI Inference Infrastructure Engineer

Modular • United States

Hybrid
USD 167,000 - 273,000
Senior AI Engineer
Senior AI Engineer

BigBear.ai • Maryland

On-site
USD 150,000 - 210,000
Full Stack Software Engineer (AI Infrastructure)
Full Stack Software Engineer (AI Infrastructure)

Bytoa • Laurel (MD)

On-site
USD 200,000 - 220,000
AI Infra Engineer (TS/SCI) - Production & Observability
AI Infra Engineer (TS/SCI) - Production & Observability

Visionist, Inc. • Laurel (MD)

On-site
USD 115,000 - 160,000
Employee ownership
Competitive retirement plan
Paid time off + holidays
+4
Senior Systems Engineer, AI Inference Infra
Senior Systems Engineer, AI Inference Infra

United States Digital Space LLC • United States

Remote
USD 150,000 - 190,000
Principal Software Engineer
Principal Software Engineer

BigBear.ai • Columbia (MD)

On-site
USD 180,000 - 240,000
Senior ML Systems Engineer — Inference & Scale
Senior ML Systems Engineer — Inference & Scale

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000