Software Engineer II - AI Infrastructure with Security Clearance

Black Eagle Defense

Fort Meade (MD)

On-site

USD 143,000 - 200,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Black Eagle Defense is seeking a Software Engineer II - AI Infrastructure in Fort Meade, MD to help build scalable AI infrastructure powering enterprise AI solutions. You will design, implement, and maintain platform capabilities for AI inference and related services, collaborating with cross-functional teams to ensure reliability and security at scale.

The role requires eight years of SWE experience, a CS degree, and strong Python, AWS, Kubernetes skills.

Qualifications

  • Eight years of SWE experience in programs/contracts of similar scope and complexity.
  • Bachelor's degree in CS or related discipline required.
  • Experience building, deploying, and maintaining production systems at scale.
  • Experience designing high-volume web app architectures for performance, scalability, and reliability.
  • Strong cloud engineering and Kubernetes deployment experience.
  • Proficiency in Python for automation and development.
  • Experience with observability and monitoring (APM, OpenTelemetry, Grafana, Prometheus).
  • Familiarity with CI/CD pipelines and DevOps practices.
  • Ability to operate in ambiguous environments and drive structure.

Responsibilities

  • Design, implement, and maintain scalable AI inference infrastructure.
  • Deploy AI services including RAG and autonomous agents.
  • Ensure reliability, security, and high performance of platform components.
  • Automate provisioning and configuration with IaC.
  • Collaborate with engineers and stakeholders to improve AI platform delivery.
  • Provide guidance and mentoring to junior engineers.

Skills

AI infra design
Python development
AWS
Kubernetes
Observability tooling
IaC
CI/CD pipelines
Security best practices
Problem solving
Communication

Education

Bachelor's degree in Computer Science or related

Tools

Kubernetes
Python
OpenTelemetry
Grafana
Prometheus
LangChain
vLLM/LiteLLM
RAG platforms
Terraform
CI/CD tooling

Job description

Job Description SALARY RANGE $143,000 - $200,000/year DUTIES As a successful candidate for the Software Engineer II - AI Infrastructure role, you will support the development, operation, and evolution of the next generation of AI infrastructure that enables innovation across the customer organization. As part of a full-stack engineering team, you will design, implement, and maintain scalable platform capabilities that serve as the foundation for AI-powered applications and services. Your efforts will focus on AI inference infrastructure while supporting a broader ecosystem that includes advanced analytics, retrieval-augmented generation (RAG), autonomous agents, and emerging AI technologies. In this role, you will independently design, develop, deploy, and optimize infrastructure components that deliver reliable, secure, and high-performance AI capabilities at scale. You will collaborate with engineers, platform teams, and stakeholders to enhance platform reliability, drive adoption of modern technologies and engineering practices, and ensure AI services remain scalable, observable, and operationally resilient. Through cloud engineering, automation, systems integration, and platform development, you will help deliver the infrastructure that powers mission-critical AI solutions across the enterprise.

Required Skills

SKILLS * Design, implement, and optimize infrastructure supporting AI model inference at scale

  • Develop, deploy, and maintain production AI services and applications, including retrieval-augmented generation (RAG), autonomous agents, and emerging AI technologies
  • Analyze ambiguous requirements and define scalable, maintainable solutions for complex systems and operational challenges
  • Drive the adoption of modern technologies, engineering standards, and best practices across development teams
  • Implement monitoring, logging, and observability capabilities to improve visibility into AI platform performance and reliability
  • Automate infrastructure provisioning, deployment, and configuration management using Infrastructure-as-Code principles
  • Ensure the availability, reliability, scalability, and performance of AI platform components and supporting services
  • Contribute to the implementation of security best practices for AI systems, services, and data environments
  • Design and integrate platform capabilities that support enterprise AI initiatives and operational requirements
  • Collaborate with engineers, platform teams, and stakeholders to improve AI infrastructure and service delivery
  • Troubleshoot complex infrastructure, platform, and application issues within production environments
  • Provide technical guidance, knowledge sharing, and informal mentorship to junior engineers
  • Support the continuous improvement and modernization of AI infrastructure, cloud environments, and platform operations
  • Contribute to the full lifecycle of AI platform development, from design and implementation through deployment and sustainment QUALIFICATIONS Eight (8) years of experience as a SWE in programs and contracts of similar scope, type, and complexity are required. A Bachelor's degree in Computer Science or a related discipline from an accredited college or university is required. Four (4) years of additional SWE experience on projects with similar software processes may be substituted for a bachelor's degree. Additional requirements: * Proven experience building, deploying, and maintaining production systems at scale
  • Experience designing and optimizing high-volume web application architectures for performance, scalability, and reliability
  • Strong background in systems integration across diverse technologies, platforms, and services
  • Hands-on experience with cloud engineering and solution deployment within AWS environments
  • Proficiency in administering and deploying applications within Kubernetes-based environments
  • Strong Python development skills for automation, infrastructure, and application development efforts
  • Experience implementing observability and monitoring solutions using technologies such as APM, OpenTelemetry, Grafana, and Prometheus
  • Familiarity with CI/CD pipelines, automation frameworks, and DevOps best practices
  • Strong understanding of infrastructure automation, deployment strategies, and operational excellence principles
  • Strong change management, stakeholder engagement, and organizational influence skills
  • Ability to operate effectively within ambiguous environments and establish structure for evolving requirements
  • Strong analytical, troubleshooting, and problem-solving skills
  • Excellent written and verbal communication skills
  • Ability to collaborate effectively across multidisciplinary engineering and operational teams
  • Experience supporting the full lifecycle of cloud-native applications and platform services from design through production operations Desired Skills NICE-TO-HAVES * Experience with AI inference serving technologies such as vLLM, LiteLLM, or similar platforms
  • Experience developing solutions using agentic AI frameworks such as LangChain or comparable technologies
  • Knowledge of vector databases, embedding models, and semantic search architectures
  • Experience designing and supporting retrieval-augmented generation (RAG) solutions and AI-enabled applications
  • Familiarity with large language model deployment, optimization, and inference workflows
  • Experience with high-performance computing environments and distributed systems architectures
  • Knowledge of scalable data processing, distributed computing, and resource optimization techniques
  • Experience supporting enterprise AI platforms and machine learning infrastructure
  • Familiarity with emerging AI technologies, frameworks, and platform capabilities
  • Experience integrating AI services and infrastructure into cloud-native environments and production systems
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer III – AI Infrastructure with Security Clearance
Senior Software Engineer III – AI Infrastructure with Security Clearance

Black Eagle Defense • Fort Meade (MD)

On-site
USD 209,000 - 266,000
Software Engineer II - AI Infrastructure
Software Engineer II - AI Infrastructure

Black Eagle Defense • Fort Meade (MD)

On-site
USD 143,000 - 200,000
Software Engineer I - Inference with Security Clearance
Software Engineer I - Inference with Security Clearance

Black Eagle Defense • Fort Meade (MD)

On-site
USD 124,000 - 181,000
Full Stack Software Engineer (AI Infrastructure)
Full Stack Software Engineer (AI Infrastructure)

Bytoa • Laurel (MD)

On-site
USD 200,000 - 220,000
Software Engineer I with Security Clearance
Software Engineer I with Security Clearance

Black Eagle Defense • Fort Meade (MD)

On-site
USD 124,000 - 181,000
Senior Software Engineer III – AI Infrastructure
Senior Software Engineer III – AI Infrastructure

Black Eagle Defense • Fort Meade (MD)

On-site
USD 209,000 - 266,000
Software Engineer II – AI Infrastructure
Software Engineer II – AI Infrastructure

Staffed4U • Maryland

On-site
USD 193,000 - 306,000
Health coverage
401(k) retirement plan
Paid time off
Sr. Systems Engineer - AI
Sr. Systems Engineer - AI

Dairy-Farmers-of-America,-Inc. • Kansas City (KS)

On-site
USD 98,000 - 139,000
Software Engineer (AI Infrastructure)
Software Engineer (AI Infrastructure)

BigBear.ai • Columbia (MD)

On-site
USD 150,000 - 190,000
AI Platform Engineer II
AI Platform Engineer II

Mindtris • Memphis (TN)

On-site
USD 110,000 - 160,000