Senior Software Engineer (AI Systems & Infrastructure)

Bytoa

Columbia (MD)

On-site

USD 235,000 - 255,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bytoa is seeking a Senior Software Engineer to join a full-stack LLM integration and delivery team in Columbia, MD. You will architect and develop scalable AI-powered applications and the infrastructure that supports them, ensuring enterprise-scale reliability and rapid iteration.

The role emphasizes ownership, learning, and building robust systems that enhance user workflows. 12+ years of experience with strong cloud, containerization, and observability skills are expected.

Qualifications

  • 12+ years of software engineering experience focused on scalable systems.
  • Strong full-stack development experience for user-facing applications.
  • Proficient in Python, Go, or Java.
  • Extensive cloud platform experience (AWS, GCP, Azure).
  • Expertise with Docker and Kubernetes.
  • Familiarity with infrastructure-as-code tools.
  • Experience with monitoring and observability tools.
  • Knowledge of CI/CD practices.
  • Excellent problem-solving and communication skills.
  • Experience delivering AI-powered features with strong UX focus.

Responsibilities

  • Lead design and development of scalable LLM-powered applications and services.
  • Architect infrastructure to enable rapid iteration and deployment of AI features.
  • Collaborate with product teams to translate user needs into solutions.
  • Build platforms to ship AI features reliably.
  • Develop automation tools to improve reliability and efficiency.
  • Implement monitoring, logging, and alerting systems.
  • Perform capacity planning and performance tuning for AI workloads.
  • Lead incident response and post-mortems.
  • Mentor junior engineers and contribute to team growth.
  • Continuously improve systems and development processes.

Skills

12+ years of software engineering
Full-stack development
Python/Go/Java
Cloud platforms (AWS/GCP/Azure)
Docker/Kubernetes
Terraform/Ansible/Puppet
Prometheus/Grafana/ELK
CI/CD pipelines
Problem solving
Communication skills
User experience focus

Education

Bachelor's degree in a technical discipline

Tools

Docker
Kubernetes
Terraform
Ansible
Puppet
Prometheus
Grafana
ELK
CI/CD tooling

Job description

Senior Software Engineer (AI Systems & Infrastructure)
  • Columbia, MD
Description:

We are seeking a highly experienced and driven Sr. Software Engineer to join a full stack LLM integration and delivery team. Ideally, you'll have a strong background in building scalable AI-powered applications and the infrastructure that supports them, with a focus on delivering exceptional user experiences. You will play a crucial role in architecting and developing the systems that power cutting-edge LLM applications, ensuring they perform reliably at enterprise scale while enabling rapid iteration and deployment.

The ideal candidate will have the grit, self-reliance, and drive to identify and fix inefficiencies, consistently leaving our codebase and infrastructure in a better state than they found it. We're looking for someone who thrives in a fast-paced environment, gets excited about seeing their code improve real user workflows, and is passionate about building and maintaining robust, scalable systems that power cutting-edge LLM applications. You should be someone who actively questions requirements, wants to learn, and brings the industrious mindset that ensures our tools not only work but truly serve our users' needs.

Salary range:$235,000-255,000
Disclaimer:Salary for this position, along with additional compensation options, will be determined on an individual basis following the interview process, considering various factors such as years of experience, skills, education/certifications, contract specifications, market conditions, etc.

Responsibilities:
  • Lead the design and development of scalable LLM-powered applications and services.
  • Architect infrastructure solutions that support rapid iteration and deployment of AI features
  • Collaborate directly with product teams to translate user needs into technical solutions.
  • Build and maintain the platforms that enable your team to ship AI features quickly and reliably.
  • Develop and manage automation tools to improve system reliability and development efficiency.
  • Implement and maintain monitoring, alerting, and logging systems.
  • Conduct capacity planning and performance tuning for AI workloads.
  • Lead and participate in incident response and post-mortem analyses.
  • Mentor junior team members and contribute to the overall growth of the engineering team.
  • Continuously identify and implement improvements to our systems and development
Skills Requirements:
  • 12+ years of experience in software engineering with focus on scalable systems.
  • Strong full-stack development experience with user-facing applications.
  • Strong programming skills in languages such as Python, Go, or Java.
  • Extensive experience with cloud platforms (e.g., AWS, GCP, Azure) and their services.
  • Proficiency in containerization technologies (Docker, Kubernetes).
  • Experience with infrastructure-as-code tools (e.g., Terraform, Ansible, Puppet).
  • Expertise in monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack).
  • Familiarity with CI/CD pipelines and practices.
  • Strong problem-solving skills and ability to troubleshoot complex systems.
  • Excellent communication skills and ability to work in a collaborative environment.
  • Experience building products that prioritize user experience and product-market fit.
Nice to Haves:
  • Experience working with Large Language Models (LLMs) and related infrastructure.
  • Experience with AI/ML model serving and optimization.
  • Background in product-focused engineering environments.
  • Familiarity with machine learning operations (MLOps) practices.
  • Experience with A/B testing and feature flagging for AI features.
  • Contributions to open-source projects.
  • Experience with distributed systems and microservices architectures.
  • Knowledge of security best practices and compliance requirements.
  • Experience with real-time data processing and streaming platforms (e.g., Apache Kafka, Apache Flink).
  • Familiarity with chaos engineering principles and tools.
Experience Requirement:

12 yrs. with a B.S. in a technical discipline or 4 additional yrs. in place of B.S.

Clearance Requirement:

Active TS/SCI with a polygraph

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Software Engineer - AI Systems - FS Poly
Sr. Software Engineer - AI Systems - FS Poly

stanleyreid • Columbia (MD)

On-site
USD 140,000 - 190,000
Health benefits
PTO + holidays
401k match
+2
Senior LLM Systems Architect — AI for Security & Automation
Senior LLM Systems Architect — AI for Security & Automation

7AI • Boston (MA)

On-site
USD 140,000 - 200,000
Junior Software Engineer – Inference with Security Clearance
Junior Software Engineer – Inference with Security Clearance

Neural Solutions • Columbia (MD)

On-site
USD 138,000 - 163,000
Staff AI Software Engineer
Staff AI Software Engineer

Harnham • San Francisco (CA)

On-site
USD 150,000 - 200,000
Senior AI Engineer
Senior AI Engineer

7AI • Boston (MA)

On-site
USD 140,000 - 200,000
R111111 Senior Machine Learning Engineer III (Raleigh, NC)*
R111111 Senior Machine Learning Engineer III (Raleigh, NC)*

LexisNexis • Raleigh (NC)

Hybrid
USD 100,000 - 150,000
Senior Software Engineer with Security Clearance
Senior Software Engineer with Security Clearance

Neural Solutions • Columbia (MD)

On-site
USD 204,000 - 247,000
AI Engineer
AI Engineer

Twenty80 llc • San Francisco (CA)

On-site
USD 120,000 - 160,000
Principal AI Engineer
Principal AI Engineer

Stellantis NV • Auburn (AL)

On-site
USD 150,000 - 230,000
Lead Java Developer – AI & LLM Systems
Lead Java Developer – AI & LLM Systems

Finance Professionals Inc. • Fort Mill (SC)

Hybrid
USD 100,000 - 140,000