Turn this role into an interview — a resume and cover letter built around what this employer wants.
Nomic is hiring a Senior Platform Engineer to own the infrastructure stack—multi-account AWS, Kubernetes, IaC, CI/CD—and to ensure production agents run fast, reliably, and at scale across many customer environments.
You will influence architectural decisions, scale the footprint across continents, and maintain high-performance, cost-effective services with strong security posture and observability.
Nomic is the domain-specific AI platform for the Architecture, Engineering, and Construction (AEC) industry. We help enterprise teams extract structured knowledge from decades of drawings, specs, and project files — combining embedding models, document parsing, and autonomous agents that reason over real-world data and take action in live environments.
Nomic is hiring a Senior Platform Engineer to own our infrastructure stack — multi-account AWS, Kubernetes, IaC, CI/CD — and the systems that make our agents run well in production: orchestrating rollouts across customer environments, keeping inference fast and reliable, and maintaining industry-leading performance as we scale.
We're already deployed across dozens of companies on three continents, delivering value in production today, and we train and deploy our own models. Your job will be to scale that footprint up dramatically — more customers, more environments, more inference — without performance or reliability slipping as we grow.
This is a senior IC role with broad ownership and real architectural influence. You'll have a wide surface area and the autonomy to shape it.
Rollout and deployment. Orchestrating how our agents and services roll out across many customer cloud environments: deployment strategies, per-customer configuration, automated health checks, and the monitoring that catches problems before customers do.
Inference and performance. Keeping model inference fast, reliable, and cost-effective at scale — serving infrastructure, GPU workloads, and the performance work that keeps agents responsive as volume grows.
Core infrastructure. Kubernetes, multi-account AWS, CI/CD, observability (traces, metrics, logs, alerting, SLOs), disaster recovery, and cost management.
Security posture. Access controls, secrets management, network security, image scanning, dependency auditing, and compliance work (SOC 2, enterprise security) as customer requirements demand.
Infrastructure as code. Defining, provisioning, and evolving all infrastructure through code — designing modules, managing state, and thinking hard about blast radius.