OXMIQ designs GPU and AI silicon for large‑scale model inference and training, and is building the system software and analysis platforms that prove out those designs before they reach customers. We are a startup driving innovation across the full stack — from atoms to agents — and we move fast on the strength of our own tools.
The Role
The Founding Principal Architect sets the architecture and technical direction for OXMIQ's AI cloud platform, making model serving, training, model lifecycle workflows, and AI infrastructure services reliable and straightforward to use. As a founding architect, this role turns fast‑moving technology and ambiguous requirements into durable decisions, clear interfaces, and production systems that engineers and customers can trust.
Key Responsibilities
- Define the architecture for reliable model serving, training, model lifecycle workflows, and the infrastructure services that support AI development and production operations.
- Set measurable standards for performance, cost efficiency, security, observability, resilience, and developer usability, and guide teams through the tradeoffs among them.
- Evaluate inference, training, scheduling, and accelerator technologies through prototypes, benchmarks, and operational evidence, then translate the findings into sound architectural decisions.
- Establish clear requirements and interfaces with compute, storage, and networking teams so AI workloads are supported coherently across the platform.
- Use modern AI‑assisted and agentic development workflows to speed up prototyping, evaluation, and architectural validation without compromising engineering rigor.
- Provide hands‑on technical leadership for the founding engineering team: mentor engineers, lead design reviews, resolve difficult cross‑cutting decisions, and help turn architecture into operable services.
Required Qualifications
- Experience with LLM inference and serving frameworks such as vLLM, along with the performance techniques and operational practices needed to run model‑serving systems in production.
- Experience with distributed training and post‑training systems, model lifecycle workflows, and the data and infrastructure concerns that connect development with production.
- Fluency in Kubernetes, workload scheduling, observability, and accelerator‑aware infrastructure in demanding production environments.
- Deep experience architecting and operating production AI/ML platforms, model‑serving systems, or large‑scale AI infrastructure services.
- Strong knowledge of inference and training behavior, including batching, parallelism, quantization, memory use, and the ways software choices interact with accelerator hardware.
- Experience evaluating and operating more than one accelerator environment, with practical judgment about portability, performance, operational complexity, and lifecycle cost.
- Fluency in Kubernetes‑based platform design, scheduling, distributed systems, security, observability, and production operations.
- A track record of setting technical direction, mentoring senior engineers, and turning ambiguous requirements into clear architecture and delivered services.
Preferred Qualifications
- 15+ years in AI infrastructure & accelerator systems, including time in architect or equivalent technical‑leadership roles.
- Experience improving the performance or cost efficiency of inference, training, or scheduling systems through rigorous measurement.
- Familiarity with model lifecycle tooling, accelerator software stacks, compilers, or distributed communication libraries across varied hardware environments.
- Meaningful open‑source contributions in AI infrastructure, distributed systems, or platform engineering.
Education
- BS/MS/PhD in Computer Science, Computer Engineering, or a related field — or equivalent practical experience.
This is a hands‑on role with founding‑team scope, working closely with compute, storage, and networking teams to execute on the technical vision for the AI cloud platform. The Founding Principal Architect owns architecture and technical direction across model serving, training, and AI infrastructure services, and provides mentorship across the founding engineering team. AI‑assisted development tools (Claude Code or equivalent) are a standard part of engineering practice at OXMIQ and are expected in daily work.
OXMIQ offers a competitive compensation package, including base salary, equity participation, comprehensive medical coverage, and the opportunity to contribute to foundational silicon and software technology.
OXMIQ is an equal opportunity employer. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, or any other legally protected characteristic.