Applied Scientist - Foundation Models, SSO

Amazon

Houston (TX)

On-site

USD 143,000 - 193,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Amazon Web Services (AWS) is seeking a highly skilled ML/AI specialist in Houston to design, train, and optimize large foundation models. You will enhance efficiency, latency, and memory usage while supporting multi-step task execution and tool use.

Collaboration with teams to deploy models on AWS accelerators and GPUs will be essential, including research and evaluation for reliability. The role emphasizes cutting-edge ML practices, distributed training, and hardware-aware optimizations with

Qualifications

  • Strong background in machine learning, deep learning, transformers, and foundation models.
  • Experience with PyTorch, JAX, or similar frameworks.
  • Experience training, fine-tuning, or serving large-scale models.
  • Knowledge of GPU architecture, CUDA, Triton, custom kernels, or hardware-aware optimization.
  • Familiarity with distributed training and model parallelism.
  • Experience with AWS Trainium, Inferentia, Neuron SDK, TPUs, or other custom accelerators is preferred.
  • PhD in machine learning, computer science, electrical engineering, or a related field is preferred.

Responsibilities

  • Design, train, fine-tune, and evaluate custom foundation models.
  • Improve model efficiency, stability, latency, throughput, and memory usage.
  • Build models that support planning, tool use, memory, self-correction, and multi-step task execution.
  • Optimize training and inference using CUDA, Triton, custom GPU kernels, and tensor operations.
  • Profile and resolve compute, memory, communication, and distributed-training bottlenecks.
  • Work with GPUs and specialized AI accelerators, including AWS Trainium and other custom tensor hardware.
  • Develop evaluations for model reliability and long-horizon agent performance.

Skills

ML & DL
PyTorch/JAX
Model training
GPU architecture
Distributed training
CUDA/Triton
Custom kernels

Education

PhD in ML/CS/EE or related
Master's + 4+ years CS/CE/ML

Tools

CUDA
Triton
PyTorch
JAX

Job description

Job ID: 10504260 | Amazon Web Services, Inc.

Qualifications
  • Strong background in machine learning, deep learning, transformers, and foundation models.
  • Experience with PyTorch, JAX, or similar frameworks.
  • Experience training, fine-tuning, or serving large-scale models.
  • Knowledge of GPU architecture, CUDA, Triton, custom kernels, or hardware-aware optimization.
  • Familiarity with distributed training and model parallelism.
  • Experience with AWS Trainium, Inferentia, Neuron SDK, TPUs, or other custom accelerators is preferred.
  • PhD in machine learning, computer science, electrical engineering, or a related field is preferred.
Key job responsibilities
  • Design, train, fine-tune, and evaluate custom foundation models.
  • Improve model efficiency, stability, latency, throughput, and memory usage.
  • Build models that support planning, tool use, memory, self-correction, and multi-step task execution.
  • Optimize training and inference using CUDA, Triton, custom GPU kernels, and tensor operations.
  • Profile and resolve compute, memory, communication, and distributed-training bottlenecks.
  • Work with GPUs and specialized AI accelerators, including AWS Trainium and other custom tensor hardware.
  • Develop evaluations for model reliability and long-horizon agent performance.
About the team
Diverse Experiences

AWS values diverse experiences. Even if you do not meet all of the qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn't followed a traditional path, or includes alternative experiences, don't let it stop you from applying.

Why AWS?

Amazon Web Services (AWS) is the world's most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating - that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.

Inclusive Team Culture

Here at AWS, it's in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon) conferences, inspire us to never stop embracing our uniqueness.

Mentorship & Career Growth

We're continuously raising our performance bar as we strive to become Earth's Best Employer. That’s why you'll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional.

Work/Life Balance

We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there's nothing we can’t achieve in the cloud.

Basic Qualifications
  • 3+ years of building models for business application experience
  • PhD, or Master's degree and 4+ years of CS, CE, ML or related field experience
  • Experience in patents or publications at top-tier peer-reviewed conferences or journals
  • Experience programming in Java, C++, Python or related language
  • Experience in any of the following areas: algorithms and data structures, parsing, numerical optimization, data mining, parallel and distributed computing, high-performance computing
Preferred Qualifications
  • Experience using Unix/Linux
  • Experience in professional software development

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

USA, TX, Houston - 142,800.00 - 193,200.00 USD annually

Important FAQs for current Government employees

Before proceeding, please review the following FAQs

https://www.amazon.jobs/en/faqs#faqs-for-us-government-employees

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied Scientist -- Foundation Models, SSO
Applied Scientist -- Foundation Models, SSO

Amazon • Austin (TX)

On-site
USD 143,000 - 193,000
Health insurance
401(k) matching
Paid time off
+1
Applied Scientist - Foundation Models, SSO
Applied Scientist - Foundation Models, SSO

Amazon • Santa Clara (CA)

On-site
USD 172,000 - 222,000
Health insurance
RSUs
401(k) matching
+1
Applied Scientist, AWS Quick
Applied Scientist, AWS Quick

Amazon • Santa Clara (CA)

Hybrid
USD 172,000 - 222,000
Applied Scientist, AWS Agentic AI
Applied Scientist, AWS Agentic AI

Amazon • Santa Clara (CA), Northern (KY)

Hybrid
USD 172,000 - 222,000
Health insurance
401(k) matching
Paid time off
+1
Applied Scientist, AWS Quick
Applied Scientist, AWS Quick

Amazon • New York (NY)

On-site
USD 172,000 - 223,000
Applied Scientist, AWS Quick
Applied Scientist, AWS Quick

Amazon Web Services (AWS) • Seattle (WA)

Hybrid
USD 172,000 - 222,000
RSUs
Health insurance
401(k) matching
Software Development Engineer II, AWS SageMaker AI
Software Development Engineer II, AWS SageMaker AI

Amazon • Factoria (WA)

On-site
USD 144,000 - 194,000
Applied Scientist, Personalize, AWS
Applied Scientist, Personalize, AWS

Amazon • Austin (TX)

On-site
USD 143,000 - 193,000
Health insurance (medical, dental, etc
401(k) matching
Paid time off
+1
Applied Scientist II, AWS Agentic AI
Applied Scientist II, AWS Agentic AI

Amazon • San Francisco (CA)

On-site
USD 171,000 - 223,000
Applied Scientist, Personalize, AWS
Applied Scientist, Personalize, AWS

Amazon • Seattle (WA)

On-site
USD 143,000 - 193,000
Health insurance
RSUs (restricted stock units)
401(k) matching
+2