Skip to content

Software Engineer- AI/ML, AWS Neuron Distributed Training

Is this job for you?

Build your CV and see how well you match this role — and every other one.

Build my CV

The role

Design, implement, and optimize distributed training solutions for large-scale ML models running on Trainium instances.
Extend and optimize popular distributed training frameworks, including FSDP, torchtitan, and Hugging Face libraries for the Neuron ecosystem.
Profile, analyze, and tune end-to-end training models and pipelines to achieve optimal performance on Trainium hardware.
Partner with hardware, compiler, and runtime teams to influence system design and unlock new capabilities.
Work directly with solution architects and customers to deploy and optimize training workloads at scale.
Operate at the intersection of cutting-edge ML research and high-performance systems, contributing to training workloads such as LLMs, Dense and Mixture-of-Experts architectures, and multimodal models.

See the full job post

Responsibilities, requirements, skills and benefits — create your free account.

At least 8 characters, one uppercase letter and one digit.

Already have an account? Log in

Similar openings

Other roles that could suit you.

See all →

Your location

Jobs and companies will be filtered on this country.

Suggested

All countries 65