Research Engineer, Artificial General Intelligence
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
The role
Design and implement scalable multimodal LLM pipelines for pre- and post-training stages.
Scale model training on large GPU clusters and AWS Trainium, optimizing distributed training and parallelism.
Tune low-level training components, including CUDA kernels, collectives, and network I/O to improve efficiency.
Prototype and evaluate novel algorithms using industry-leading frameworks (NeMo, Megatron Core, PyTorch, Jax, vLLM, TRT).
Collaborate across teams in an Agile environment, delivering robust features while adapting to new scientific advances.
Contribute to system architecture decisions and establish best practices for scalable AI infrastructure.
Scale model training on large GPU clusters and AWS Trainium, optimizing distributed training and parallelism.
Tune low-level training components, including CUDA kernels, collectives, and network I/O to improve efficiency.
Prototype and evaluate novel algorithms using industry-leading frameworks (NeMo, Megatron Core, PyTorch, Jax, vLLM, TRT).
Collaborate across teams in an Agile environment, delivering robust features while adapting to new scientific advances.
Contribute to system architecture decisions and establish best practices for scalable AI infrastructure.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
or
Already have an account?
Log inSimilar openings
Other roles that could suit you.
? Fellow, AI Performance Software Engineer ? Applied Machine Learning Engineer, Circuit Design - New College Grad 2026 ? Applied Machine Learning Engineer - VLSI Design ? Machine Learning Engineer II, Sponsored Products and Brands ? Software Engineer- AI/ML, AWS Neuron Distributed Training ? Applied Scientist, LLM Code Agents, Kiro Science