Research Engineer, Artificial General Intelligence
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
Track your applications on mobile The free Whileresume app, on iPhone and Android.
The role
Design and implement scalable multimodal LLM pipelines for pre- and post-training stages.
Scale model training on large GPU clusters and AWS Trainium, optimizing distributed training and parallelism.
Tune low-level training components, including CUDA kernels, collectives, and network I/O to improve efficiency.
Prototype and evaluate novel algorithms using industry-leading frameworks (NeMo, Megatron Core, PyTorch, Jax, vLLM, TRT).
Collaborate across teams in an Agile environment, delivering robust features while adapting to new scientific advances.
Contribute to system architecture decisions and establish best practices for scalable AI infrastructure.
Scale model training on large GPU clusters and AWS Trainium, optimizing distributed training and parallelism.
Tune low-level training components, including CUDA kernels, collectives, and network I/O to improve efficiency.
Prototype and evaluate novel algorithms using industry-leading frameworks (NeMo, Megatron Core, PyTorch, Jax, vLLM, TRT).
Collaborate across teams in an Agile environment, delivering robust features while adapting to new scientific advances.
Contribute to system architecture decisions and establish best practices for scalable AI infrastructure.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
or
Already have an account?
Log inYou might also like these jobs
No closely matching jobs yet — here are the most recent ones.
? Data Scientist, GTM Analytics ? Senior Software Engineer, ML Platform ? Research Scientist II, Amazon Industrial Robotics ? Applied Scientist, Skill Builder, Training & Certification ? Staff Machine Learning Engineer - Infrastructure ? Software Engineer, AI Studio iOS Mobile - New York or Mountain View