Research Engineer, Artificial General Intelligence
职位介绍
Design and implement scalable multimodal LLM pipelines for pre- and post-training stages.
Scale model training on large GPU clusters and AWS Trainium, optimizing distributed training and parallelism.
Tune low-level training components, including CUDA kernels, collectives, and network I/O to improve efficiency.
Prototype and evaluate novel algorithms using industry-leading frameworks (NeMo, Megatron Core, PyTorch, Jax, vLLM, TRT).
Collaborate across teams in an Agile environment, delivering robust features while adapting to new scientific advances.
Contribute to system architecture decisions and establish best practices for scalable AI infrastructure.
Scale model training on large GPU clusters and AWS Trainium, optimizing distributed training and parallelism.
Tune low-level training components, including CUDA kernels, collectives, and network I/O to improve efficiency.
Prototype and evaluate novel algorithms using industry-leading frameworks (NeMo, Megatron Core, PyTorch, Jax, vLLM, TRT).
Collaborate across teams in an Agile environment, delivering robust features while adapting to new scientific advances.
Contribute to system architecture decisions and establish best practices for scalable AI infrastructure.
查看完整职位
工作职责、任职要求、技能与福利 — 免费创建账号即可查看。
或
已有账户?
登录您可能也感兴趣的职位
暂无高度相似的职位 — 以下是最新职位。