Skip to content
Whileresume
Leaderboard Build my CV Hire Log in

Staff Software Engineer, ML Frameworks & Efficiency

Is this job for you?

Build your CV and see how well you match this role — and every other one.

Build my CV
Track your applications on mobile The free Whileresume app, on iPhone and Android.

The role

Optimize distributed ML systems for high performance on TPU and GPU clusters, applying SPMD, MPMD, and FSDP to scale model training.
Improve accelerator FLOPS efficiency by refining compiler optimizations (XLA), authoring low-level kernels (Pallas, Triton), and enabling low-precision computation.
Develop new neural model architectures (e.g., sparse architectures) and decoding strategies (e.g., speculative decoding) for improved training and inference on modern hardware.
Evaluate and integrate open-source and state-of-the-art technologies to enhance the performance and scalability of ML workloads.
Promote best practices for distributed systems architecture and contribute to technical leadership within the team.
Qualifications include a strong CS/math background, experience with ML frameworks (TensorFlow, JAX, XLA), Python and C++, and proficiency with profiling tools.

See the full job post

Responsibilities, requirements, skills and benefits — create your free account.

At least 6 characters. The longer, the safer.
or

Already have an account?

You might also like these jobs

No closely matching jobs yet — here are the most recent ones.

Your location

Jobs and companies will be filtered on this country.

Suggested

All countries 65