Skip to content

Staff Software Engineer, ML Frameworks & Efficiency

Is this job for you?

Build your CV and see how well you match this role — and every other one.

Build my CV

The role

Optimize distributed ML systems for high performance on TPU and GPU clusters, applying SPMD, MPMD, and FSDP to scale model training.
Improve accelerator FLOPS efficiency by refining compiler optimizations (XLA), authoring low-level kernels (Pallas, Triton), and enabling low-precision computation.
Develop new neural model architectures (e.g., sparse architectures) and decoding strategies (e.g., speculative decoding) for improved training and inference on modern hardware.
Evaluate and integrate open-source and state-of-the-art technologies to enhance the performance and scalability of ML workloads.
Promote best practices for distributed systems architecture and contribute to technical leadership within the team.
Qualifications include a strong CS/math background, experience with ML frameworks (TensorFlow, JAX, XLA), Python and C++, and proficiency with profiling tools.

See the full job post

Responsibilities, requirements, skills and benefits — create your free account.

At least 8 characters, one uppercase letter and one digit.

Already have an account? Log in

Similar openings

Other roles that could suit you.

See all →

Your location

Jobs and companies will be filtered on this country.

Suggested

All countries 65