active · Approximate location
Engineering Manager, ML Performance Optimization
full timeFoster City, CA
Create a free profile to personalize salary, commute and role fit.
AT A GLANCE
TechnologyManager level3+ years experienceFull Time
Required
- 3+ years experience
- People Leadership
Zoox is on a mission to reimagine transportation and build autonomous robotaxis from the ground up that are safe, reliable, clean, and enjoyable for everyone. With bidirectional driving capabilities and four-wheel steering, our vehicle allows us to maneuver through compact spaces
About the role
Zoox is on a mission to reimagine transportation and build autonomous robotaxis from the ground up that are safe, reliable, clean, and enjoyable for everyone. With bidirectional driving capabilities and four-wheel steering, our vehicle allows us to maneuver through compact spaces and change directions without needing to reverse. We are at a critical inflection point as we scale our robotaxi deployment, and it is a great time to join Zoox and have a significant impact on executing our mission.
Our growing ML performance engineering leadership team is looking for an Engineering Manager, ML Performance Optimization. The centralized ML performance team at Zoox plays a crucial role in enabling innovations across all our ML research teams to develop and deploy models across our robotaxi and cloud infrastructure and to advance cutting-edge training and inference optimization techniques.
The Opportunity
We are working on many interesting challenges to enable rapid experimentation and scale our multi-modal Foundation models and RL infrastructure, and ensure these models run efficiently on our vehicles, meeting our latency targets. You will get to work across all ML teams within Zoox - Perception, Prediction, Planner, Simulation and our Advanced Hardware Engineering group, and have the opportunity to significantly push the boundaries of ML scaling and acceleration at Zoox.
You will lead a team of strong ML performance engineers and this team has many growth opportunities as we expand our geofences at U.S markets and venture into new ML domains. If you want to learn more about our ML Infrastructure, here is one of our past talks at re:Invent.
Requirements
- 8+ years of relevant experience, including 3+ years of management experience managing engineers.
- Strong technical background in ML performance optimization, such as distributed training strategies (data, tensor, pipeline parallelism, FSDP/ZeRO), mixed-precision training, kernel-level optimization (CUDA, Triton), compiler stacks (torch.compile, XLA, TVM), quantization, and profiling/benchmarking across GPU and embedded accelerators.
- Experience building user-friendly ML Infrastructure that enabled large-scale model training and high-throughput, low-latency serving use cases.
- Experience with training frameworks like PyTorch, JAX, etc., leveraging GPUs for distributed model training.
- Experience with GPU-accelerated inference using TensorRT, Ray Serve, or similar frameworks.
- Proven track record of extensive cross-functional collaboration, partnering with research, product, hardware, and platform teams to align priorities, influence technical direction, and deliver measurable performance improvements across organizational boundaries.
Other Jobs From This Employer
Embedded Software Engineer - Drive SystemsFoster City, CA
Salary not listedFinancial Analyst, Production ManufacturingFoster City, CA
Salary not listedQuality EngineerFremont, CA
Salary not listedRecruiter, Business & OperationsFoster City, CA
Salary not listedRelease ManagerFoster City, CA
Salary not listed