active · Approximate location
Staff Software Engineer, HPC
full timeFoster City, CA
Create a free profile to personalize salary, commute and role fit.
AT A GLANCE
TechnologySeniorFull Time
Required
- Scheduling
Zoox is looking for an experienced Staff Software Engineer to build, scale, and operate our custom High-Performance Computing infrastructure. As Zoox scales its autonomous vehicle development, our HPC platform must keep pace with rapidly growing compute, storage, and scheduling d
About the role
Zoox is looking for an experienced Staff Software Engineer to build, scale, and operate our custom High-Performance Computing infrastructure. As Zoox scales its autonomous vehicle development, our HPC platform must keep pace with rapidly growing compute, storage, and scheduling demands across the company. You will modernize our HPC platform—built on industry-leading technologies like Ray.io, SLURM, and Kubernetes—with a focus on reliability, scalability, and world-class developer velocity.
These HPC services form the backbone of development workflows across all Zoox software teams, from data engineering to training our AI models in Perception, Planner, Prediction, to Simulation, and more. You will have a direct impact on the productivity and effectiveness of every engineering team at Zoox.
The position comes with a high degree of independence and the opportunity to define Zoox's HPC platform strategy, both technically and organizationally. You will work closely with stakeholders in Autonomy and Software teams to understand their workload requirements and translate them into robust, scalable infrastructure.
Requirements
- Experience designing and operating large-scale distributed systems in production
- Experience with Ray.io, particularly Ray Core and Ray Data (or equivalent technologies)
- Experience with Kubernetes, particularly for heterogeneous workloads
- Experience with cloud infrastructure on AWS or similar providers
- Track record of shipping and operating reliable, highly available scalable infrastructure
- Demonstrated ability to prioritize development work and build cross-functional consensus around technical tradeoffs
- Proficiency with Python
- Exposure to machine learning workloads (training, inference, data generation)
- Experience with Kubernetes or SLURM at scale (>10k+ nodes)
- Experience with SLURM workload manager and advanced scheduling policies
- Background in algorithmic optimization or operations research
- Experience building developer tools and platforms used by large engineering organizations
Other Jobs From This Employer
Senior Vehicle Attributes EngineerFoster City, CA
Salary not listedSenior Strategic Sourcing ManagerBoston, MA
Salary not listedStaff Compensation ManagerFoster City, CA
Salary not listedManufacturing Test EngineerFoster City, CA
Salary not listedSenior Machine Learning Engineer - Perception 3D SegmentationFoster City, CA
Salary not listed