SWE, AI Infrastructure
Palo Alto · Full-time
About the Role
You will build the platform our ML team works on. Training runs, datasets, experiments, compute, and the path from a checkpoint to a model running on a robot all sit on infrastructure someone has to own.
The measure of this role is iteration speed. When an ML engineer has an idea, how long until they know whether it worked.
What You'll Do
Own the data platform: storage, versioning, and lineage for large multimodal datasets, so any model can be traced back to the data that produced it.
Build dataset assembly and serving, including the interfaces ML engineers use to slice, filter, and load training data.
Own our ML compute, on-premise and cloud: cluster provisioning, GPU node health, drivers, networking, and storage.
Own scheduling and quota across the team, so compute is allocated sensibly and utilization stays high.
Keep the cluster reliable, including monitoring, failure recovery, and being the person who figures out why a multi-day run died.
Own training orchestration: job scheduling, multi-node runs, checkpointing, and recovery.
Build experiment tracking and comparison, so results are reproducible and a run from six months ago can be rerun.
Own the model registry and the release path onto robots, including versioning, rollout, and rollback.
Own CI and test infrastructure for the ML stack.
Work with ML engineers as internal customers, and treat their iteration speed as the metric you are accountable for.
What We're Looking For
Strong software engineering with depth in distributed systems and data infrastructure. Python, plus Go, Rust, or C++.
Experience building ML platforms that other engineers depended on daily.
Hands-on GPU cluster operations: provisioning, drivers, interconnect, and diagnosing hardware and network failures under load.
Experience with large-scale data infrastructure: object storage, columnar formats, and pipelines handling terabytes or more.
Familiarity with cluster scheduling: Kubernetes, Slurm, Ray, or equivalent.
Track record of reducing the time between an idea and a result.
Pragmatism about building versus adopting, and the judgment to know when a tool is good enough.
Bonus Points
Experience running on-premise GPU clusters, including procurement, datacenter logistics, and capacity planning.
Experience with multimodal data at scale: video, sensor, and time-series.
Experience deploying models to edge or embedded targets, including versioning and rollback on devices in the field.
Experience with data versioning and lineage tooling.
Robotics or autonomous vehicle infrastructure experience.
Compensation: Competitive base + equity