Software Engineer: ML Infra
Why This Role Stands Out
This Software Engineer role offers a unique opportunity to contribute to groundbreaking advancements in embodied AI and robotics, shaping the future of human-machine collaboration. If you're a mid-senior engineer passionate about building robust ML infrastructure and eager to work alongside a world-class team from leading AI and robotics labs, this position is an excellent fit for your career growth. Apply now to be part of a mission to make general intelligence useful for everyone!
Quick Overview
Job Description
About Generalist
At Generalist, we are on a mission to build general intelligence for the physical world and make it useful to everyone. We believe the industries and homes of the future will depend on humans and machines working together in new ways. Robots can help us build more and get more done.
We build embodied foundation models, starting with a focus on dexterity. This requires advancing the frontiers of data, models, and hardware, to enable robots to intelligently interact with the physical world.
The company embraces both large-scale AI and robotics as core to its DNA. Our team of researchers, roboticists, and company builders come from OpenAI, Boston Dynamics, Google DeepMind, and other frontier labs—with a track record of shipping AI breakthroughs. Before Generalist, we pioneered large embodied multimodal models and vision-language-action models (PaLM-E, RT-2, Gemini Robotics), launched and scaled ChatGPT and GPT-4 to hundreds of millions of users, engineered the foundations of autonomous driving, built next-generation robots (Atlas, Spot, Stretch) and pushed the limits of what they can do (from parkour to manipulation, and testing robustness).
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
About the Role
Generalist trains very large robot foundation models. This requires utilizing very large numbers of the latest generation GPU hardware and infrastructure (currently Nvidia) to run distributed training jobs and researcher experiments. We have extreme requirements on storage and data loading infrastructure that requires maximizing cloud infrastructure and custom solutions.
You will also own inference infrastructure. For our robots this is a fleet of on-prem GPUs attached to robots that have extreme real-time and latency budgets in compute constrained environments.
You’ll be responsible for:
Owning our GPU compute fleets
Ensure our GPUs are easy for researchers to use and maximally utilized
Optimizing and improving ML data loading transport and storage in highly distributed fully utilized environments.
Orchestration of robot inference fleets
You might thrive in this role if you:
Have managed large fleets of GPUs doing large-scale, long-term, highly distributed training runs or inference
Deep experience in Slurm or Kubernetes for ML workload orchestration
Have build high-scale ML data loaders and preparation systems
Deeply understand every layer of the ML hardware, storage, and networking stacks
Have experience in the NVidia GPU ecosystem
Similar jobs
- US
Software Engineer III
NewUSAA
SAN ANTONIO, TX🇺🇸HybridYesterdaySQLJavaPythonTechnology - GR
Software Developer (Systems Software) with Security Clearance
NewGRVTY
Chantilly, VA🇺🇸HybridYesterdayAWSAgileJavaTechnology - AS
Sr. Network Engineer with Security Clearance
NewAstrion
Columbia, MD🇺🇸$145k - $165k/yrOn-siteYesterdayAnsibleAzurePython+1Technology - PE
Senior Systems Engineer, TS/SCI w/Poly with Security Clearance
NewPeraton
Fort Meade, MD🇺🇸$176k - $282k/yrHybridYesterdayTechnology - SR
Software Developer II with Security Clearance
NewScientific Research Corporation
North Charleston, SC🇺🇸HybridYesterdayDockerOWASPActive Directory+10Technology - LM
Computing Systems/Software Engineer Sr with Security Clearance
NewLockheed Martin
Mount Laurel, NJ🇺🇸HybridYesterdayDockerTCP/IPAgile+7Technology