Why This Role Stands Out
This role offers an exciting opportunity to shape the future of AI infrastructure, working with cutting-edge technology and diverse partners to solve complex distributed systems challenges. You'll thrive here if you are passionate about building robust, scalable systems and eager to develop your expertise in scheduling, orchestration, and fault tolerance within a collaborative environment. Apply now to be at the forefront of AI innovation!
Quick Overview
Job Description
About us
Gimlet is building the first multi-silicon neocloud designed for fast, efficient AI inference.
We combine large-scale compute infrastructure with an execution platform that partitions AI workloads and maps each stage to the hardware best suited to run it.
We work with foundation labs, hyperscalers, and AI-native companies, giving our team access to technical problems spanning frontier models, production infrastructure, and emerging hardware.
About the role
As a Member of Technical Staff, you will build the systems that schedule, route, and coordinate AI workloads across Gimlet’s infrastructure.
Different stages of an inference pipeline may run on different hardware, scale independently, and exchange state across the system. Your work will determine how those workloads are placed, coordinated, routed, recovered, and operated in production.
You will work across scheduling, orchestration, control planes, APIs, and fault tolerance. You will design systems that make distributed infrastructure easier to operate, enable workloads to run reliably across a heterogeneous fleet, and partner with compiler, ML systems, networking, and infrastructure engineers to connect the full execution stack.
What success looks like
In the first 12-18 months, you will:
Build scheduling and orchestration systems for heterogeneous compute
Design systems that manage independently scalable stages of distributed inference pipelines
Improve the reliability and fault tolerance of production AI infrastructure
Develop control planes and APIs that simplify how workloads are deployed and managed
Improve resource management and scheduling as Gimlet expands across new accelerator types, nodes, and data centers
Help the platform scale across additional hardware, nodes, and data centers
You may be a good fit if you have
Experience building or operating distributed systems in production
Strong software-engineering and systems fundamentals
The ability to reason about concurrency, consistency, failure modes, and system tradeoffs
Experience with scheduling, resource management, RPC, or asynchronous messaging
A bachelor’s degree in a relevant field or equivalent practical experience
Strong candidates may also have
Experience with Kubernetes or Kubernetes-adjacent systems beyond basic usage
Experience designing service-oriented architectures using RPC or asynchronous messaging
Familiarity with scheduling, queues, or resource management systems
Experience building reliable APIs and operating systems under high load
Software development experience in languages commonly used for systems development (e.g., Go, C++, Python)
Why join now?
Gimlet is expanding from its core technology into a production neocloud spanning new hardware, customers, and data centers.
Solve hard problems.
Own meaningful work.
Build for production.
Help define what’s next.
Agency Policy: Gimlet Labs does not accept unsolicited resumes from recruitment agencies or search firms. Any unsolicited resumes submitted without a signed agreement will be considered the property of Gimlet Labs, and no fees will be paid.
Similar jobs
- AI
Senior Staff Software Engineer, Trust
NewAirbnb
Remote - US🇺🇸Remote13 hours agoScalaJavaKotlin+1Technology - AI
Staff Software Engineer, Event logging
NewAirbnb
United States🇺🇸13 hours agoC++JavaKotlinTechnology - AB
Tech Lead Manager - Product Engineering (Identity Security)
NewAbnormal
Hybrid - San Francisco🇺🇸$179.8k - $258.5k/yr8 hours agoMicroservicesAWSMachine Learning+2 - MT
Software Developer - C++ / Ada / GMD Fire Control with Security Clearance
NewMoseley Technical Services, Inc.
Huntsville, AL🇺🇸$70k - $90k/yrOn-siteYesterdayDockerMATLABAgile+9Technology - RT
Principal Software Engineer, EW Real-time Embedded with Security Clearance
NewRTX
Goleta, CA🇺🇸HybridYesterdayFPGAMachine LearningAgile+5Technology - RT
Software Engineer Intern with Security Clearance
NewRTX
Tucson, AZ🇺🇸On-siteYesterdayAgileC#C+++1Technology