Why This Role Stands Out
This role offers an exceptional opportunity to shape the future of AI infrastructure at a rapidly growing, well-funded company with a reputation for innovation. You'll thrive here if you're a skilled platform engineer eager to tackle complex challenges, contribute to groundbreaking technology, and join a team of accomplished professionals. Apply today to be at the forefront of the next computing revolution!
Quick Overview
Job Description
About Us:
AI needs a new infrastructure layer. We're building it at Modal.
Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now.
Our customers include category-defining companies like Lovable, Ramp, Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale.
We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September.
Our team includes creators of popular open-source projects (e.g.,Seaborn,Luigi), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience.
The Role:
At Modal, we sell cloud services atop which our customers run their critical production systems. As a rapidly growing new cloud infrastructure company, we seek to improve our reliability dramatically while scaling the size of our platform, customer base, and our team.
This role is for people who are deep systems thinkers, love stacking nines, and thrive from making others move faster at scale. Responsibilities include:
Identifying architectural changes to improve reliability and performance.
Fostering a culture of reliability across Modal’s engineering organization.
Defining and implementing operational processes such as deployments, upgrades, etc.
Operating systems like Kubernetes, Postgres, Redis, etc.
Participating in on-call rotations, and responding to production incidents.
Requirements:
5+ years of experience writing high-quality production code.
2+ years of on-call experience for critical production services.
Strong cloud skills, and deep familiarity with at least one hyperscaler cloud (AWS preferred).
Familiarity with auto scaling, fleet management, and capacity planning at scale.
Experience operating databases, monitoring, CI/CD, and other infrastructure, at scale
Experience owning and scaling Kubernetes clusters to thousands of nodes a plus.
Experience with systems safety research (e.g. STAMP) and control theory a plus.
Ability to work in-person in our NYC or Stockholm offices.
Similar jobs
- AS
Software Engineer II
NewApex Systems
South San Francisco, CA🇺🇸On-site14 hours agoSpringJavaPython+1Technology - AS
Information Technology - Software Developer Oracle (IT)
NewApex Systems
Chicago, IL🇺🇸Hybrid14 hours agoOraclePL/SQLSOAP+3Technology - GE
Application Programmer
Genesis10
Pennington, NJ🇺🇸$58 - $66/hrHybrid3 days agoSQLSpringSpring Boot+4 - CG
Senior Lead Network Engineer with Security Clearance
NewContact Government Services, LLC
Miami, FL🇺🇸$80k - $200k/yrHybrid14 hours agoAgileJiraPenetration Testing+1Technology - SO
SST NETWORK ENGINEER with Security Clearance
NewSolution One Industries, Inc
Chaffee, AR🇺🇸Hybrid14 hours agoTCP/IPVMwareTechnology - BR
Senior Software Engineer - 1115 (Hybrid Brooklyn)
NewBraintrust
New York, NY🇺🇸On-site14 hours agoPHPAngularReactTechnology