Why This Role Stands Out
This role offers an exceptional opportunity to significantly impact engineering velocity by building cutting-edge infrastructure for agentic development, making it ideal for proactive and innovative DevOps professionals. You will thrive here if you possess a relentless problem-solving approach and a passion for rapid execution within the exciting AI domain. Apply now to join a forward-thinking team focused on accelerating innovation.
Quick Overview
Job Description
Member of Technical Staff — DevOps
We look for infrastructure engineers who are obsessed with engineering velocity. Everything we do — research, training, serving, customer deployments — moves only as fast as our tooling. Your mission is to design, build, and operate the platform underneath it all, including the internal infrastructure for agentic development, so every engineer and researcher can iterate rapidly with agents doing more of the work.
Responsibilities
Measure developer velocity, and partner with researchers, product engineers, and forward-deployed engineers to find what slows them down and remove it
Build internal infrastructure for agentic development: sandboxes, tool access, context, and eval harnesses that let autonomous research and coding agents run end to end
Self-host and serve open-source LLMs for internal agent use
Own CI/CD, build and test infrastructure, reproducible dev environments, and infrastructure-as-code so research and product code ship quickly and safely
Build the monitoring, alerting, and on-call for our product and customer-facing services to catch failures before users do
Own cost across the dev stack: track and cut spend on compute, tools, and inference
What we're looking for
We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains.
AI-pilled: uses coding agents daily and has built internal infrastructure for agentic development, not just used it
Experience self-hosting and serving open-source LLMs (e.g. vLLM, SGLang)
Strong systems background: Linux, networking, containers and orchestration (e.g. Docker, Kubernetes), infrastructure-as-code (e.g. Terraform)
Experience building CI/CD and developer tooling (e.g. GitHub Actions, Buildkite, Bazel, Nix)
Knowledge of cloud platforms (GCP, AWS, or Azure) and best practices for observability and security
Track record of measurably improving developer velocity
Owns deliverables end-to-end, from requirements through autonomous execution
Similar jobs
- AM
Site Reliability Engineer (SRE)
NewApplication Management Services LLC
Pittsburgh, PA🇺🇸On-site20 hours agoDockerShellAWS+7Technology - AI
Datascience platform and Devops Engineer
NewAsterism IT Solutions
Juno Beach, FL🇺🇸Hybrid20 hours agoAWSCDKCloudFormation+3Technology - IS
Sr Infrastructure Engineer
NewIFLOWSOFT Solutions Inc.
Austin, TX🇺🇸On-site20 hours agoAzurePowerShellVMwareTechnology - IS
DevSecOps Engineer
NewInternational Solutions Group
United States🇺🇸Hybrid20 hours agoEngineering - FI
Senior Platform Engineer
Fisher Investments
Camas, WA🇺🇸$120k - $165k/yrHybrid4 days agoSQLMLOpsMachine Learning+9Technology - KG
Senior Database Reliability Engineer (DBRE)
NewK&K Global Talent Solutions
United States🇺🇸Remote20 hours agoMySQLRubyAWS+13Technology