MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling)
Why This Role Stands Out
As an MLOps Engineer focused on LLM Systems, you'll play a pivotal role in a leading AI lab, contributing to the development of foundational AI models with significant growth potential. If you have hands-on experience in areas like GPU kernel programming, performance profiling, or inference serving, and thrive in a remote, collaborative environment, this opportunity offers competitive compensation and the chance to shape the future of AI. Apply now to join a cutting-edge team and accelerate your career in the dynamic field of artificial intelligence.
Quick Overview
Job Description
This role is for one of our clients
Compensation: $90-$120 per hour
Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking MLOps Engineers with hands-on experience in large language model infrastructure across any of four areas: GPU kernel programming, performance profiling and trace analysis, debugging accelerated and distributed workloads, and high-throughput inference serving. This role involves AI model training and evaluation work, including writing and assessing MLOps and ML systems tasks and solutions to generate high-quality training data for frontier AI systems.
Key Responsibilities
- Design challenging, domain-relevant tasks across four areas, GPU kernels, performance profiling, debugging, and inference serving, and write accurate, well-structured solutions to them.
- Guide research and engineering teams to close knowledge gaps and improve AI model performance on ML systems, training infrastructure, and framework-level topics.
- Evaluate MLOps and ML systems tasks and solutions, and provide clear, written technical feedback that stands up to reviewer scrutiny.
- Develop guidelines and detailed rubrics or evaluation frameworks covering kernel-level optimization, profiler output interpretation, distributed systems reasoning, and serving throughput and latency trade-offs.
- Collaborate with other subject matter experts to keep training data consistent and accurate.
Core Qualifications
- 2+ years of hands-on professional experience in ML systems, ML infrastructure, model serving, or GPU and accelerator performance engineering. This is a hands-on systems role rather than an applied modelling or data science one.
- Practical experience in at least one of the following, with more than one a strong plus: writing or optimizing custom GPU kernels (CUDA, Triton, Pallas); performance profiling and trace analysis (Kineto, torch.profiler, Nsight, XLA or JAX profiler); debugging distributed or accelerator-bound workloads; serving large language models at scale (vLLM, SGLang, TensorRT-LLM, Ray Serve, KV cache, paged attention, continuous batching).
- Working production experience with JAX and/or PyTorch. Framework-level depth is a strong plus: custom operators, distributed training (FSDP, DDP, DeepSpeed, Megatron), or compiler and graph-level work.
- Familiarity with modern accelerators such as A100, H100, B200 or TPU, and the ability to reason about throughput, latency and memory trade-offs.
- Demonstrable career progression.
- Ability to engage reliably for at least 40 hours/week during weekdays.
- Strong written communication skills and the ability to explain complex technical decisions clearly.
Similar jobs
- ED
Machine Learning Scientist - Hybrid - Hove, UK
NewEDF
hove🇬🇧Remote3 hours agoMachine LearningScikit-learnComputer Vision+6Technology - TR
Machine Learning Scientist- Mid-Level
NewTripadvisor
London🇬🇧22 hours agoSQLETLMachine Learning+6Technology - MG
Machine Learning Engineer (Voice) - London
NewMicrotech Global Ltd
South East London, London🇬🇧Hybrid15 hours agoMachine LearningLLMTechnology - NE
Senior ML Engineer (AI Research/ Portability)
NewNebius
Israel; Remote - Europe; United Kingdom🇬🇧RemoteYesterdayRustMachine LearningPython+1Technology - MP
Lead AI / ML Engineer
NewMichael Page
Knaphill, Surrey🇬🇧£80k - £95k/yrHybrid21 hours agoSQLMLOpsMLflow+4Technology - AB
Machine Learning Engineer II - Behavioral Security Products
NewAbnormal
Remote - UK🇬🇧RemoteYesterdaySQLAWSMLOps+8Technology