Haystack
← Back to Jobs
Full time
Engineering
TI

Kernel Engineer (Internship and Full-time)

TilderesearchSan Francisco🇺🇸United StatesPosted Jul 14, 2025

Why This Role Stands Out

This role offers an exceptional opportunity to push the boundaries of AI by optimizing high-performance GPU kernels, directly impacting the advancement of foundational AI models. You'll thrive here if you have a passion for deep learning, demonstrated expertise in ML kernels, and a drive to collaborate with researchers on cutting-edge infrastructure. Apply now to contribute to a leading AI research lab and accelerate your career in this exciting field.

Quick Overview

Seniority
Entry Level
Employment type
Full Time
Work mode
On Site
Location
San Francisco, United States
Posted
1 year ago

Job Description

Tilde Research is a moonshot AI lab advancing mechanistic interpretability, new architectures, and pretraining science. We build foundational understanding of models to advance the frontier of intelligence.


About the role:

As a Kernel Engineer at Tilde, you'll design, implement, and optimize high-performance GPU kernels that are critical to scaling our training and inference workloads. Your work will enable faster iteration cycles, higher throughput, and lower latency. You'll work closely with ML researchers and engineers to co-design models and infrastructure that are deeply performance-aware, and help push the limits of what current hardware can support.

What you might work on:

  • Design, develop, and tune custom GPU kernels for core model operations

  • Work with ML engineers to prototype and scale novel model architectures

  • Contribute to system-wide efforts to improve efficiency and throughput, beyond just kernel-level optimizations


You're a good fit if you:

  • Have experience in deep learning or related research areas

  • Have demonstrated exceptional capability in working on ML kernels. This can include:

    • Strong open source contributions

    • Thoughtful technical blog posts/work logs

    • Previous experience working on hardware-aligned algorithms

  • Deep familiarity with PyTorch, Triton/TK/TileLang (>1 of), basic familiarity with CUDA, and knowledge of GPU architecture.

  • Communicate clearly and effectively, both verbally and in writing

  • Strong algorithmic thinker

  • Are able to learn quickly

Similar jobs