Haystack
← Back to Jobs
Remote
Technology
IT

AI Engineer

ISite Technologies IncUnited States🇺🇸United StatesPosted 8 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
Yesterday
RustCUDAC++LLMPyTorchPython

Job Description

Job Role: AI Engineer
Experience: 10 Years
Location: Remote
Visa: Any

Job Description:


You will own the question: “How well do our AI models run on our hardware, and how do we make them run better?”  
You’ll benchmark and profile training and inference workloads, identify compute, memory, and I/O bottlenecks, and recommend (and implement) optimizations — from model-level techniques like quantization and batching to system-level tuning of GPU utilization, memory bandwidth, and data pipelines.

Responsibilities
Benchmark AI models (LLMs, vision, multimodal) across hardware configurations; measure latency, throughput, utilization, memory behavior, and scaling efficiency

Profile workloads end-to-end using tools such as Nsight Systems/Compute, PyTorch Profiler, and system telemetry (nvidia-smi, DCGM) to isolate bottlenecks

Build roofline/performance models to quantify achieved vs. theoretical performance and prioritize the highest-impact optimizations

Apply and evaluate optimizations: quantization, pruning, distillation, operator/kernel fusion, graph compilation, KV-cache management, batching strategies, speculative decoding

Recommend hardware/system configurations (GPU selection, memory sizing, interconnect, storage/network I/O) for given model workloads

Establish performance baselines, SLAs, and regression testing so models stay fast as they evolve

Write clear analyses and recommendations for engineering and leadership audiences

Required Qualifications
BS/MS in CS, Computer Engineering, EE, or equivalent practical experience

Strong Python; working proficiency in at least one systems language (C++/Rust/C)

Hands-on experience with a deep-learning framework (PyTorch preferred), including model execution, export, and profiling

Demonstrated experience delivering measurable performance improvements in DL training or inference

Solid grounding in computer architecture: memory hierarchy, bandwidth vs. compute limits, parallelism

Ability to reason quantitatively about latency, throughput, batching, memory footprint, and utilization under real workloads

Fluency with Linux and GPU computing environments

Preferred Qualifications
GPU programming (CUDA, Triton, ROCm/HIP) and low-level libraries (cuBLAS, cuDNN, CUTLASS)

Inference runtimes/serving engines: TensorRT(-LLM), ONNX Runtime, vLLM, SGLang, Triton Inference Server

LLM inference mechanics: attention, KV caching, prefill vs. decode, continuous batching, speculative decoding

Distributed training/inference: data/tensor/pipeline parallelism, NCCL, InfiniBand/RoCE

Model compression research or MLPerf-style benchmarking experience

Edge/on-device deployment (Jetson, NPUs, Core ML) if your systems include edge hardwareb

Similar jobs