Haystack
← Back to Jobs
Other
IW

AI Performance Engineer

Info Way SolutionsPhoenix, AZ🇺🇸United StatesPosted 3 Sept 2026

Quick Overview

Salary
$55 - $60/hr
Seniority
Mid Senior
Work mode
Hybrid
Location
Phoenix, AZ, United States
Posted
Yesterday
RustCUDAC++LLMPyTorchPython

Job Description

AI Performance Engineer

Client: Infosys/Microsoft

Pheonix, AZ// Seattle, WA

Rate: $55-60/hr C2C

 

AI Model tuning:
AI Performance Engineer — Model Optimization & Systems
About the role You will own the question: "How well do our AI models run on our hardware, and how do we make them run better?" You'll benchmark and profile training and inference workloads, identify compute, memory, and I/O bottlenecks, and recommend (and implement) optimizations — from model-level techniques like quantization and batching to system-level tuning of GPU utilization, memory bandwidth, and data pipelines.
Responsibilities
Benchmark AI models (LLMs, vision, multimodal) across hardware configurations; measure latency, throughput, utilization, memory behavior, and scaling efficiency
Profile workloads end-to-end using tools such as Nsight Systems/Compute, PyTorch Profiler, and system telemetry (nvidia-smi, DCGM) to isolate bottlenecks
Build roofline/performance models to quantify achieved vs. theoretical performance and prioritize the highest-impact optimizations
Apply and evaluate optimizations: quantization, pruning, distillation, operator/kernel fusion, graph compilation, KV-cache management, batching strategies, speculative decoding
Recommend hardware/system configurations (GPU selection, memory sizing, interconnect, storage/network I/O) for given model workloads
Establish performance baselines, SLAs, and regression testing so models stay fast as they evolve
Write clear analyses and recommendations for engineering and leadership audiences
Required qualifications
BS/MS in CS, Computer Engineering, EE, or equivalent practical experience
Strong Python; working proficiency in at least one systems language (C++/Rust/C)
Hands-on experience with a deep-learning framework (PyTorch preferred), including model execution, export, and profiling
Demonstrated experience delivering measurable performance improvements in DL training or inference
Solid grounding in computer architecture: memory hierarchy, bandwidth vs. compute limits, parallelism
Ability to reason quantitatively about latency, throughput, batching, memory footprint, and utilization under real workloads
Fluency with Linux and GPU computing environments
Preferred qualifications
GPU programming (CUDA, Triton, ROCm/HIP) and low-level libraries (cuBLAS, cuDNN, CUTLASS)
Inference runtimes/serving engines: TensorRT(-LLM), ONNX Runtime, vLLM, SGLang, Triton Inference Server
LLM inference mechanics: attention, KV caching, prefill vs. decode, continuous batching, speculative decoding
Distributed training/inference: data/tensor/pipeline parallelism, NCCL, InfiniBand/RoCE
Model compression research or MLPerf-style benchmarking experience
Edge/on-device deployment (Jetson, NPUs, Core ML) if your systems include e

Similar jobs