Quick Overview
Job Description
AI Performance Engineer
Client: Infosys/Microsoft
Pheonix, AZ// Seattle, WA
Rate: $55-60/hr C2C
AI Model tuning:
AI Performance Engineer — Model Optimization & Systems
About the role You will own the question: "How well do our AI models run on our hardware, and how do we make them run better?" You'll benchmark and profile training and inference workloads, identify compute, memory, and I/O bottlenecks, and recommend (and implement) optimizations — from model-level techniques like quantization and batching to system-level tuning of GPU utilization, memory bandwidth, and data pipelines.
Responsibilities
Benchmark AI models (LLMs, vision, multimodal) across hardware configurations; measure latency, throughput, utilization, memory behavior, and scaling efficiency
Profile workloads end-to-end using tools such as Nsight Systems/Compute, PyTorch Profiler, and system telemetry (nvidia-smi, DCGM) to isolate bottlenecks
Build roofline/performance models to quantify achieved vs. theoretical performance and prioritize the highest-impact optimizations
Apply and evaluate optimizations: quantization, pruning, distillation, operator/kernel fusion, graph compilation, KV-cache management, batching strategies, speculative decoding
Recommend hardware/system configurations (GPU selection, memory sizing, interconnect, storage/network I/O) for given model workloads
Establish performance baselines, SLAs, and regression testing so models stay fast as they evolve
Write clear analyses and recommendations for engineering and leadership audiences
Required qualifications
BS/MS in CS, Computer Engineering, EE, or equivalent practical experience
Strong Python; working proficiency in at least one systems language (C++/Rust/C)
Hands-on experience with a deep-learning framework (PyTorch preferred), including model execution, export, and profiling
Demonstrated experience delivering measurable performance improvements in DL training or inference
Solid grounding in computer architecture: memory hierarchy, bandwidth vs. compute limits, parallelism
Ability to reason quantitatively about latency, throughput, batching, memory footprint, and utilization under real workloads
Fluency with Linux and GPU computing environments
Preferred qualifications
GPU programming (CUDA, Triton, ROCm/HIP) and low-level libraries (cuBLAS, cuDNN, CUTLASS)
Inference runtimes/serving engines: TensorRT(-LLM), ONNX Runtime, vLLM, SGLang, Triton Inference Server
LLM inference mechanics: attention, KV caching, prefill vs. decode, continuous batching, speculative decoding
Distributed training/inference: data/tensor/pipeline parallelism, NCCL, InfiniBand/RoCE
Model compression research or MLPerf-style benchmarking experience
Edge/on-device deployment (Jetson, NPUs, Core ML) if your systems include e
Similar jobs
- FO
AI Engineer III
NewFedEx Office
Memphis, TN🇺🇸$9.2k/moRemoteYesterdayDockerExpressSQL+22Technology - SM
Sr.AI Engineer
NewSmartIMS Inc.
Minnetonka, MN🇺🇸HybridYesterdayAzurePythonTerraformTechnology - WH
Gen AI Engineer
Whiztek Corp
Chicago, IL🇺🇸Hybrid7 weeks agoMachine LearningDeep LearningGPT+5Technology - EC
AI Test Engineer with Security Clearance
ECS
Fairfax, VA🇺🇸$125k - $150k/yrOn-site1 week agoSQLMachine LearningSelenium+3Technology - AS
Gen AI Engineer Lead
NewApex Systems
Ridgeland, MS🇺🇸HybridYesterdayDockerFastAPIFlask+27Technology - OR
AI Architect
NewOrangePeople
Los Angeles, CA🇺🇸HybridYesterdayAWSAzureCircleCI+4Technology