Quick Overview
Seniority
Mid Senior
Work mode
On Site
Location
Dallas, TX, United States
Posted
Yesterday
KubernetesPython
Job Description
Hello,
Hiring!! C2C role||
Job Title: AI Inference Engineer / Inference SME
Job Location: Plano/Dallas, TX or Bedminster, NJ
Interview: Final, in-person client interview at any of the locations given
Must have skills: vLLM, model inference/serving, transformer architecture, GPU workloads, Python. AI/model-serving SME.
Job Nice to Haves: NVIDIA Triton Inference Server, NVIDIA NIMs, GPU-based inference, Model deployment/hosting, Kubernetes, Python/software development, GPU/inference monitoring
About the Position / Current Initiatives: This is the person who deeply understands how AI models are actually hosted and served at scale. They'll work directly with the inference engines and help manage/model the actual serving environment not simply consume OpenAI or other APIs.
Position Summary: This role is crucial for understanding how AI models are hosted and served at scale. The candidate will work directly with inference engines, managing and modeling the serving environment beyond simply consuming APIs from providers like OpenAI. The position requires deep technical expertise, particularly in vLLM, AI/model inference, and transformer-based architecture. While enterprise experience is not mandatory, a strong technical background is essential.
What type of experience does the right candidate have:
Extensive knowledge of vLLM and AI/model inference
Experience with transformer-based architectures
Familiarity with GPU-based inference and related technologies
What the responsibilities are of the right candidate:
Directly manage inference engines and the model serving environment
Monitor and troubleshoot GPU utilization, traffic, and memory
Ensure model performance and address any misbehavior or tier 3 debugging issues
Collaborate on deploying models at scale using technologies like NVIDIA Triton Inference Server and Kubernetes
Similar jobs
- ME
AI Engineer
NewMetaRPO
United States🇺🇸HybridYesterdayMachine LearningScrumAgile+4Technology - AS
AI Developer
NewApex Systems
Scottsdale, AZ🇺🇸HybridYesterdayAWSAzureGenerative AI+4Technology - VI
Senior Applied AI Engineer - Software Engineering with Security Clearance
NewVisionist, Inc.
Columbia, MD🇺🇸$170k - $240k/yrHybridYesterdaySQLLLMPython+1Technology - SI
Sr AI Engineer
NewSpar Information Systems
Seattle, WA🇺🇸HybridYesterdayMachine LearningPythonTechnology - BT
AI Architect with Security Clearance
NewBespoke Technologies Inc.
Tysons, VA🇺🇸HybridYesterdayAWSETLAirflow+5Technology - RI
Staff Software Engineer, AI Developer Productivity - Agent Platform & Evaluation
Rivian
Palo Alto, CA🇺🇸$206.5k - $258.1k/yrHybrid1 week agoRustAWSDatabricks+5Technology