Haystack
← Back to Jobs
Other
IC

Senior LLM Engineer – Generative AI- 10+ yrs- New York, United States- Onsite

iMedhas Consulting ServicesNew York, NY🇺🇸United StatesPosted Sep 29, 2026

Why This Role Stands Out

This onsite role at iMedhas Consulting Services offers a unique opportunity to lead the development and deployment of cutting-edge Generative AI solutions within a secure, on-premises environment, perfect for experienced LLM Engineers passionate about deep technical challenges and enterprise-level AI innovation. You'll gain invaluable experience fine-tuning open-weight models and optimizing local serving engines, contributing directly to impactful AI advancements.

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
New York, NY, United States
Posted
17 hours ago
DockerMachine LearningScikit-learnCUDAData PrivacyDeep LearningGenerative AIHugging FaceKubernetesLLMPyTorchPythonTensorFlow

Job Description

Job Description :

Role OverviewWe are seeking an experienced LLM Engineer with 7+ years of software engineering experience, including 4+ years dedicated to AI/ML. You will design, develop, fine-tune, and deploy state-of-the-art Large Language Models (LLMs) and Generative AI applications directly within our on-premises, air-gapped enterprise infrastructure.In this role, you will lead the end-to-end lifecycle of local GenAI solutions—from self-hosted model serving and custom prompt engineering to fine-tuning open-weight models (e.g., Llama 3, Mistral, Qwen) while ensuring strict enterprise data privacy, security, and low latency.Required Qualifications & Technical SkillsExperience: 7+ years of overall software development experience, with 4+ years of hands-on experience in Machine Learning, Deep Learning, and AI.Python Mastery: Expert-level Python skills and deep familiarity with core AI ecosystems: PyTorch, TensorFlow, Hugging Face (transformers, peft, datasets, accelerate), spaCy, and Scikit-Learn.Self-Hosted / Open-Source LLMs: Hands-on experience working with open-weight foundation models (Llama, Mistral, Gemma, DeepSeek, Qwen) and local serving engines (vLLM, Ollama, TensorRT-LLM, Triton).On-Prem Infrastructure & Orchestration: Solid understanding of Linux, Docker/Kubernetes (OpenShift, Rancher, microK8s), local GPU orchestration, and CUDA driver configurations.Deployments: Proven track record of deploying at least one end-to-end GenAI application in a production environment.Education & Core Competencies: Bachelor’s or Master’s degree in Computer Science, Data Science, AI, or a related quantitative field. Strong problem-solving, analytical, and cross-functional communication skills.

Similar jobs