Haystack
← Back to Jobs
Technology
MS

ChatGPT Enterprise / GenAI Platform Operations Lead (W2 Candidate Only)

Metalight Solutions IncChicago, IL🇺🇸United StatesPosted 21 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
Chicago, IL, United States
Posted
Yesterday
AWSSAMLSSOAzureDatadogGenerative AIGoogle CloudLLMStakeholder Management

Job Description

Job Title: ChatGPT Enterprise / GenAI Platform Operations Lead

Location: Chicago, IL (3days On-Site, 2days Remote)
Employment Type: Contract (W2 Candidate only)

Position Overview

We are looking for an experienced GenAI Platform Operations / SRE Lead to manage and operate enterprise-scale AI platforms, with a strong focus on ChatGPT Enterprise, OpenAI APIs, Codex, AI governance, security, monitoring, and platform reliability.

The ideal candidate will combine a strong background in SRE, CloudOps, IT Operations, or Platform Engineering with hands-on experience supporting Generative AI / LLM platforms. This role will be responsible for enterprise ChatGPT administration, AI platform security, Codex workspace operations, usage and credit optimization, monitoring, governance, and leading distributed/offshore teams.

Key Responsibilities

ChatGPT Enterprise Platform Administration & Security

  • Manage the enterprise ChatGPT administration environment, including domain verification, SSO/SAML, user provisioning, onboarding, and access management.
  • Implement and maintain security, privacy, compliance, and data protection controls to minimize enterprise data leakage risks.
  • Manage RBAC and permissions for custom GPTs, APIs, Codex/developer workspaces, and other AI platform capabilities.
  • Establish and maintain governance standards for enterprise AI usage.
  • Partner with security and compliance teams to implement responsible AI and risk mitigation practices.

Codex Agent & Workspace Operations

  • Provision, monitor, and manage local and cloud-based development environments used by autonomous Codex agents.
  • Support AI agents performing code migration, refactoring, testing, and multi-file development workflows.
  • Establish governance around agent workflows, hooks, discovery surfaces, and third-party integrations.
  • Monitor agent execution, token consumption, resource utilization, execution timelines, and failure patterns.
  • Establish controls to prevent infinite loops, excessive resource consumption, or uncontrolled AI execution.

AI Usage, Financial & Credit Management

  • Track AI platform usage, adoption, seat utilization, and consumption metrics.
  • Analyze standard versus usage-based seats and optimize workspace allocation.
  • Monitor and optimize AI platform credits and organizational consumption.
  • Develop dashboards and reports measuring AI adoption, usage patterns, productivity, and platform health.
  • Identify opportunities to reduce unnecessary consumption and improve platform ROI.

AI Operations / SRE

  • Lead day-to-day AI platform monitoring, incident triage, troubleshooting, root-cause analysis, and operational support.
  • Define operational thresholds, usage limits, policies, alerts, and escalation procedures.
  • Apply ITSM, ITIL, and SRE principles to enterprise AI operations.
  • Develop operational runbooks, incident management processes, service-level objectives, and continuous improvement frameworks.
  • Monitor AI platform health and performance using tools such as Dynatrace, Datadog, or similar platforms.

AI Process & Prompt Optimization

  • Design reusable frameworks for AI operational processes and platform optimization.
  • Guide engineering and business teams on prompt engineering and effective GenAI usage.
  • Analyze AI usage patterns and recommend improvements to workflows, policies, and platform configuration.
  • Support AI evaluation, responsible AI, risk mitigation, and quality improvement initiatives.

Team & Stakeholder Leadership

  • Lead and mentor distributed/offshore AI Operations, SRE, CloudOps, or Platform Engineering teams.
  • Establish engineering standards, operational procedures, and code/workflow quality practices.
  • Act as a liaison between business stakeholders, engineering teams, security/compliance teams, and OpenAI/vendor teams.
  • Communicate platform health, incidents, adoption, risks, and operational metrics to technical and business leadership.

Required Qualifications

  • 7 10 years of experience in SRE, IT Operations, CloudOps, DevOps, Platform Engineering, or related disciplines.
  • 1 2+ years of hands-on experience with Generative AI / LLM platforms.
  • Experience with OpenAI, ChatGPT Enterprise, OpenAI APIs, or Codex.
  • Experience administering or operating enterprise cloud/platform environments.
  • Experience with SSO/SAML, RBAC, identity/access management, and enterprise security controls.
  • Strong understanding of ITSM, ITIL, SRE, incident management, monitoring, and operational processes.
  • Experience leading offshore/distributed teams.
  • Strong stakeholder management and communication skills.

Technical Skills

Required / Strongly Preferred

  • ChatGPT Enterprise
  • OpenAI APIs
  • Custom GPTs
  • OpenAI Assistants / agent-based APIs
  • Codex
  • GenAI / LLM platforms
  • AI Operations / LLMOps
  • SRE / Platform Engineering
  • ITSM / ITIL
  • RBAC / SSO / SAML
  • Cloud platforms Azure preferred, AWS or Google Cloud Platform
  • Monitoring / Observability Dynatrace, Datadog, or similar

AI Frameworks

Experience with one or more:

  • LangChain
  • LangGraph
  • CrewAI
  • Other agentic AI / LLM orchestration frameworks

Nice to Have

  • Azure OpenAI
  • LLMOps frameworks
  • AI evaluation frameworks/tools
  • Responsible AI / AI governance
  • AI security and risk management
  • OpenAI / Azure certifications
  • Experience managing enterprise AI adoption and FinOps/AI consumption

Similar jobs