Haystack
← Back to Jobs
Remote
Technology

Senior AI Platform Engineer

NextGen IT Inc.United States🇺🇸United StatesPosted 4 Aug 2026

Quick Overview

Work Type
Remote
Level
Mid Senior

Job Description

Role: Senior AI Platform Engineer

Location: Remote

Duration: Long term contract

About the Role

We are looking for a Senior AI Platform Engineer to build the foundational infrastructure required to securely deploy, govern, observe, and scale enterprise AI capabilities. The role will start with a focused 3 6 month mandate to enable coding agents, including GitHub Copilot and Claude Code using BYOK deployment patterns, with model access served through Parsons-controlled cloud environments across Azure AI Foundry and AWS Bedrock.

In parallel, you will extend that same foundation into a broader AI platform capability for enterprise AI applications that require governed model access, MCP-based access to tools and data, secure agent runtime patterns, and reusable agent development platforms in Azure AI Foundry and AWS AgentCore.

Phased Mandate: Phase 1 focuses on standing up the secure AI enablement foundation for coding agents. Phase 2 extends the same platform patterns to broader AI applications, enterprise MCP access, and agent development platforms in Azure AI Foundry and AWS AgentCore.

What You Will Do: Phased Roadmap
Phase 1: Build the Coding-Agent Foundation, First 3 6 Months
  • Deploy governed BYOK patterns: implement reusable infrastructure patterns for GitHub Copilot and Claude Code where model calls are routed through Parsons-controlled cloud tenancy rather than unmanaged vendor defaults.
  • Administer GitHub Copilot SaaS where required: govern GitHub Copilot SaaS administration for users who remain on the standard SaaS model while most users are enabled through approved BYOK patterns, including license assignment, access policies, usage visibility, policy configuration, and coordination with GitHub/M365 administrators.
  • Integrate Azure AI Foundry and AWS Bedrock: support approved model hosting, model routing, model access policies, and secure connectivity across Azure and AWS government/commercial boundary patterns as approved by architecture, security, and compliance.
  • Create reusable onboarding patterns: build infrastructure templates, sample repositories, deployment pipelines, and setup guides so development teams can adopt approved coding-agent capabilities quickly, consistently, and securely.
  • Design for regulated developer environments: support secure developer workstations, GovCloudC High constraints, private networking, identity federation, secrets management, and least-privilege access patterns.
  • Create the bridge to broader AI platform reuse: design the coding-agent foundation as reusable enterprise infrastructure, not a one-off implementation, so model access, policy, observability, and MCP patterns can be extended to other AI applications.
2. Establish AI Gateway and MCP Gateway Capabilities
  • Azure AI Gateway / APIM: design and operate gateway controls for model routing, authentication, authorization, model allow-lists, token budgets, rate limits, circuit breakers, content filtering, and usage attribution.
  • Enterprise MCP Gateway: build a secure broker layer that allows coding agents to access enterprise data, APIs, developer tools, repositories, ticketing systems, and approved SaaS services without exposing direct unmanaged connections.
  • Connector governance: define onboarding, registration, approval, versioning, and decommissioning patterns for MCP servers, APIs, tools, and data connectors.
  • Trust-boundary enforcement: make data classification, approved data types, tenant boundaries, cross-cloud routing, and egress controls explicit and auditable.
3. Implement Governance, Security, and Compliance Controls
  • Policy enforcement: implement technical controls for model eligibility, developer access, approved use cases, prompt/response handling, data loss prevention, and tool-call authorization.
  • Compliance alignment: partner with Security, Legal, Risk, and Federal stakeholders to design controls aligned to NIST 800-53, NIST 800-171, CMMC, FedRAMP, FISMA, and DoD impact-level requirements where applicable.
  • CUI-aware architecture: design patterns that support Controlled Unclassified Information handling requirements, including tenant selection, authorized model endpoints, encryption, audit logging, and controlled egress.
  • Human oversight: implement review, approval, exception, escalation, and rollback workflows for high-risk coding-agent use cases and sensitive integrations.
4. Build Observability, AgentOps, and FinOps
  • Unified telemetry: capture model calls, prompts/responses metadata, tool calls, data access, latency, errors, cost, token usage, policy decisions, and user/team attribution.
  • Common AI FinOps dashboard: build shared reporting for token consumption, model usage, API calls, cost attribution, budget thresholds, anomaly detection, adoption trends, forecast vs. actual spend, and showback/chargeback across teams, applications, coding agents, and enterprise AI agents.
  • Evaluation and quality controls: define test harnesses, regression evaluations, red-team scenarios, hallucination/failure detection, and release gates for coding-agent workflows.
  • Production readiness: establish SLOs, runbooks, alerting, incident response, root-cause analysis, and continuous hardening practices for AI gateway and MCP gateway services.
5. Enable Developers and Platform Adoption
  • Developer onboarding: create clear onboarding guides, sample projects, secure defaults, approved model-selection guidance, and training materials for development teams.
  • Reusable engineering patterns: publish reference architectures for coding-agent workflows, secure prompt/data handling, repository access, PR/code-review workflows, and tool-use guardrails.
  • Cross-functional delivery: work with Enterprise Architecture, Cloud Engineering, Security, SRE, Developer Productivity, Legal/Risk, and business-unit engineering teams.
  • Broader AI application enablement: support product and engineering teams that need governed access to foundation models, enterprise tools, APIs, data sources, and agentic workflows beyond software development use cases.
  • Standards leadership: set engineering standards for agentic code development, BYOK usage, MCP design, gateway policy, and observability across AI engineering teams.
Phase 2: Extend to Enterprise AI Applications and Agent Platforms
  • Agent development platforms: build and operate reusable agent development platforms in Azure AI Foundry and AWS AgentCore, including secure templates, deployment pipelines, runtime patterns, testing/evaluation workflows, and operating standards.
  • Model-access foundation: provide governed model access for AI applications through approved gateways, policy enforcement, usage metering, tenant isolation, data classification controls, and cost-management patterns.
  • MCP-enabled enterprise integration: enable AI applications and agents to access approved enterprise systems through governed MCP patterns, including tool registration, access approval, logging, decommissioning, and trust-boundary enforcement.
  • Platform roadmap ownership: partner with Enterprise Architecture, Security, Cloud, Data, and application teams to evolve the platform from coding-agent enablement into a shared enterprise AI application foundation.
Required Qualifications
  • 8+ years of software, cloud, platform, DevOps, or infrastructure engineering experience, including 3+ years building enterprise cloud/platform services and 1 2+ years supporting AI/ML, LLM, or agentic systems.
  • Strong hands-on engineering skills with Python plus infrastructure-as-code experience using Terraform, Bicep, CloudFormation, or equivalent.
  • Experience deploying cloud-native services using containers, Kubernetes or serverless patterns, CI/CD, secrets management, private networking, and enterprise identity controls.
  • Hands-on experience with Azure AI Foundry, Azure API Management, Azure Monitor/Application Insights, Microsoft Entra ID, Key Vault, Private Link, Azure Policy, and related governance/security services.
  • Working knowledge of AWS Bedrock, AWS GovCloud, IAM, CloudWatch, networking, and secure cross-cloud or multi-cloud operating patterns.
  • Experience designing or operating LLM gateways, model-routing layers, API gateways, MCP servers/gateways, tool/function calling, RAG pipelines, or agent orchestration frameworks.
  • Experience administering or governing GitHub Copilot SaaS, including access management, policy configuration, licensing, usage reporting, and coordination with GitHub or Microsoft 365 platform administration teams.
  • Strong understanding of security architecture: least privilege, RBAC/ABAC, conditional access, encryption, egress controls, audit logging, data classification, and secrets management.
  • Ability to design controls for regulated environments, including NIST, CMMC, FedRAMP, FISMA, and CUI handling requirements.
  • Experience building observability for distributed systems, including logs, metrics, traces, dashboards, alerting, SLOs, incident response, and operational runbooks.
  • Excellent communication skills with the ability to explain architecture, risk, cost, compliance, and operational trade-offs to technical and non-technical stakeholders.
Preferred Qualifications
  • Experience with GitHub Copilot Enterprise, GitHub Copilot BYOK, Claude Code, Anthropic SDKs, OpenAI/Azure OpenAI APIs, AWS Bedrock model access, or comparable coding-agent platforms.
  • Experience with MCP, A2A, LangGraph, Semantic Kernel, LlamaIndex, LangChain, AgentCore, or related orchestration/tooling frameworks.
  • Experience building AI gateways with model allow-lists, content safety, prompt-injection defenses, usage metering, model routing, rate limits, circuit breakers, and cost attribution.
  • Experience with Microsoft Purview, Defender for Cloud, Sentinel/SIEM integration, policy-as-code, OpenTelemetry, Langfuse, Arize, LangSmith, or other AI observability/evaluation tools.
  • Experience supporting Defense Industrial Base, Federal, government cloud, GCC High, Azure Government, AWS GovCloud, or environments processing CUI.
  • Experience implementing chargeback/showback models for cloud, token, model, API, or developer-platform usage.
  • Experience with AI FinOps, cloud cost management, tokenomics, cost allocation, budget guardrails, consumption forecasting, anomaly detection, or showback/chargeback reporting for AI platforms.
  • Experience defining product/platform roadmaps, operating models, onboarding processes, SLOs, runbooks, and stakeholder communication for shared enterprise platforms.
Preferred Certifications
  • Microsoft Certified: Azure Solutions Architect Expert, Azure AI Engineer Associate, Azure Security Engineer Associate, or comparable Azure certification.
  • AWS Certified Solutions Architect, AWS Security Specialty, AWS Machine Learning, or comparable AWS certification.
  • GitHub Copilot, GitHub Actions, GitHub Advanced Security, or related GitHub platform certification or demonstrated equivalent experience.
  • FinOps Certified Practitioner, Terraform/Kubernetes certification, CISSP, CCSP, CompTIA Security+, or comparable cloud, platform, FinOps, or security certification.

Skills

AWS
Encryption
Machine Learning
Azure
CloudFormation
GitHub Actions
Kubernetes
LLM
Python
Terraform
Vault

Similar jobs