Haystack
← Back to Jobs
Technology
DT

Senior AI Platform Engineer - Nearshore

Diligente TechnologiesToronto, ON🇺🇸United StatesPosted Oct 1, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Toronto, ON, United States
Posted
Yesterday
SSOHelmJWTKubernetesLLMVault

Job Description

Position Title: Senior Platform Engineer - Nearshore

Location: Canada (Remote)

Type: Contract

  • Build, deploy, and operate the shared AI platform services that application teams depend on LLM gateway and proxy layers, model routing and fallback, MCP servers and registries, agent runtimes, and the supporting APIs around them with telemetry, health signals, and cost attribution designed in rather than added later.
  • Onboard and enable new model providers and model versions across environments, including routing rules, rate limits and quotas, model tiering and cost-aware selection, fallback behavior, and version deprecation paths, and make each provider's usage, latency, error, and spend profile visible and comparable.
  • Design and operate the MCP layer that exposes enterprise systems as tools to agents server deployment, tool registration and discovery, connection and session handling, safe tool-permission boundaries, and tracing of tool calls, retries, and failures.
  • Support the agent lifecycle on the platform publishing and versioning, registry and discovery, memory and session management, scaling and resilience and instrument agent steps, state transitions, and evaluation outcomes so behavior can be explained after the fact.
  • Run platform services on Kubernetes using GitOps practices: Helm and Kustomize manifests, declarative delivery, autoscaling, resource tuning, health checks, and zero-downtime or progressive rollout.
  • Implement authentication, authorization, and tenancy for the platform OAuth2/OIDC and SSO integration, JWT issuance and validation, role and group based access, virtual key and API key management, and secrets handling through a managed key vault.
  • Apply guardrails and policy enforcement, including PII detection with masking or blocking, prompt and response filtering, content policy, audit logging, and retention rules for prompts, completions, and traces.
  • Partner with engineering and developer teams to define observability standards for AI applications, LLM integrations, and agentic workflows, and make those standards the default path rather than an extra step.
  • Implement and maintain telemetry patterns in Langfuse, including traces, spans, prompts, completions, feedback, evaluation metadata, latency, errors, token usage, and model/provider context.
  • Instrument the platform itself with OpenTelemetry, metrics, and structured logs so gateway, MCP, and agent-runtime behavior is traceable end to end alongside application-level AI traces.
  • Build dashboards, reports, and alerts that help teams understand AI reliability, performance, quality, evaluation outcomes, and production behavior.
  • Design cost and usage observability across LLM vendors such as OpenAI, Anthropic, Google Gemini, and other providers, including attribution by application, team, user, model, workflow, and environment.
  • Create showback or chargeback-ready metrics for token usage, request volume, model mix, latency, cache behavior, evaluation runs, and vendor spend, and feed what they show back into routing, tiering, and capacity decisions.
  • Support AI evaluation practices by helping teams define test sets, golden datasets, scoring strategies, prompt and version comparisons, regression checks, and release readiness signals.
  • Analyze execution patterns across agent tooling such as Claude Code, in-house agent frameworks, LangChain, LangGraph, CrewAI, and Google ADK, and turn what you find into platform fixes and guidance.
  • Support retrieval and context-grounding infrastructure embedding models, vector and graph stores, ingestion and refresh pipelines, and retrieval quality measurement and tuning.
  • Build the developer-facing self-service surface: onboarding automation, provisioning workflows, templates, and internal portals or CLIs that let teams get access, register agents and tools, and ship without manual tickets.
  • Extend observability into the software delivery layer where relevant AI-assisted development telemetry, pipeline and pull-request lifecycle metrics, and DORA-style delivery signal

Similar jobs