Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
Toronto, ON, United States
Posted
Yesterday
SSOHelmJWTKubernetesLLMVault
Job Description
Position Title: Senior Platform Engineer - Nearshore
Location: Canada (Remote)
Type: Contract
- Build, deploy, and operate the shared AI platform services that application teams depend on LLM gateway and proxy layers, model routing and fallback, MCP servers and registries, agent runtimes, and the supporting APIs around them with telemetry, health signals, and cost attribution designed in rather than added later.
- Onboard and enable new model providers and model versions across environments, including routing rules, rate limits and quotas, model tiering and cost-aware selection, fallback behavior, and version deprecation paths, and make each provider's usage, latency, error, and spend profile visible and comparable.
- Design and operate the MCP layer that exposes enterprise systems as tools to agents server deployment, tool registration and discovery, connection and session handling, safe tool-permission boundaries, and tracing of tool calls, retries, and failures.
- Support the agent lifecycle on the platform publishing and versioning, registry and discovery, memory and session management, scaling and resilience and instrument agent steps, state transitions, and evaluation outcomes so behavior can be explained after the fact.
- Run platform services on Kubernetes using GitOps practices: Helm and Kustomize manifests, declarative delivery, autoscaling, resource tuning, health checks, and zero-downtime or progressive rollout.
- Implement authentication, authorization, and tenancy for the platform OAuth2/OIDC and SSO integration, JWT issuance and validation, role and group based access, virtual key and API key management, and secrets handling through a managed key vault.
- Apply guardrails and policy enforcement, including PII detection with masking or blocking, prompt and response filtering, content policy, audit logging, and retention rules for prompts, completions, and traces.
- Partner with engineering and developer teams to define observability standards for AI applications, LLM integrations, and agentic workflows, and make those standards the default path rather than an extra step.
- Implement and maintain telemetry patterns in Langfuse, including traces, spans, prompts, completions, feedback, evaluation metadata, latency, errors, token usage, and model/provider context.
- Instrument the platform itself with OpenTelemetry, metrics, and structured logs so gateway, MCP, and agent-runtime behavior is traceable end to end alongside application-level AI traces.
- Build dashboards, reports, and alerts that help teams understand AI reliability, performance, quality, evaluation outcomes, and production behavior.
- Design cost and usage observability across LLM vendors such as OpenAI, Anthropic, Google Gemini, and other providers, including attribution by application, team, user, model, workflow, and environment.
- Create showback or chargeback-ready metrics for token usage, request volume, model mix, latency, cache behavior, evaluation runs, and vendor spend, and feed what they show back into routing, tiering, and capacity decisions.
- Support AI evaluation practices by helping teams define test sets, golden datasets, scoring strategies, prompt and version comparisons, regression checks, and release readiness signals.
- Analyze execution patterns across agent tooling such as Claude Code, in-house agent frameworks, LangChain, LangGraph, CrewAI, and Google ADK, and turn what you find into platform fixes and guidance.
- Support retrieval and context-grounding infrastructure embedding models, vector and graph stores, ingestion and refresh pipelines, and retrieval quality measurement and tuning.
- Build the developer-facing self-service surface: onboarding automation, provisioning workflows, templates, and internal portals or CLIs that let teams get access, register agents and tools, and ship without manual tickets.
- Extend observability into the software delivery layer where relevant AI-assisted development telemetry, pipeline and pull-request lifecycle metrics, and DORA-style delivery signal
Similar jobs
- EG
Dev Ops Engineer
NewEliassen Group
Plano, TX🇺🇸$60 - $66/hrOn-siteYesterdayEngineering - DD
Kubernetes Engineer
NewDexian DISYS
Philadelphia, PA🇺🇸$55 - $60/hrOn-siteYesterdayKubernetesTechnology - RA
Sr Snowflake Admin - with DeVops & AWS-Docker experience -Remote Work
NewRapidIT, Inc
United States🇺🇸RemoteYesterdayDockerSQLAWS+5Technology - ET
Senior DevOps Cloud Engineer
NewEsvee Technologies Inc
Plano, TX🇺🇸HybridYesterdayAzureBashGit+8Technology - NI
Devops Engineer
NewNimbusAITech LLC
Columbus, OH🇺🇸On-siteYesterdaySAMLSSOBash+1Technology - GL
Senior Cloud DevOps Engineer
NewGXO Logistics
Fort Worth, Texas🇺🇸Hybrid1 hour agoDockerShellAWS+15Technology