Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
Toronto, ON, United States
Posted
Yesterday
Machine LearningAzure.NET
Job Description
Join a growing financial technology organization where your engineering expertise will make a meaningful difference in how schools across North America manage their financial operations.
As a Staff Site Reliability Engineer, you'll have the opportunity to take real ownership of platform reliability, shape the future of infrastructure strategy, and help build resilient systems that thousands of educational institutions depend on every day. Working across Microsoft Azure, .NET, legacy applications, and modern cloud technologies, you'll tackle interesting technical challenges while driving improvements in observability, automation, incident management, and overall system performance.
You'll also have the freedom to explore emerging AI-driven solutions, collaborate with talented DevOps and engineering teams, and mentor others as the organization continues to grow and evolve. If you're someone who loves solving complex problems, enjoys staying close to the technology, and wants the opportunity to leave a lasting mark on both the platform and the engineering culture, this is an exciting opportunity to do exactly that.
Required Skills & Experience
As a Staff Site Reliability Engineer, you'll have the opportunity to take real ownership of platform reliability, shape the future of infrastructure strategy, and help build resilient systems that thousands of educational institutions depend on every day. Working across Microsoft Azure, .NET, legacy applications, and modern cloud technologies, you'll tackle interesting technical challenges while driving improvements in observability, automation, incident management, and overall system performance.
You'll also have the freedom to explore emerging AI-driven solutions, collaborate with talented DevOps and engineering teams, and mentor others as the organization continues to grow and evolve. If you're someone who loves solving complex problems, enjoys staying close to the technology, and wants the opportunity to leave a lasting mark on both the platform and the engineering culture, this is an exciting opportunity to do exactly that.
Required Skills & Experience
- 10+ years of experience in Site Reliability Engineering, Platform Engineering, or infrastructure operations, with proven expertise establishing reliability strategies, defining SLIs/SLOs and error budgets, leading production incident response, and implementing effective root cause analysis and postmortem practices.
- Strong hands-on experience with Microsoft Azure, .NET environments, legacy .NET Framework applications, and IIS, including the ability to improve availability, resilience, and operational performance across both modern and established production systems.
- Advanced expertise building observability and monitoring capabilities across metrics, logging, distributed tracing, and alerting, combined with strong scripting and automation skills to eliminate operational toil, improve incident detection, and develop maintainable production-grade tooling.
- Advanced experience in capacity planning, performance optimization, load testing, and proactive infrastructure scaling, with the ability to identify system bottlenecks and prevent reliability issues before they affect production services.
- Exposure to AI-driven operations and intelligent observability, including applying machine learning or AI-assisted tooling to anomaly detection, incident triage, operational analytics, and automated remediation workflows.
- Demonstrated technical leadership in coaching senior engineers, improving on-call practices, establishing reliability engineering standards, and influencing cross-functional architecture and operational decisions across DevOps and product engineering teams.
- Hands-On Engineering: 70%
- Team Collaboration & Cross-Functional Work: 30%
You will receive the following benefits:
Medical, Dental, and Vision Insurance
Vacation Time
Current Vacancy: Yes
Use of AI in Hiring: No
Applicants must be currently authorized to work in Canada on a full-time basis now and in the future.
Accommodation will be provided in all parts of the hiring process as required under Motion Recruitment's Employment Accommodation policy. Applicants need to make their needs known in advance.
#LI-AC1
Similar jobs
- TC
DevSecOps Engineer
NewTECHNEPTUNE CONSULTING INC
United States🇺🇸RemoteYesterdaySAFeAgileEngineering - TB
Site Reliability Engineer (SRE)
NewThe Brixton Group
Benton Harbor, MI🇺🇸RemoteYesterdayDockerDynamoDBNode.js+15Technology - ED
DevOps Engineer III
NewAuto ApplyEnable Dental
Austin, Texas🇺🇸Remote10 hours agoDockerAWSEncryption+17Technology - TH
DevOps Engineer
NewAuto ApplyTheIncLab
Colorado Springs, Colorado🇺🇸Hybrid5 hours agoDockerShellAWS+22Technology - SG
DevOps Engineer
NewAuto ApplyStafford Gray
Lansing, Michigan🇺🇸Hybrid10 hours agoShellAuditingAzure+7Technology - JG
Azure DevOps Engineer
NewJudge Group, Inc.
Berkeley Heights, NJ🇺🇸On-siteYesterdayEncryptionAgileAnsible+4Technology - JG
Sr DevOps Engineer
NewJudge Group, Inc.
Lone Tree, CO🇺🇸HybridYesterdayDockerMicroservicesAnsible+2Technology - ST
Senior Azure DevOps Engineer
NewStefanini
Fort Myers, FL🇺🇸RemoteYesterdayAzureGitHub ActionsTerraformTechnology - MR
Azure DevOps Engineer
NewMotion Recruitment Partners, LLC
Toronto, ON🇺🇸HybridYesterdayService MeshArgoCDAzure+6Technology - HP
MLOps certified - Cloud ML Ops/SRE
NewHPTech Inc.
Sunnyvale, CA🇺🇸On-siteYesterdayMLOpsKubernetesTechnology - SY
Site Reliability Engineer
NewSynkriom
United States🇺🇸HybridYesterdayRubyAWSAnsible+4Technology - SB
Cloud Infrastructure DevOps Engineer Remote Location
NewSierra Business Solution LLC
United States🇺🇸RemoteYesterdayDockerAWSSplunk+12Technology