GuruLink - 124 Jobs
Toronto, ON
Job Details:
Location: REMOTE / Toronto, Ontario
This job allows you to work remotely.
About the role
Our client is a well-established financial technology company processing large-scale, mission-critical transaction volume across cloud and on-premises environments. They are hiring a Manager, Site Reliability Engineering to lead a team responsible for the availability, performance, and resiliency of their core platforms - shaping SRE maturity and driving measurable improvements in system health as the organization scales its reliability practice.
What you'll do
- Lead and manage a team of SRE engineers supporting the reliability, availability, and performance of business-critical applications and platforms
- Implement and operationalize SRE practices: SLIs, SLOs, error budgets, incident response, and post-incident reviews
- Oversee production operations including on-call rotations, incident management, escalations, and problem management
- Partner with Development and DevOps teams to embed reliability principles into system design and delivery pipelines
- Drive observability strategy across monitoring, logging, and alerting standards
- Reduce operational toil through automation and self-healing system design
- Lead capacity planning, resiliency testing, and disaster recovery readiness
- Recruit, mentor, and develop SRE talent, fostering a culture of continuous improvement
What's offered
- Comprehensive total rewards: performance-based bonus, flexible benefits from day one, HSA/PSA choice
- Retirement support: profit-sharing with company match, defined contribution pension
- Growth opportunities: Coursera access, mentorship, internal mobility
- Hybrid flexibility and generous time-off programs
Must Have Skills:
- 8+ years in senior technical roles supporting distributed systems
- 3+ years leading and developing technical teams
- Strong grounding in SRE principles: SLOs, SLIs, error budgets
- Hands-on experience with cloud platforms (Azure preferred), Kubernetes, infrastructure as code, and automation
- Experience with enterprise observability tooling (Dynatrace, Datadog, New Relic, or AppDynamics)
- Strong scripting/programming skills for automation and operational efficiency
- Solid Linux systems and production infrastructure experience
Nice to Have Skills:
- Experience in payment processing, fintech, or PCI-regulated environments
- Familiarity with change management and compliance frameworks
- Understanding of SDLC and modern delivery practices
- Bachelor's degree in Computer Science, Software Engineering, or equivalent experience