Rightsline Inc - 5 emplois
Toronto, ON
Détails de l'emploi :
Rightsline is hiring a DevOps Manager, based in Toronto, to lead the team that owns our cloud platform, delivery pipelines, and production reliability across North America and Europe. You will manage and develop a team of five — a Senior DevOps Engineer, a DevOps Engineer, and two Senior Systems Engineers — and hold end-to-end accountability for the infrastructure our product runs on. You will report to the Senior DevOps Architect and partner closely with Engineering, QA, Security and Compliance, and Support.
Mission & ImpactRightsline runs a multi-region AWS platform: production in two regions serving customers in the US and the EU, twelve environments spanning developer sandbox through pre-production mirror to production, infrastructure defined in Terraform, and a templated CI/CD fleet that builds and deploys every component into each of them. The team owns our AWS accounts end to end, along with a mixed Linux and Windows server estate. The team that runs this platform is deeply capable technically. Your mission is to pair that capability with the operational discipline of a mature platform organization: a named owner for every system and environment, infrastructure changes that land predictably and completely across all the places they belong, documentation current enough to be trusted during an incident, and a team that holds itself accountable for follow-through.
You will be accountable for the platform as a product that the rest of the company depends on — its reliability, its cost, its security posture, and how quickly engineers can get work safely into production. Success is measured by production availability and incident recovery in both regions; by environment changes that complete the first time, with a record of what changed and why; by runbooks, diagrams, and environment inventories that are current; by disaster recovery and audit evidence produced as routine output rather than as a scramble; and by a team where ownership is unambiguous, people are growing, and commitments are met.
What You Will OwnTeam Leadership & AccountabilityManage, coach, and develop a team of four infrastructure engineers, including senior individual contributors — people with deep expertise you will lead through judgment and clarity rather than positional authority.
Establish unambiguous ownership: every system, environment, pipeline, and recurring operational duty has a named owner and a documented backup.
Run the performance cycle — goals, regular one-on-ones, written feedback, and development plans calibrated against our engineering level framework — and address both strong performance and gaps directly and fairly.
Hold the team to its commitments, and build the visibility that makes commitments legible: a single prioritized queue of work, current ticket hygiene, and clear closure on action items.
Recruit, onboard, and retain, building enough bench depth that no system depends on one person being available.
Own the change management process for infrastructure and environments end to end: proposal, peer review, scheduled execution, verification, and a durable record of what changed and why.
Define and hold the line on what “done” means for environment work — the change applied consistently across every environment and region it belongs in, code and state in agreement, monitoring updated, documentation updated, and the ticket closed.
Drive infrastructure work through version control and code review rather than console changes, and systematically close configuration drift between environments and regions.
Make documentation a deliverable: architecture diagrams, runbooks, escalation paths, and environment inventories that an on-call engineer can rely on at three in the morning.
Own incident practice — on-call rotation, severity definitions, blameless postmortems, and remediation items that are tracked to closure.
Own our AWS accounts outright — account structure and organization, IAM and access boundaries, service quotas, networking, billing and tagging, and the guardrails that keep each account consistent with the next.
Own everything running in those accounts across both regions — containerized services, load balancing, relational databases, search, messaging, and serverless components — including capacity, performance, patching, and security posture.
Own the server estate on both Linux and Windows: build standards, configuration management, patching cadence, hardening, and lifecycle, so neither platform becomes the one nobody wants to touch.
Own infrastructure as code in Terraform: module design, environment composition, state management, and the review standards that keep it maintainable as the fleet grows.
Own the CI/CD estate — a templated pipeline-per-component-per-environment fleet across twelve environments and two regions — along with the release promotion model and production approval gates.
Own observability and log management so that alerts are actionable, noise is low, and the data needed to investigate an issue is still there when the question gets asked.
Partner with Engineering to reduce deployment risk and shorten lead time to production, and with QA on environment availability, parity, and refresh.
Own availability and performance objectives for production in both regions, and the monitoring that demonstrates they are being met.
Run our recurring disaster recovery test cycle end to end — including evidence against our 15-minute RPO and 4-hour partial / 24-hour full RTO objectives — and drive every finding to closure before the next cycle.
Produce infrastructure evidence for SOC 2 and PCI DSS as a by-product of how the team already works: access reviews, change records, backup and restore verification, and vulnerability management.
Partner with Security on hardening, secrets management, remediation timelines, and audit readiness, and keep our EU data residency commitments intact.
Act as the escalation point for infrastructure issues affecting customers, communicating status credibly to engineering leadership and, when needed, to customers.
Own cloud spend and infrastructure tooling costs, with a clear view of cost per environment and a plan for where it should go.
Coordinate effectively with engineering colleagues in the United States and India, including planning work and releases across time zones.
7+ years in DevOps, SRE, or infrastructure engineering, including 2+ years directly managing a DevOps or infrastructure team, with hiring, performance, and development responsibility. Management experience specifically in this space is required — leading infrastructure engineers is a different job from leading application developers.
Experience leading senior engineers and architects — specialists who know their domains better than you do — and a track record of earning their trust.
A demonstrated record of introducing process rigor to a technically strong team without slowing it down: change management, documentation standards, and on-call and incident practice that the team genuinely adopted and sustained.
Comfort holding people accountable — setting expectations explicitly, following up consistently, and having the direct conversation early rather than late.