GuruLink - 129 Jobs
Toronto, ON
Job Details:
Location: Vancouver, British Columbia
This role suits someone systems-minded, pragmatic, security-aware, and averse to repetitive manual work. You'll step in when the platform is on fire, but your focus is prevention over reaction.
You'll own real infrastructure at real scale, with authority over the IaC, GKE, and CI/CD architecture every team depends on, and you'll help shape how an AI-first engineering org ships software safely.
This is a Hybrid role:
• 3 days a week
• Office located downtown Vancouver
What You'll Do:
• Build and extend a Pulumi component library so teams get correct, disaster-recovery-ready infrastructure by default instead of hand-rolling stacks across multiple GCP environments.
• Run GKE as a product, handling cluster and node pool upgrades, in-cluster services, workload right-sizing, and policy enforcement.
• Own the core building blocks of CI/CD workflows, cut CI time and flake rate fleet-wide, and standardize the golden path for shipping services so pipelines stay out of engineers' way.
• Improve the internal developer platform surfaces that show engineers what's happening under the hood in their own systems.
• Own the health of the observability stack, with consistent instrumentation, actionable alerts, and less on-call noise, so production stays easy to understand.
• Drive GCP and LLM cost reduction through attribution, right-sizing, commitment planning, and removal of orphaned resources.
• Extend software supply chain controls, own infrastructure identity and access, and build SOC 2 and other compliance requirements into automated, evidenced controls so the secure path is also the fast path and audits are a byproduct of how the team builds.
Tech Stack:
GCP, Kubernetes, Pulumi (IaC), Postgres, Temporal
Must Have Skills:
• Deep, hands-on experience in DevOps, SRE, platform, or infrastructure engineering, operating production systems that paying customers depend on, including on-call ownership.
• Deep, practical Kubernetes knowledge.
• Infrastructure as code treated as software (Pulumi or Terraform).
• Strong GCP or AWS fundamentals.
• Background in systems engineering and writing software.
• CI/CD owned as a product: you've built and maintained pipelines used by other teams. GitHub Actions or Buildkite preferred; GitLab CI or CircleCI experience transfers.
• Hands-on experience with security tooling as an engineering practice.
• Compliance engineering for SOC 2 or enterprise partner security programs, implemented as automated controls; threat modeling and related practices.
• Ownership under ambiguity: you scope a vague problem, ship it incrementally, and tell people clearly what changed and why.
• Fluency with AI coding tools and experience working on AI-first teams.
Nice to Have Skills:
• Ownership of platform migrations and decommissions.
• Built an internal developer platform or self-service tooling that other engineers adopted.
• Comfort with LLM infrastructure (gateways, token cost attribution, rate limiting, provider failover)