Titre du poste ou emplacement

DevOps Engineer, Agentic Operation & Live Games

Big Viking Games - 7 emplois

Toronto, ON

Publié il y a 28 jours

Détails de l'emploi :

Temps plein
Expérimenté

About Big Viking Games

Big Viking Games is a Canadian gaming company focused on building, operating, and growing long-standing online game communities. Our games have entertained players for years, supported by loyal audiences, live operations, evolving content systems, product innovation, and deep player-driven economies.

Our flagship titles, YoWorld and FishWorld, have served millions of players over their lifetime. These are enduring live-service virtual worlds with rich in-game economies, virtual goods, social interaction, and long-term player engagement at their core.

We are entering a new phase of modernization and growth, with a focus on stronger infrastructure, better automation, practical AI adoption, improved reliability, stronger security practices, and scalable systems that help our games and teams perform at a higher level.

About the Role

Big Viking Games is hiring a Senior DevOps Engineer to modernize the infrastructure behind our live-service games — and to do it by building automation that thinks, not just scripts that run.

Our games run on mature tech stacks with large data volumes and player counts — real systems, with real players on them, where the constraints are genuine and the consequences are visible. There is meaningful room to automate how they are operated, and we want someone who builds that: tooling that takes routine work off people's hands and runs safely against production without compromising uptime or data integrity.

This is a hands-on senior role for an engineer who is fluent in modern cloud infrastructure — AWS, containers, IaC, CI/CD, observability, production reliability — and who has started using agentic coding tools and tool-calling systems to do infrastructure work that previously required a person. If you have wired a coding agent into your CI/CD, built an MCP server so tooling could act on your infrastructure safely, or replaced a manual runbook with something that diagnoses and remediates on its own, this role is aimed at you.

The defining trait is self-direction, and it works in two directions. Given a backlog, you improve on how the work gets done — bringing agentic approaches to problems that were scoped as manual ones, and solving the underlying issue rather than the individual ticket. Left to your own judgement, you find the repetitive work nobody has flagged, decide what's worth automating, and build it. In both cases you own the guardrails that make automation safe to run against production.

This is a hybrid role based in Toronto, with an expectation of working in office three days per week. Live-service games require operational awareness outside regular business hours, including periodic on-call and incident response availability.

Requirements

What You'll Do

Build agentic automation for infrastructure work

  • Design, build, and operate agentic tooling that performs real DevOps work — diagnosis, remediation, provisioning, routine maintenance — rather than only summarizing or suggesting.
  • Build and maintain the integration layer that lets tooling act safely on our systems: MCP servers, API integrations, webhook-driven orchestration, serverless functions, and the permissions and audit trails around them.
  • Convert manual runbooks, SOPs, and recurring operational chores into automation that runs unattended, with sensible escalation when it shouldn't proceed alone.
  • Establish the guardrails that make automated action against production defensible: least-privilege scopes, dry-run and approval paths for destructive operations, logging of what automation did and why, and clear rollback.
  • Use agentic coding tools to accelerate your own infrastructure work — IaC authoring, migration scripting, log and incident analysis, documentation — and improve how the wider engineering team does the same.

Own and modernize live production infrastructure

  • Monitor, maintain, and improve cloud infrastructure across AWS, Netlify, Vercel, and related platforms supporting our live games, data systems, internal tools, and operational workflows.
  • Drive infrastructure modernization while maintaining uptime for live games with active player communities — every improvement ships while the plane is flying.
  • Implement and maintain Infrastructure as Code using Terraform, CloudFormation, CDK, or similar.
  • Improve CI/CD pipelines, release workflows, and deployment reliability so teams can ship safely and frequently.
  • Operate containerized workloads and GitOps-based deployment.

Reliability, observability, and security

  • Rebuild observability across the stack: logging, metrics, alerting, dashboards, and operational visibility — with particular attention to early detection of silent failures. Our previous monitoring stack is no longer functional, so this is a build, not a handover.
  • Maintain and monitor data pipelines between game source databases (MariaDB), the Snowflake data warehouse, and downstream analytics and reporting systems — ensuring pipeline health, freshness, and alerting when data stops flowing.
  • Own secrets and credential lifecycle management across platforms — API key rotation, access controls, environment variable governance, and least-privilege practices.
  • Support incident response, root cause analysis, remediation planning, and post-incident improvements — and automate the parts of that loop that repeat.
  • Help manage cloud spend, infrastructure usage, resource tagging, and environment efficiency.
  • Create documentation and runbooks that are executable wherever possible rather than purely descriptive.
What You Bring

Experience

  • 7+ years in DevOps, infrastructure engineering, cloud engineering, site reliability engineering, or a similar role.
  • 1+ year building with agentic coding tools and tool-calling systems to automate, scale, or secure infrastructure and DevOps work. We mean systems that take action — agents wired into pipelines, MCP servers or tool integrations you built, automated diagnosis or remediation you put into production. Using an AI assistant to write code faster is useful but is not what this requirement is asking for.

Core technical

  • Strong hands-on experience with AWS or similar cloud platforms.
  • Experience designing, maintaining, and improving production infrastructure — including comfort with legacy systems that predate modern cloud-native patterns.
  • Proficiency with Infrastructure as Code tools such as Terraform, CloudFormation, CDK, Pulumi, or similar, including managing shared state across a team.
  • Experience with containerized applications (Docker) and container orchestration — Kubernetes, ECS, EKS, or similar — including debugging real cluster problems in production.
  • Experience with GitOps and declarative deployment (ArgoCD, Flux, or equivalent).
  • Experience with CI/CD tooling, version control, deployment automation, and modern release workflows.
  • Strong understanding of Linux systems, networking, cloud security, monitoring, logging, and operational troubleshooting.
  • Experience establishing observability — choosing and deploying the stack, deciding what warrants an alert, and tuning alert noise down — not only consuming dashboards someone else built.
  • Experience supporting live production systems where uptime, reliability, and performance matter, and that cannot tolerate extended downtime.
  • Experience with relational databases (MariaDB, MySQL, Postgres) — replication, backup and verified restore, and schema change against systems that stay online.
  • Security-aware mindset with practical experience in secrets management, credential rotation, access control, vulnerability reduction, and least-privilege practices.

How you work

  • Self-starting. You identify the work rather than wait for it to be assigned, and you can justify why you picked one problem over another.
  • Proactive about toil. You notice repetitive work and treat it as a defect to be engineered away, not a cost of doing business.
  • Inventive but pragmatic. You reach for novel approaches where they genuinely help and recognize when a cron job and forty lines of Python is the better answer.
  • Strong problem-solving skills and the ability to investigate complex infrastructure or production issues, including silent failures and data pipeline outages.
  • Comfort working directly with engineers to improve build, deploy, and operational workflows.
  • Practical ownership mindset with the ability to prioritize, execute, and close loops.
Nice to Have
  • Experience supporting live games, virtual worlds, multiplayer systems, or other real-time online products.
  • Experience with GitHub Actions specifically, or similar CI/CD platforms.
  • Experience with Datadog, Grafana, Prometheus, CloudWatch, ELK, OpenTelemetry, or similar observability tooling.
  • Experience with Redis, Memcached, queues, workers, or event-driven systems.
  • Experience with Snowflake, data warehouse connectivity, ETL monitoring, or data pipeline reliability.
  • Experience with serverless platforms (Netlify Functions, Vercel, AWS Lambda) and multi-platform hosting environments.
  • Experience with disaster recovery, backup strategies, incident management, load testing, and performance tuning.
  • Experience improving cloud cost management, tagging, resource optimization, or infrastructure governance.
  • Experience evaluating where automation should not be applied — cases where you deliberately kept a human in the loop, and why.
  • Experience working in small, high-leverage engineering teams where infrastructure ownership is broad and hands-on.
Ideal Candidate Profile

The ideal candidate is a senior infrastructure engineer who keeps live systems stable while systematically removing the manual work involved in keeping them that way.

They are not tool-collectors. They understand uptime, production risk, cloud cost, security, release quality, and operational discipline — and they treat automation as an engineering discipline with its own failure modes, not as a way to look modern. They know that automation acting on production needs guardrails, observability, and a rollback path, and they build those first.

They work independently. Given a mature live game and a broad remit, they can identify what matters, sequence it sensibly, and improve systems incrementally without disrupting what is already working. They are comfortable being the person who decides what gets automated next.

This role suits someone who wants deep ownership of production infrastructure and a genuine mandate to reinvent how a live-service game is operated.

Benefits

Compensation

Compensation range: $140,000 to $180,000 CAD, determined based on experience, technical depth, infrastructure ownership breadth, agentic automation experience, and overall fit.

Benefits
  • Group Retirement Savings Plan matching and participation.
  • Comprehensive benefits package, including health, dental, and vision coverage.
  • Health and Wellness spending account.
  • Generous time off policies.
  • Opportunity to support long-running live-service games with established player communities.
  • Deep ownership of cloud modernization, DevOps automation, security improvement, and agentic infrastructure workflows.
  • A high-impact role with meaningful ownership over reliability, performance, and engineering operations.
Accessibility and Accommodation

Big Viking Games is committed to creating an inclusive and accessible environment for all candidates. We welcome applications from individuals of all abilities and will provide accommodations throughout the hiring process as needed.

If you require accommodation during the hiring process, please contact [email protected] so we can work with you to support your needs.

Partager un emploi :