Collabera Canada Inc - 2 Jobs
Toronto, ON
Job Details:
Role: Senior Azure Databricks / PySpark Developer
Contract: 12 months (with possible extension)
Weekly Hours: 37.5
Location: Toronto, ON
Hybrid: Core on-site 5 days/month - every Monday, one Friday a month
Role/Scope
- Modernizing the data platform, which is Edge platform
- This team is a regulatory reporting engine responsible for risk analysis reporting. This team cares about accuracy and auditability. They look at the regulatory data for regulatory bodies like OSFI.
- Their Book of Record sits on their legacy platform. They extract data here then transform to Edge for OSFI reporting. To run risk calculations, they must pull for last 2 years from their legacy platform
- The business is more active in asking for more types of reporting and the view is that technology is advanced enough now, to run more ad hoc reporting. Additionally, Bank of Canada can/is asking for different types of reports and they want to know what inputs were pulled for finalized generated reports.
- The program that the new Director of App Management is running with is the Economic Credit loss program, a two year program being built from ground zero. The program is related to Data Lineage, Auditing, and Traceability.
- This senior resource should know how modern ingestion works, how to clean data and populate data in Databricks
- Expectation for resource: Take initiative, figure-it-out mentality
Day to Day
- Data Pipeline Development
- Design and develop scalable data pipelines using Azure Databricks, Python, and PySpark.
- Implement ETL/ELT workflows for structured, semi-structured, and unstructured data.
- Develop and maintain data processing workflows using Delta Lake architecture.
- Ensure efficient data ingestion, transformation, and loading processes.
- Data Management & Optimisation
- Optimize performance of data pipelines and processing workloads.
- Ensure data quality, reliability, and consistency across data platforms.
- Write and optimize complex SQL queries for data processing and validation.
- Implement data governance and access controls.
- Cloud & Data Platform Integration
- Work with Azure Data Lake for storing and managing large datasets.
- Manage data cataloging and governance using Unity Catalog.
- Integrate data workflows with CI/CD pipelines for automated deployments.
- Support scalable and secure cloud-based data architectures.
- Collaboration & Migration Projects
- Work closely with cross-functional teams to deliver data-driven solutions.
- Participate in data migration and modernization initiatives.
- Support troubleshooting and resolution of data pipeline issues.
- Document data architecture, pipelines, and processes.
- Guide/train intermediate resource
Must Haves
- 10 years of IT experience and at least 5+ years of hands-on experience in Databricks (including RBAC, permission, catalog management, governance)
- Design and develop scalable data pipelines using Azure Databricks, SCALA, Python, PySpark, and Delta Lake.
- Have a good understanding of RDBMS and proficiency in writing complex SQL (1st Preference – Oracle / PLSQL).
- Implement ETL/ELT workflows for structured, semi-structured, and unstructured data.
- Optimize performance of data processing and ensure data quality standards.
- Work with Azure Data Lake, Unity Catalog, and integrate with CI/CD pipelines. (Unity Catalog preferred, if not Hive Meta Store or Metadata Management)
- Strong knowledge of Databricks technologies, specifically Delta Lake and Apache Spark, with a focus on implementing data engineering best practices.
- Working experience on a migration project.
Nice to Haves
- Experience with orchestration tools like Airflow or Azure Data Factory
Thanks,