Manager, Network Reliability and Resiliency
ServiceNow · Toronto, Ontario, Canada
- Senior
- Full-time
- Posted 2026-09-15
- Confirmed live on 25 September 2026
Job description
Description du poste
Due to Government of Canada regulatory requirements, this position requires the successful completion of a Government of Canada Reliability Status screening as a condition of employment. The screening process requires 5 years of verifiable background history. This includes identity verification, education verification, a criminal record check, and a credit check. Candidates must be eligible to obtain and maintain Reliability Status, which generally requires Canadian citizenship or Canadian permanent resident status. Employment is contingent upon successful completion and maintenance of the required screening.
What you get to do in this role:
We are seeking a Manager, Network Reliability and Resiliency to lead a team responsible for the reliability and day-to-day operation of production network services supporting ServiceNow's cloud platform. This is a technical people-manager role. You will develop engineers and manage team priorities while staying actively engaged in complex troubleshooting, high-severity incidents, customer escalations, operational readiness, and reliability improvement.
You will apply SRE principles to network operations by using service indicators and objectives, error-budget thinking, observability, post-incident learning, and automation to improve availability, reduce operational toil, and make execution safer and more consistent. While this is not an individual contributor role, you must have the technical depth and judgment to guide investigations, challenge assumptions, make risk-based decisions, and help the team reach durable solutions.
Lead and develop the team
• Manage, coach, and develop network reliability engineers through clear goals, regular feedback, performance reviews, and career development.
• Set priorities and ownership for operational work, reliability initiatives, technical debt, and project commitments.
• Build sustainable on-call and escalation practices and promote calm, accountable execution during high-pressure events.
• Hire and onboard new team members and ensure they gain the technical context, operating practices, and support needed to succeed.
Provide technical and incident leadership
• Actively engage in complex production troubleshooting and customer-impacting escalations by reviewing evidence, guiding technical hypotheses, identifying risk, and coordinating the right subject-matter experts.
• Lead or support major incident response, including mitigation decisions, stakeholder communication, escalation management, and restoration of service.
• Ensure post-incident reviews identify contributing factors and result in clear, prioritized, and completed preventive actions.
• Review high-risk changes and operational plans for technical soundness, rollback readiness, monitoring coverage, and customer impact.
Improve reliability through SRE practices
• Partner with engineering and service owners to define and use meaningful SLIs and SLOs for network services.
• Use error budgets, incident trends, capacity signals, and operational data to balance service reliability, delivery pace, and risk.
• Improve observability, alert quality, dashboards, runbooks, and operational readiness so the team can detect and resolve issues efficiently.
• Track practical reliability outcomes such as availability, recurring incidents, change success, alert effectiveness, and time to detect and recover.
Embed automation in daily operations
• Create a strong automation mindset across the team and identify repetitive, error-prone, or slow operational activities that should be eliminated or automated.
• Prioritize automation that improves change safety, validation, triage, remediation, reporting, and operational consistency.
• Work with engineering and automation partners to move useful tools and workflows into production with clear ownership, documentation, monitoring, and support models.
• Measure whether automation reduces toil and operational risk rather than treating automation delivery alone as the outcome.
Partner across the organization
• Collaborate with network engineering, SRE, security, platform, data center, customer support, and other partner teams to resolve issues and improve service reliability.
• Represent the team's technical assessment, customer impact, risks, dependencies, and recovery plan clearly to technical and business stakeholders.
• Ensure new technologies, services, and automations meet operational acceptance criteria before the team assumes production ownership.
• Improve incident, change, problem-management, and escalation processes based on operational evidence and team feedback.
Qualifications
To be successful in this role you have:
• Five or more years of relevant experience in network engineering, network reliability, cloud infrastructure, SRE, or large-scale production operations.
• Experience managing or formally leading engineers, including prioritization, coaching, performance feedback, and delivery account
Interview problems reported for ServiceNow
Reported by candidates and public write-ups, not by ServiceNow. Practise each one here:
- Longest Substring Without Repeating Characters — Medium
- Number of Islands — Medium
- Container With Most Water — Medium
- Longest Repeating Character Replacement — Medium
- Valid Parentheses — Easy
- Two Sum — Easy
- Merge Two Sorted Lists — Easy
- Longest Palindromic Substring — Medium
- Coin Change — Medium
- Maximum Subarray — Medium
- Set Matrix Zeroes — Medium
- Group Anagrams — Medium
- Best Time to Buy and Sell Stock — Easy
- Top K Frequent Elements — Medium
- Product of Array Except Self — Medium
- Reverse Linked List — Easy
- Combination Sum — Medium
- Pacific Atlantic Water Flow — Medium
- House Robber II — Medium
- Longest Common Subsequence — Medium
- Merge Intervals — Medium
- Maximum Product Subarray — Medium
More at ServiceNow
- Senior Analyst, US International Tax · Santa Clara, CALIFORNIA, United States
- Advisory Solution Consultant - State Government · Sydney, NSW, Australia
- Senior in-Market Engineer · Tokyo, , Japan
- Principal Software Engineer · San Diego, CALIFORNIA, United States
- Staff Software Engineer · Santa Clara, CALIFORNIA, United States
- Staff Software Engineer · San Diego, CALIFORNIA, United States
- Principal Software Engineer · Santa Clara, CALIFORNIA, United States
- Pricing Operations Senior Analyst · Remote; San Francisco de Heredia, Heredia, Costa Rica