HireFT
Browse JobsHow it worksPricingAboutSuccess Stories
    Back to jobs
    SE

    ServiceNow

    Software

    Manager, Network Reliability and Resiliency

    Toronto, CanadaHybridFull-time5+ yrs experiencePosted 2w ago
    All ServiceNow jobs

    Job description

    Due to Government of Canada regulatory requirements, this position requires the successful completion of a Government of Canada Reliability Status screening as a condition of employment. The screening process requires 5 years of verifiable background history. This includes identity verification, education verification, a criminal record check, and a credit check. Candidates must be eligible to obtain and maintain Reliability Status, which generally requires Canadian citizenship or Canadian permanent resident status.  Employment is contingent upon successful completion and maintenance of the required screening.

    What you get to do in this role:

    We are seeking a Manager, Network Reliability and Resiliency to lead a team responsible for the reliability and day-to-day operation of production network services supporting ServiceNow's cloud platform. This is a technical people-manager role. You will develop engineers and manage team priorities while staying actively engaged in complex troubleshooting, high-severity incidents, customer escalations, operational readiness, and reliability improvement.

    You will apply SRE principles to network operations by using service indicators and objectives, error-budget thinking, observability, post-incident learning, and automation to improve availability, reduce operational toil, and make execution safer and more consistent. While this is not an individual contributor role, you must have the technical depth and judgment to guide investigations, challenge assumptions, make risk-based decisions, and help the team reach durable solutions.

    Lead and develop the team

    • Manage, coach, and develop network reliability engineers through clear goals, regular feedback, performance reviews, and career development.
    • Set priorities and ownership for operational work, reliability initiatives, technical debt, and project commitments.
    • Build sustainable on-call and escalation practices and promote calm, accountable execution during high-pressure events.
    • Hire and onboard new team members and ensure they gain the technical context, operating practices, and support needed to succeed.

    Provide technical and incident leadership

    • Actively engage in complex production troubleshooting and customer-impacting escalations by reviewing evidence, guiding technical hypotheses, identifying risk, and coordinating the right subject-matter experts.
    • Lead or support major incident response, including mitigation decisions, stakeholder communication, escalation management, and restoration of service.
    • Ensure post-incident reviews identify contributing factors and result in clear, prioritized, and completed preventive actions.
    • Review high-risk changes and operational plans for technical soundness, rollback readiness, monitoring coverage, and customer impact.

    Improve reliability through SRE practices

    • Partner with engineering and service owners to define and use meaningful SLIs and SLOs for network services.
    • Use error budgets, incident trends, capacity signals, and operational data to balance service reliability, delivery pace, and risk.
    • Improve observability, alert quality, dashboards, runbooks, and operational readiness so the team can detect and resolve issues efficiently.
    • Track practical reliability outcomes such as availability, recurring incidents, change success, alert effectiveness, and time to detect and recover.

    Embed automation in daily operations

    • Create a strong automation mindset across the team and identify repetitive, error-prone, or slow operational activities that should be eliminated or automated.
    • Prioritize automation that improves change safety, validation, triage, remediation, reporting, and operational consistency.
    • Work with engineering and automation partners to move useful tools and workflows into production with clear ownership, documentation, monitoring, and support models.
    • Measure whether automation reduces toil and operational risk rather than treating automation delivery alone as the outcome.

    Partner across the organization

    • Collaborate with network engineering, SRE, security, platform, data center, customer support, and other partner teams to resolve issues and improve service reliability.
    • Represent the team's technical assessment, customer impact, risks, dependencies, and recovery plan clearly to technical and business stakeholders.
    • Ensure new technologies, services, and automations meet operational acceptance criteria before the team assumes production ownership.
    • Improve incident, change, problem-management, and escalation processes based on operational evidence and team feedback.

    To be successful in this role you have:

    • Five or more years of relevant experience in network engineering, network reliability, cloud infrastructure, SRE, or large-scale production operations.
    • Experience managing or formally leading engineers, including prioritization, coaching, performance feedback, and delivery accountability.
    • Sufficient hands-on technical background to guide production troubleshooting across Linux-based systems and network services. You can interpret logs, metrics, alerts, and packet-level evidence and make sound operational decisions.
    • Working knowledge of networking concepts and technologies such as TCP/IP, routing, DNS, load balancing or ADCs, firewalls, cloud networking, and network observability. Deep expertise in every area is not required.
    • Experience leading or coordinating significant incidents and customer-impacting escalations in an always-on service environment.
    • Working knowledge of SRE practices, including SLIs, SLOs, error budgets, monitoring and alerting, incident management, and post-incident improvement.
    • An automation mindset and experience using scripting, workflow automation, or engineering partnerships to reduce manual operational work and improve consistency.
    • Experience working with geographically distributed teams and cross-functional partners in software, platform, infrastructure, or cloud services.
    • Strong written and verbal communication skills, sound judgment under pressure, and consistent attention to detail.
    • Experience using or evaluating AI-assisted tools to improve analysis, decision-making, automation, or team workflows, with appropriate attention to accuracy, security, and operational risk.

    Preferred qualifications

    • Experience operating networking for a global SaaS, large enterprise, cloud provider, or similarly complex production environment.
    • Familiarity with BGP or OSPF, data center fabrics, load balancers or ADCs, DDoS protection, VPNs, firewalls, or public-cloud networking.
    • Experience improving observability, change safety, capacity management, or operational readiness for production services.
    • Experience with IT service management practices, including incident, change, and problem management.
    • Relevant certifications such as CCNA, CCNP, Azure/AWS/GCP related

    Work Personas

    We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

    Equal Opportunity Employer

    ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity,  veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.

    Accommodations

    We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact [email protected] for assistance.

    Export Control Regulations

    For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.

    From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.

    Job details are sourced from the employer's original posting.

    Open job posting
    SE

    About the company

    ServiceNow

    Moveworks: Moveworks is the Agentic AI Assistant platform that empowers the entire workforce. Our platform enables employees to converse with all of their business systems through natural language to quickly find answers and automate tasks. Powered by the world's most advanced LLMs, our proprietary models, and a sophisticated Agentic AI platform, we're transforming how work gets done by allowing AI to take initiative, streamline complex workflows, and continuously learn and adapt. Moveworks is trusted by over 5.5 million employees at more than 350 of the world’s largest companies, including 10% of the Fortune 500, to automate everyday tasks and streamline business operations. Recognized on the Forbes Cloud 100 and AI 50 lists, Moveworks was also named one of Fast Company’s 2025 Most Innovative Companies and Inc’s Best in Business, in the Best in Innovation category. Moveworks was also recognized at Microsoft’s 2025 Partner of the Year and in 2024, received the AI Breakthrough Award. In December 2025, Moveworks was acquired by ServiceNow, marking a pivotal milestone in our journey to create a single front door to work for all business systems. By combining ServiceNow’s leading workflow automation with Moveworks’ Reasoning Engine and natural language capabilities, we deliver the AI platform for every person and every workflow. Built to go beyond basic summaries to deliver meaningful business impact. Together, our AI acts across enterprise systems to turn conversations into completed work. By joining our team, you’ll be at the forefront of the AI transformation, backed by the global scale of ServiceNow and the agility of a high-growth company. We are looking for world-class talent to help us extend agentic AI to every employee across every corner of the business. Come join us! ServiceNow: It all started in sunny San Diego, California in 2004 when a visionary engineer, Fred Luddy, saw the potential to transform how we work. Fast forward to today — ServiceNow stands as a global market leader, bringing innovative AI-enhanced technology to over 8,100 customers, including 85% of the Fortune 500®. Our intelligent cloud-based platform seamlessly connects people, systems, and processes to empower organizations to find smarter, faster, and better ways to work. But this is just the beginning of our journey. Join us as we pursue our purpose to make the world work better for everyone.

    View all ServiceNow jobs
    Industry
    Software
    Founded
    2004
    Open roles
    626

    Interested in this role?

    Apply with HireFT

    Free to start — no card required.

    Your fit

    How well do you match?

    Sign in to see how your résumé lines up with this role.