Director, Site Reliability Engineering & Service Enablement
ServiceNow · Santa Clara, CALIFORNIA, United States
- Senior
- Full-time
- $221,200 – $387,100
- Posted 2026-09-16
- Confirmed live on 25 September 2026
Job description
Job Description
Team:
Our Site Reliability Engineering (SRE) team consists of highly skilled engineers responsible for maintaining and enhancing the reliability, scalability, and performance of the ServiceNow infrastructure. Our SRE’s are empowered to resolve technical issues across the entire technology stack, from hardware to applications. Additionally, they work to improve the platform's operability, aiming to reduce the number of incidents and minimize Mean Time to Recovery (MTTR). To achieve this, the team combines software development, networking, database, and systems engineering skills to tackle complex problems, striving to maintain our platform operating for our customers.
Role:
We are looking for a Director of Site Reliability Engineering to lead the next phase of our reliability transformation as ServiceNow modernizes toward a cloud-agnostic, cloud-ready production platform.
This leader will own key elements of the SRE operating model across Reliability Engineering, Service Enablement, Service Registry, SLI/SLO standards, reliability governance, automation, AI-enabled operations, and production readiness. The role will lead a global engineering organization and partner across Product Engineering, Infrastructure, Architecture, Security, Release Engineering, and Customer Support to establish consistent reliability practices across ServiceNow products and services.
The Director will play a critical role in evolving the organization from reactive operations toward an engineering-led SRE model focused on prevention, automation, resilience, and continuous improvement.
What you get to do in this role:
• Define and execute the SRE strategy and operating model across reliability engineering, service enablement, observability, automation, incident learning, and production readiness.
• Lead and develop a global organization of engineering managers, technical leaders, and SREs.
• Establish enterprise reliability standards for service ownership, tiering, golden signals, SLIs/SLOs, error budgets, alerting, on-call practices, and service health reviews.
• Lead the Service Enablement strategy by establishing minimum reliability requirements and maturity standards for critical services.
• Own the Service Registry strategy, improving service ownership, dependency visibility, maturity tracking, and impact-aware operational decision-making.
• Drive adoption of SLIs, SLOs, error budgets, and burn-rate alerting across critical services, ensuring teams consistently use reliability signals to manage customer impact.
• Build a culture of engineering away toil by turning recurring operational work and incident patterns into automation, self-service, and systemic fixes.
• Establish the AI-enabled SRE roadmap, including change-risk assessment, operational insights, remediation recommendations, and policy-driven automation.
• Drive reliability and production-readiness strategy across AWS, Azure, and GCP by establishing cloud-agnostic patterns while addressing hyperscaler-specific operational requirements.
• Partner with product and platform engineers to design, launch, and operate reliable services throughout the production lifecycle.
• Establish launch and production-readiness practices that validate availability, latency, performance, capacity, dependencies, rollback, and recovery before customer impact.
• Drive sustainable operations by scaling self-service capabilities, automation platforms, and systemic reliability improvements across engineering teams.
• Lead incident response, blameless postmortems, and corrective actions that convert production failures into lasting reliability improvements.
• Measure reliability through SLIs, SLOs, error budgets, golden signals, change failure rate, MTTR, capacity health, and toil reduction.
• Influence architecture and platform direction to simplify operating models and improve reliability across ServiceNow's global infrastructure.
• Partner with executive and engineering leaders to prioritize reliability investments and drive adoption beyond the direct SRE organization.
Qualifications
To be successful in this role you have:
• Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
• 12 years of significant leadership experience in Site Reliability Engineering, Production Engineering, Platform Engineering, Cloud Infrastructure, or large-scale distributed systems with a Bachelor's degree; or 8 years and a Master's degree; or a PhD with 5 years experience; or equivalent experience.
• Proven success leading managers and senior technical leaders across geographically distributed engineering organizations.
• Demonstrated success leading SRE, infrastructure, or reliability transformation at scale.
• Strong understanding of SLIs/SLOs, er
Interview problems reported for ServiceNow
Reported by candidates and public write-ups, not by ServiceNow. Practise each one here:
- Longest Substring Without Repeating Characters — Medium
- Number of Islands — Medium
- Container With Most Water — Medium
- Longest Repeating Character Replacement — Medium
- Valid Parentheses — Easy
- Two Sum — Easy
- Merge Two Sorted Lists — Easy
- Longest Palindromic Substring — Medium
- Coin Change — Medium
- Maximum Subarray — Medium
- Set Matrix Zeroes — Medium
- Group Anagrams — Medium
- Best Time to Buy and Sell Stock — Easy
- Top K Frequent Elements — Medium
- Product of Array Except Self — Medium
- Reverse Linked List — Easy
- Combination Sum — Medium
- Pacific Atlantic Water Flow — Medium
- House Robber II — Medium
- Longest Common Subsequence — Medium
- Merge Intervals — Medium
- Maximum Product Subarray — Medium
More at ServiceNow
- Senior Analyst, US International Tax · Santa Clara, CALIFORNIA, United States
- Advisory Solution Consultant - State Government · Sydney, NSW, Australia
- Senior in-Market Engineer · Tokyo, , Japan
- Principal Software Engineer · San Diego, CALIFORNIA, United States
- Staff Software Engineer · Santa Clara, CALIFORNIA, United States
- Staff Software Engineer · San Diego, CALIFORNIA, United States
- Principal Software Engineer · Santa Clara, CALIFORNIA, United States
- Pricing Operations Senior Analyst · Remote; San Francisco de Heredia, Heredia, Costa Rica