Senior Manager, Data & Storage Reliability Engineering

ServiceNow · Dublin, , Ireland

  • Senior
  • Full-time
  • Posted 2026-09-22
  • Confirmed live on 25 September 2026

Apply at ServiceNow

Job description

Job Description
ServiceNow is seeking an experienced Sr Manager for our Data & Storage Reliability Engineering team.
This leader will drive engineering excellence across prevention engineering, reliability engineering, observability, incident learnings, diagnostics, automation, capacity planning, and platform risk reduction. The role requires deep technical expertise in database and distributed systems architectures, large-scale SaaS production environments, and customer-facing operations, combined with strong people leadership and execution rigor.
The ideal candidate has experience leading high-performing engineering teams responsible for identifying recurring production patterns, converting incident and escalation learnings into durable engineering improvements, strengthening observability, and building sustainable solutions that improve platform resilience at scale.
You should have experience with large-scale web applications, database platforms, distributed systems, Linux-based production environments, and a strong problem-solving mindset for reliability, automation, diagnostics, observability, and prevention. Qualified candidates will be responsible for leading a team of highly skilled engineers that push the limits of scalability, resiliency, and operational excellence.
Do you

• Have experience leading teams of engineers and developing people?

• Enjoy problem solving and using an analytical mindset to understand why systems fail and how to prevent repeat issues?

• Have a technical background in roles including database engineering, reliability engineering, systems/cloud engineering, SRE, DevOps, or production engineering?

• Know Linux operating systems, databases, observability, diagnostics, and production troubleshooting well enough to guide engineers through complex investigations?

• Have an attitude of continuous improvement and a passion for removing inefficient, repetitive, or reactive processes through automation and engineering prevention?
Answer 'yes' to these questions and we want to hear from you. Hit the Apply button and let's have a chat about the role and your skills and experiences.
Let’s start with the role
As a Sr Manager of the Data & Storage Reliability Engineering team your responsibilities will be:

• Define and execute the team-level strategy for prevention engineering, reliability, observability, resilience, and operational risk reduction across large-scale production environments.

• Lead initiatives that turn production signals, incident learnings, customer escalations, migration outcomes, and platform telemetry into durable engineering improvements.

• Partner closely with SWAT and Customer & Production Engineering to establish a continuous feedback loop between production operations and platform improvement.

• Drive improvements in observability, diagnostics, automation, reliability reviews, resiliency validation, migration readiness, and engineering guardrails.

• Identify recurring failure patterns, reliability risks, observability gaps, operational inefficiencies, scalability constraints, and performance bottlenecks, and drive action to reduce future customer impact.

• Establish reliability, resilience, observability, automation, and prevention goals for critical database and storage services.

• Champion proactive monitoring, production analytics, and automation to improve operational health and reduce repetitive manual work.

• Lead deep root cause analysis and ensure sustainable corrective actions are implemented for recurring issues and customer-impacting events.

• Partner with engineering leaders to influence database, storage, reliability, observability, and platform architecture priorities based on production evidence.

• Build and develop a world-class team of reliability, prevention, observability, and platform engineers.

• Own team management, recruitment, career development, objective setting, project prioritization, onboarding, and performance reviews.

• Manage an engineering team that supports production-facing work, including on-call or escalation participation where required.

• Drive a culture of intolerance for repetitive manual activities by promoting automation, self-service diagnostics, guardrails, and scalable engineering practices.

• Drive initiatives with partner teams to improve the reliability, resilience, scalability, and operational efficiency of the ServiceNow application and platform.

• Act as part of the escalation and crisis management ecosystem by helping convert immediate recovery learnings into sustainable engineering prevention.

• Analyze and evaluate existing processes to drive continuous improvement, operational efficiency, and prevention-oriented engineering practices.

• Provide training, documentation, dashboards, playbooks, and support to partner teams that interface with the Data & Storage Reliability Engineering team.

• Onboard new hires, new technologies, new systems, and new automation

Interview problems reported for ServiceNow

Reported by candidates and public write-ups, not by ServiceNow. Practise each one here:

More at ServiceNow

All open software jobs