Staff Software Engineer, Reliability
LinkedIn · Mountain View, CA, United States
- Senior
- Full-time
- $156,000 – $255,000
- Posted 2026-09-24
- Confirmed live on 25 September 2026
Job description
Job Description
This role will be based in Mountain View, CA.
At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. The work location of this role is hybrid, meaning it will be performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team.
Site Health Platform sits at the core of LinkedIn’s Reliability Infrastructure organization, with a primary focus on the end-to-end incident management ecosystem. Our mission is for every member and customer to experience LinkedIn as "always on", every engineer to benefit from a more insightful and proactive site-wide reliability ecosystem, and every business and product owner to be well-informed about service disruptions as they occur.
We own the full incident lifecycle across thousands of services and multiple regions, from incident response and mitigation, through problem management and post-incident learning. The platforms we build are the backbone of how LinkedIn detects issues, coordinates incident response, captures context, and turns outages and near misses into structured, actionable insights.
By transforming incidents into data and learnings, we enable teams to systematically improve reliability over time. Our work informs engineering priorities, infrastructure investments, capacity planning, and executive decision-making, ensuring the network is dependable when it matters most.
You will be exposed to many different technologies, architectures, and systems hosted in state-of-the-art data centers across the globe.
Responsibilities:
• Designing and evolving the core incident management platforms that power LinkedIn’s full incident lifecycle, from detection and response to problem management and prevention, across thousands of services and teams.
• Serving in a critical on-call rotation, providing expert incident triage and coordination during high-severity outages. Partnering closely with service owners and product teams to diagnose issues quickly, mitigate member impact, and drive timely resolution under pressure.
• Transforming raw, unstructured incident data into clear, actionable intelligence using AI and LLM-based systems, including automated summarization, classification, root cause signals, and mitigation recommendations.
• Building analytics and insights that surface systemic reliability risks, recurring failure patterns, and cross-service dependencies, enabling org-level prioritization rather than isolated, service-by-service fixes.
• Building platforms and tools that enable realistic, fleet-wide stress testing of data center and regional capacity, validating incident readiness across dependencies, traffic patterns, and growth scenarios before they impact a significant production outage.
• Driving consistency, clarity, and quality in how incidents are declared, managed, reviewed, and learned from, raising the reliability bar across a large, fast-moving engineering organization.
• Influencing service architecture, SLOs, and reliability standards through platforms, data, and technical leadership, ensuring improvements are durable, measurable, and adopted at scale.
Qualifications
Basic Qualifications:
• Bachelor’s degree in Computer Science, Engineering, or related technical field or equivalent practical experience. Many postings also prefer or require an advanced degree (MS/PhD) for Staff-level roles.
• 6+ years of professional experience in software development, distributed systems, or reliability engineering.
• Experience leading technical projects/providing architectural leadership
• Experience building products and operating large-scale distributed systems.
• Experience with two or more backend languages such as Go, Python or Java with a track record of owning complex production systems.
• Full-stack engineering experience, including building user-facing web applications and operational dashboards using modern frontend frameworks such as React.js, along with backend APIs and data pipelines.
• Understanding of web development fundamentals including API design, performance, accessibility and building intuitive interfaces for engineers and operational users.
• Understanding of reliability engineering principles, incident management, observability and operating systems under failure conditions.
• Demonstrated ability to lead technical design across teams, influence architecture beyond direct ownership and drive adoption through well-designed platforms.
• Debugging and root cause analysis skills, with the ability to communicate complex technical findings clearly to engineers, partners and leadership.
Preferred Qualifications:
• Experience applying AI or LLM-based techniques to operational or incident data, including automated summarization, classification, root cause hypothesis generation or reliability recommendations.
• Familiarity with vector databases and retrieval-based systems
Interview problems reported for LinkedIn
Reported by candidates and public write-ups, not by LinkedIn. Practise each one here:
- Maximum Subarray — Medium
- Valid Parentheses — Easy
- Number of Islands — Medium
- Maximum Product Subarray — Medium
- Minimum Window Substring — Hard
- Search in Rotated Sorted Array — Medium
- Lowest Common Ancestor of a Binary Search Tree — Medium
- Maximum Depth of Binary Tree — Easy
- Serialize and Deserialize Binary Tree — Hard
- Merge Intervals — Medium
- Two Sum — Easy
- Merge Two Sorted Lists — Easy
- Merge K Sorted Lists — Hard
- Binary Tree Level Order Traversal — Medium
- Insert Interval — Medium
- Product of Array Except Self — Medium
- House Robber — Medium
- Palindromic Substrings — Medium
- Graph Valid Tree — Medium
- Number of Connected Components in an Undirected Graph — Medium
- Longest Substring Without Repeating Characters — Medium
- Same Tree — Easy
- Reorder List — Medium
- Invert Binary Tree — Easy
- Combination Sum — Medium
- Course Schedule — Medium
- House Robber II — Medium
- Word Break — Medium
- Unique Paths — Medium
- Best Time to Buy and Sell Stock — Easy
- Longest Consecutive Sequence — Medium
- Find Minimum in Rotated Sorted Array — Medium
- Reverse Linked List — Easy
- Linked List Cycle — Easy
- Validate Binary Search Tree — Medium
- Kth Smallest Element in a BST — Medium
- Longest Palindromic Substring — Medium
More at LinkedIn
- Learning Designer · San Francisco, CA, United States
- Senior Associate , Decision science · Bengaluru, KA, India
- Senior Account Executive, Marketing Solutions · Tokyo, 13, Japan
- Sales Strategy and Operations Associate · New York, NY, United States
- Staff Technical Program Manager · Mountain View, CA, United States
- Associate Engineer, Data Center · Manassas, VA, United States
- Sr. Program Manager, AI Initiatives, Go-To-Market Enablement · Sunnyvale, CA, United States
- Sr. Director, Software Engineering - Product Platform & Infrastructure · Mountain View, CA, United States