Senior Software Engineer,  SRE

Roku · Bengaluru, India

  • Senior
  • Full-time
  • Posted 2026-02-12
  • Confirmed live on 25 September 2026

Apply at Roku

Job description

Teamwork makes the stream work.

Roku is changing how the world watches TV

Roku is the #1 TV streaming platform in the U.S., Canada, and Mexico, and we've set our sights on powering every television in the world. Roku pioneered streaming to the TV. Our mission is to be the TV streaming platform that connects the entire TV ecosystem. We connect consumers to the content they love, enable content publishers to build and monetize large audiences, and provide advertisers unique capabilities to engage consumers.

From your first day at Roku, you'll make a valuable - and valued - contribution. We're a fast-growing public company where no one is a bystander. We offer you the opportunity to delight millions of TV streamers around the world while gaining meaningful experience across a variety of disciplines.

What does the team work on?

The Platform Infrastructure team ensures that all Roku systems run smoothly. These systems support over 100M+ users and billions in transaction revenue per year. We are a group of highly skilled infrastructure and software engineers who help build and operate systems at internet scale, including Platform (Kubernetes, Istio, Envoy, operators, and more) and Observability (OSS/CNCF-supported observability projects). We engage with multiple teams to achieve company-impacting results.

What is the role?

We are seeking a talented and experienced SRE (Site Reliability Engineering) Senior Software Engineer to help architect, build, and operate large-scale systems that stay reliable, secure, and cost-effective at internet scale. The ideal candidate takes end-to-end ownership of outcomes—treating reliability, security, cost, operability, and supportability as part of the job, not just delivering code. They bring calm, decisive incident leadership, separating mitigation from root-cause investigation and running blameless reviews that produce lasting improvements. Strong judgment and prioritization are essential, balancing roadmap delivery against operational debt, security, and compliance while clearly explaining trade-offs. This engineer pairs deep technical depth in distributed systems with broad systems thinking, and turns ambiguous objectives into executable roadmaps, epics, and backlogs. Just as important is the ability to build influence through credibility and sound reasoning, coach other engineers, and raise the operational capability of the whole team. If you enjoy solving intriguing system challenges, are innovative at heart, and thrive on making a measurable impact across teams, this role might be a great fit for you.

How will I use AI at Roku?

At Roku, we don't just use AI, we work with it. AI agents and smart tools help power drafts, analysis, and repetitive workflows, while our people bring direction, judgment, and accountability.

We value curious, adaptable builders, who can show how they've used AI, agents, or automation to move faster, improve quality, and scale their impact. Strong candidates know how to frame problems, guide AI-assisted work, check the output, and learn quickly. Above all, they bring curiosity, adaptability, and sound judgment.

What are the responsibilities of the role?

Ownership & Incident Leadership

• p]:inline" data-streamdown="list-item">Take responsibility for service outcomes end to end, including reliability, security, cost, operability, and supportability

• p]:inline" data-streamdown="list-item">Lead major incidents with composure when information is incomplete, separating mitigation from root-cause investigation and communicating impact, status, risks, and next steps without speculation

• p]:inline" data-streamdown="list-item">Facilitate comprehensive, blameless post-incident reviews that identify root causes and contributing factors, and follow corrective actions through to completion

• p]:inline" data-streamdown="list-item">Track incident trends to surface systemic issues and prioritize reliability improvements

• p]:inline" data-streamdown="list-item">Implement chaos engineering, game days, and disaster recovery exercises to validate resilience and build confidence in recovery procedures

SRE Process & Principles Implementation

• p]:inline" data-streamdown="list-item">Establish and evolve SRE principles, frameworks, and methodologies across the organization, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets

• p]:inline" data-streamdown="list-item">Manage Error Budgets as a data-driven mechanism for balancing feature velocity against reliability, and facilitate risk-tolerance conversations between engineering and product teams

• p]:inline" data-streamdown="list-item">Use SLOs, error budgets, incident data, and operational metrics to guide where the team invests

Reliability Engineering & Infrastructure

• p]:inline" data-streamdown="list-item">Reduce toil by identifying repetitive operational work and eliminating it through infrastructure-as-code, automation frameworks,

Interview problems reported for Roku

Reported by candidates and public write-ups, not by Roku. Practise each one here:

More at Roku

More like this

All open software jobs