Staff Site Reliability Engineer (SRE)
EarnIn · Mountain View, US
- Senior
- Full-time
- $252,000 – $308,000
- Posted 2026-09-10
- Confirmed live on 25 September 2026
Job description
About EarnIn
As one of the first pioneers of earned wage access, our passion at EarnIn is building products that deliver real-time financial flexibility for those with the unique needs of living paycheck to paycheck. Our community members access their earnings as they earn them, with options to spend, save, and grow their money without mandatory fees, interest rates, or credit checks.
We’re fortunate to have an incredibly experienced leadership team, combined with world-class funding partners like A16Z, Matrix Partners, DST, Ribbit Capital, and a very healthy core business with a tremendous runway. We’re growing fast and are excited to continue bringing world-class talent onboard to help shape the next chapter of our growth journey.
WHY this role exists
EarnIn’s products must deliver speed, reliability, resilience, and trust to community members who depend on them. As EarnIn grows, we cannot rely on heroics, tribal knowledge, manual investigation, or isolated SRE expertise. We must embed reliability practices that scale across product engineering teams, enhance customer experience, and enable rapid shipping without increasing operational risk.
This role exists to lead EarnIn’s next stage of reliability maturity: an AI-first operating model that uses AI to actively detect, investigate, respond to, learn from, and prevent production issues. As a Staff Site Reliability Engineer, you will guide technical direction for reliability across critical services, relying on AI-assisted workflows as key tools to reduce toil, speed incident response, improve production readiness, and enhance the operational quality of the engineering organization.
The base salary range for this full-time position is $252,000-$308,000, plus equity and benefits. Our salary ranges are determined by role, level, and location. This is a hybrid position in Mountain View (Headquarters) and will require in-office work 2 days a week.
HOW you will create impact
• Operate as a Staff-level technical leader: set standards, architect solutions, mentor engineers, influence work across teams, and build reusable systems and practices that multiply your impact beyond what you ship yourself
• Bring AI-first thinking to reliability practices, using AI to speed up alert triage and incident investigation, automate runbooks, surface operational knowledge, improve postmortem quality, track corrective actions, quantify reliability through scorecards, catch capacity risks early, and analyze architectural risk
• Keep human ownership and engineering judgment at the center of operations. AI helps engineers gather context faster, think more clearly, and repeat themselves less, but accountability stays with people
• Partner with SRE, product engineering, infrastructure, security, and leadership to build reliability into how teams work, making it easy to adopt and impossible to ignore
WHAT you will own
Reliability strategy and standards
• Define and evolve reliability standards across critical services, including SLIs, SLOs, error budgets, production readiness, observability, incident response, and resilience patterns.
• Establish a reliability operating model that clarifies service ownership, operational expectations, and decision-making around reliability tradeoffs for product engineering teams.
• Use AI-assisted analysis to interpret reliability trends, detect weak operational signals, highlight capacity risks using pattern recognition, and generate actionable reliability scorecards for teams, clearly delineating where AI automates data gathering and insight generation.
AI-first incident response and operational workflows
• Overhaul key stages of the incident lifecycle to achieve faster detection, sharper triage, richer context retrieval, clearer communication, and stronger follow-through.
• Command high-severity incidents as Incident Commander and reinforce the systems, tools, and practices that simplify incident management.
• Design and implement workflows in which AI assists with alert correlation, signal enrichment, root-cause exploration, runbook retrieval, postmortem drafting, and corrective-action tracking.
• Ensure AI-assisted incident workflows remain reviewable, auditable, and safe by requiring human verification at all critical steps and maintaining clear operational ownership with humans accountable for final decisions.
On-call quality and toil reduction
• Elevate on-call quality by silencing noisy alerts, automating repetitive investigations, and enabling responders to rapidly digest service context.
• Build tools that gather context from systems like Datadog, CloudWatch, incident.io, Slack, runbooks, deployment history, and service metadata.
• Transition teams from reactive paging to proactive reliability enhancement.
Architecture and resilience
• Steer service designs for graceful degradation, failure isolation, robust capacity planning, and operational safety throughout EarnIn’s AWS environment.
• Apply production data, i
Interview problems reported for EarnIn
Reported by candidates and public write-ups, not by EarnIn. Practise each one here:
- Two Sum — Easy
- Longest Palindromic Substring — Medium
- Merge Intervals — Medium
More at EarnIn
- Senior Full-Stack Engineer · Remote, US
- Senior AI Platform Engineer · Mountain View, US
- Senior Data Platform Engineer · Bengaluru, India
- Senior Front End Engineer · Vancouver, Canada
- Senior Staff Software Engineer · Mountain View, US
- Senior Product Manager, Home Screen · Mountain View, US
- Manager, Software Engineering · Bengaluru, India
- Senior HRIS Analyst (Workday) · Mountain View, US