Make Your Resume Now

Senior Site Reliability Engineer

Posted September 15, 2026
Full-time Mid-Senior Level

Job Overview

Responsibilities

  • Operate against our existing SLIs, SLOs and error budgets,  hold services to them, review and adjust targets as systems evolve, and use budget burn to arbitrate between reliability work and feature velocity.

  • Own the incident lifecycle end to end   detection, response, mitigation, blameless postmortems with tracked follow-through   troubleshooting across the whole stack (OS, application, database, cache, network), and mature the on-call rotation around it: alerts tuned for signal, runbooks kept current, MTTD and MTTR trending down.

  • Extend and improve our observability stack   instrument new services, close coverage gaps in metrics, logs and traces, and raise dashboard and alert quality so teams can diagnose their own services.

  • Design and maintain safe release processes: canary, progressive rollout, automated rollback across dev, staging and production.

  • Eliminate toil through automation; treat repetitive manual operations as bugs to be engineered away.

  • Own infrastructure as code   provisioning, configuration, policy, and the documentation around it.

  • Capacity planning, performance tuning and cloud cost efficiency   forecast growth, model headroom, upgrade before saturation.

  • Implement and maintain infrastructure security controls   secrets and credential management in Vault, access control, monitoring and response.

  • Conduct production readiness reviews and resilience testing for new and existing services.

  • Operate our MCP gateway and internal AI tooling as production infrastructure   availability, access control, rate limiting, cost and usage visibility   and apply AI-assisted automation to operational work such as incident triage, log summarisation and runbook generation.

Ready to Apply?

Take the next step in your career journey

Stand out with a professional resume tailored for this role

Build Your Resume – It’s Free!