Site Reliability Engineer (SRE)
Full Time £60,000 - £65,000 / yearJob Overview
This role offers a hybrid work offering to be present in our London office twice a week.
Reward Gateway|Edenred is a leading digital platform for services and payments for people at work, connecting 52 million users and 2 million partner merchants in 45 countries via close to 1 million corporate clients.
Reward Gateway|Edenred is a leading digital platform for services and payments for people at work, connecting 52 million users and 2 million partner merchants in 45 countries via close to 1 million corporate clients.
Our shared mission of ‘Making the World a Better Place to Work' and ‘Enriching connections, For good’, guides our every action and charts a sustainable path to a better future.
Due to expansion, an opportunity has become available for a Site Reliability Engineer to join our team to help us transform our existing operational workloads to an SRE approach.
Due to expansion, an opportunity has become available for a Site Reliability Engineer to join our team to help us transform our existing operational workloads to an SRE approach.
Some of your responsibilities & core duties will include:
- Integrating tightly with our Product Engineering teams
- Following SRE practices and maintaining high standards of compliance
- Implementing a new standard of observability utilising SLI/SLO/Error Budgets
- Continually evolving our observability platforms for greater coverage
- Using a code-first approach to build and changes to reduce TOIL
- Advocating a strong focus on availability, reliability and uptime
- Liaising and embedding with the Engineering teams for the constant evolution of metrics
- Working towards planned roadmap goals
- Actively taking part in the daily stand-ups and keeping sprints on track
- Keeping up-to-date documentation in the JIRA & Confluence tools
- Taking part in SRE Incident Management processes
- Acting as a key Incident Commander within the Incident Management process
- Taking part in SRE On Call
- Ensuring a focus on cost efficiency for the platforms & services
- Working with team members to foster collaboration and ongoing communication with stakeholders
The experience and key skills you will have:
- Good experience in DevOps or SRE, with a keen interest to learn and grow as a Site Reliability Engineer
- Observability product experience (eg Datadog)
- Managing services using SLI/SLO & Error Budgets
- Experience with AWS or other cloud providers
- Experience in HA environments
- Automation skills through Terraform, Python, Bash or similar
- Good SRE skills with a good understanding of SRE practices
- Some understanding of SQL, PHP, Kubernetes, CI/CD advantageous
- Ability to work both independently and as part of a team
- Ability to work under pressure and be highly reliable
- Adaptability and flexibility to change in a fast-moving environment
- An ability to learn new tools and processes quickly and impart that knowledge
Make Your Resume Now