Site Reliability Engineer
As a Site Reliability Engineer, the successful candidate will shape the stability of the systems behind every click, query and live change. The Site Reliability team protects and improves the availability, performance and resilience of the systems that support the global product.
This role combines software engineering, automation and incident response to reduce toil, sharpen observability and strengthen service health across a complex technical estate. The engineer will work with Open Telemetry, logging, telemetry and automation to surface issues faster and improve operational control. The role also includes using AI tools, LLM platforms and coding assistants to boost productivity, support autonomous operations and improve system insight.
Working across SRE, development and IT Operations, the engineer will help embed reliability throughout the software development lifecycle, lead technical work and share knowledge that lifts standards across the wider engineering community. This role is eligible for inclusion in the company’s hybrid work from home policy.
- Develop and maintain resilient tools, operational APIs and automation for effective system management.
- Use orchestration and scripting to remove manual activity, reduce toil and improve operational consistency.
- Write and contribute to code, telemetry and instrumentation that improve service reliability and observability.
- Build dashboards and operational views using telemetry from Grafana, Splunk, New Relic and related platforms.
- Configure and manage Cloudflare edge services using Infrastructure as Code and integrate edge telemetry with observability platforms.
- Diagnose incidents end to end, trace issues from the edge through to origin systems and coordinate effective remediation.
- Participate in live incident response, post-mortems and root-cause analysis to prevent recurrence.
- Maintain and administer monitoring, alerting, APM and analytics toolsets, including PagerDuty workflows.
- Drive initiatives that improve reliability, observability, performance and continuous improvement across teams.
- Mentor colleagues, share knowledge and work with IT Operations to deliver tooling that increases business value.
- Software engineering background with Python, Golang, JavaScript or similar language (required).
- Knowledge of modern development practices, including testing, source control and delivery lifecycles (required).
- An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management (required).
- Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty (required).
- Proficiency in shell scripting for automation and system management (required).
- Experience with Infrastructure as Code, including Terraform and Ansible (required).
- Knowledge of Cloudflare or a comparable edge platform, including DNS, CDN, WAF, DDoS protection and traffic management (required).
- Ability to troubleshoot distributed systems across edge, network, platform, application, dependency and origin layers (required).
- Experience working in a large-scale, 24/7 enterprise where uptime, performance and stability are critical (required).
- Practical experience using LLM platforms and coding assistants safely to improve productivity, quality and root-cause analysis (required).
bet365 is one of the world's leading online gambling companies, founded in 2000 by Denise Coates CBE. The company employs over 9,000 people and serves more than 100 million customers in 27 languages, with a market-leading position built on its In-Play betting product. It offers betting across 96 sports and hundreds of thousands of streaming events, handling millions of requests and bets at peak times. Headquartered in Stoke-on-Trent, England, bet365 is known for its software innovation and continues to develop its online betting and gaming platform.