SRE III
Entain is seeking a talented and motivated SRE Engineer III to join its dynamic team. In this role, the individual will execute a range of site reliability activities, ensuring optimal service performance, reliability, and availability.
Entain requires the successful candidate to collaborate with cross-functional engineering teams to develop scalable, fault-tolerant, and cost-effective cloud services.
- Implement automation tools, frameworks, and CI/CD pipelines, promoting best practices and code reusability.
- Enhance site reliability through process automation, reducing mean time to detection, resolution, and repair.
- Identify and manage risks through regular assessments and proactive mitigation strategies.
- Develop and troubleshoot large-scale distributed systems in both on-prem and cloud environments.
- Deliver infrastructure as code to improve service availability, scalability, latency, and efficiency.
- Monitor support processing for early detection of issues and share knowledge on emerging site reliability trends.
- Analyze data to identify improvement areas and optimize system performance through scale testing.
- Infrastructure as Code (IaC) – Automate infrastructure setup using Terraform, Ansible, or Kubernetes (required).
- Monitoring & Observability – Monitor systems using logs, metrics, and traces to identify issues quickly (required).
- Security & Compliance – Manage access, maintain audit logs, and protect data through encryption (required).
- Incident Management – Handle production issues quickly and reduce downtime (required).
- Performance Optimization – Improve application speed, API response time, and resource usage (required).
- Dependency Management – Improve microservices reliability using retries and circuit breakers (required).
- CI/CD & Releases – Automate deployments and ensure safe rollbacks when issues occur (required).
- Capacity & Scalability – Plan infrastructure based on traffic and business growth (required).
- Chaos Engineering – Test system reliability by simulating failures (required).
- Cross-Team Collaboration – Work closely with Engineering, DevOps, Security, and Compliance teams (required).
- Production Support – Take ownership of production issues, perform initial troubleshooting, and work with engineering teams to resolve them quickly (required).
- Proficiency in monitoring tools such as Datadog, Prometheus, Grafana, or New Relic (required).
- Experience with incident management tools like PagerDuty or Opsgenie (required).
- Knowledge of automation and configuration tools including Terraform, Ansible, or Puppet (required).
- Familiarity with CI/CD tools such as Jenkins, GitHub Actions, or Argo CD (required).
- Experience with logging and tracing tools like ELK, OpenTelemetry, or Jaeger (required).
- Understanding of security tools including Vault, AWS IAM, or Snyk (required).
- Safe home pickup and home drop.
- A regular bonus and great pension.
- 24 days annual leave.
- Extra paid leave, including wellbeing and development days.
- Life assurance and Income Protection.
- Private healthcare and wellbeing support.
- INR 3,000 per month Communication allowance.
- Up to INR 16,000 per year in Crèche expenses (children under 3).
Entain is one of the world's largest sports betting and gaming entertainment groups and a FTSE 100 company. Formed when GVC Holdings rebranded as Entain in December 2020, its brands trace their history back to the 1880s and include bwin, Coral, Foxy, Gala, Ladbrokes and partypoker. Through its joint venture with MGM Resorts International, it powers BetMGM in the United States with its proprietary technology. Headquartered in London, Entain employs over 30,000 people with offices across 19 countries.
