Site Reliability Engineer
The Site Reliability Engineer will be responsible for the technical design, planning, implementation, and the highest level of performance tuning and recovery procedures for mission critical systems.
Other responsibilities include the troubleshooting of hardware, software, and networking issues, as well as ensuring that all computing operations run with optimal performance and security.
- Ensure high availability and acceptable levels of performance of mission critical infrastructure resources.
- Continuously analyse system performance and stability in production and proactively identify areas in need of optimisation.
- Respond to support escalations.
- Develop procedures to maintain security and protect systems from unauthorised use.
- Develop procedures, programs and documentation for backup and restoration of host operating systems and host-based applications.
- Develop and maintain adequate system monitoring and alerting to maximise system uptime and minimise impacts on the end user experience.
- Own and resolve escalated service issues.
- Provide technical services and support for internal systems, cloud, and network infrastructure.
- Manage virtualisation technologies: VMware, Proxmox, Microsoft, LXC.
- Perform remote monitoring and management of system alerts and notifications.
- Participate in the administration and maintenance of the remote monitoring and management system: update agent scripts, respond to alerts, monitor dashboards, and periodic system audits and review.
- Participate in post incident reviews.
- Perform any other task/responsibility which may be related and/or connected to the role of Site Reliability Engineer and/or DevOps Engineer.
- Proven experience as a Site Reliability Engineer, DevOps, System Administrator, or similar role for at least 3 years (required).
- Working knowledge of Cloud environments and respective technologies (e.g. AWS, Oracle) (required).
- Experience with databases (MySQL, PostgreSQL) (required).
- Experience with proxies, networks (LAN, WAN), and patch management tools (required).
- Proficiency with log management and visualisation (ELK stack) (required).
- Proficiency with monitoring solutions such as Prometheus and Grafana (required).
- Proficiency with Git and CI/CD pipelines (GitLab) (required).
- Proficiency with configuration management tools such as Terraform (required).
- Proficiency with containerisation and orchestration (Docker and Kubernetes, etc) (required).
- Working knowledge of server operating systems, particularly Linux and Windows (required).
- Good knowledge on network protocols such as the HTTP and TCP stacks (required).
- Knowledge of system security (e.g. intrusion detection systems) and data backup/recovery (required).
- Ability to create scripts and automate processes using tools such as Rundeck, Ansible, Bash and Puppet (required).
- Working experience using Nginx, PHP-FPM, SSL, DNS and Cloudflare (nice-to-have).
- A degree in Information Technology, Computer Science, or a related discipline (required).
- Professional certifications (nice-to-have).
- Excellent critical thinking and problem-solving skills, and attention to detail (required).
- Ability to work on own initiative, and as a team player (required).
- Organised, patient and professional demeanour, with a can-do attitude (required).
- Availability outside of working hours to resolve emergency issues promptly (required).
- Private health insurance
- Wellness allowance – up to €300 per year
- Fresh and healthy Breakfast&Lunch prepared everyday in our penthouse kitchen
- Birthday leave
- Company and team-building events
- Relocation package to Malta, including flight and two weeks of accommodation
Videoslots is an online casino operator founded in 2011 and now part of the Immense Group. It offers one of the industry's largest game libraries, with more than 13,000 casino games. The company operates in regulated markets and is known for its broad games portfolio and player tools. Videoslots is headquartered in Pieta, Malta.
