Senior DevOps Engineer (Platform & Reliability)
Hybrid | London
Paying up to 95k + 20% bonus (realistic, can be significantly higher)
Hybrid Kubernetes – Linux – Automation – Production databases
*Extremely exciting* enterprise-scale, global entertainment organisation seeking an experienced, Senior DevOps Platform Engineer to support and enhance a modern hybrid infrastructure environment across cloud and on-premise platforms.
You’ll be working with closely with a small, highly dedicated, highly capable team.
Hands-on, senior role focused on Kubernetes, databases, Linux administration, automation, monitoring and platform reliability. You will help keep business-critical systems secure, scalable and available and support engineering teams across a global organisation.
Not CI/CD-only DevOps; this is production platform ownership.
Some workloads run in the cloud, some stay on-prem (for latency and control). Load is uneven and highly bursty.
Failover, capacity and clear runbooks are critical.
The role
- Build, manage and scale Kubernetes clusters across cloud and on-premise environments, including GKE and self-managed clusters
- Support and administer MySQL, PostgreSQL and MongoDB (application teams own product schema; you co-own operational)
- Manage monitoring, alerting and logging platforms
- Automate infrastructure with Infrastructure as Code and scripting
- Maintain high availability and resilience across critical systems
- Manage access controls and permissions in cloud platforms
- Troubleshoot, performance-tune and harden platforms
- Develop technical standards, best practices and documentation
- Provide technical support and guidance to internal engineering teams in more than one region
Experience
Note, we want to speak to people with experience in some or all of the following – you do not need every line.
- Building and managing Kubernetes in production, including GKE and/or self-managed clusters
- Designing and supporting highly available, resilient infrastructure
- Monitoring and observability — Grafana, Prometheus, ELK Stack, rsyslog or equivalent
- Infrastructure automation using Terraform, Ansible, Packer and Bash
- Strong Linux administration, particularly Ubuntu
- Exposure to medium or large-scale environments
- Ability to work independently manging competing priorities
- Collaboration with teams across locations and time zones
Beneficial:
- On-prem and hybrid Kubernetes
- Hands-on MySQL or PostgreSQL administration
- MongoDB — sharded, on-prem clustering would be particularly useful but is not a must have