01 Zakres zadań
Project
Join a global Site Reliability Engineering team responsible for ensuring the reliability, availability, and performance of an enterprise Kubernetes platform.
You will work on production-critical infrastructure, supporting Kubernetes deployments, troubleshooting complex incidents, improving platform resilience, and driving reliability practices across a global engineering organization.
You will
- Ensure the reliability, availability, and performance of the Kubernetes infrastructure platform
- Support the deployment, configuration, and maintenance of Kubernetes
- Diagnose and resolve infrastructure incidents, performance issues, and integration failures
- Perform root cause analysis and implement long-term reliability improvements
- Analyze logs, monitoring data, and platform metrics to identify potential issues
- Collaborate with engineering and infrastructure teams to improve platform resilience
- Participate in a 24/7 on-call rotation, including weekend support
- Coordinate a local team of 5 engineers, including monthly schedules and on-call rotations
- Work closely with the Global SRE Team Lead to coordinate the local team's contribution to global workloads
