01 Zakres zadań
We are looking for an enthusiastic and collaborative Site Reliability Engineer to join our growing strategic team. As an SRE, you will contribute to the stability, scalability, and automation of our cloud-based production environments and play a key role in ensuring our services are reliable and high-performing. This position is ideal for candidates at least 3-5 years of SRE experience who are passionate about both software engineering and operations.
What will you do:
- Maintain and support production systems, ensuring high availability, reliability and scalability
- Assist in implementing SRE best practices including monitoring, alerting, SLOs/SLIs, and incident response
- Deploy, configure, and manage containerized applications using Docker and Kubernetes
- Contribute to CI/CD pipeline development and configuration to streamline development and deployment processes
- Automate repetitive tasks and processes using scripting languages such as Python, Bash or Go
- Collaborate with development, QA, and operations teams to address issues and drive improvements across the stack
- Participate in on-call support rotation and help resolve incidents in a timely manner
- Document procedures, configurations, and post-incident reviews to foster a learning culture
