01 Zakres zadań
You will
- Lead service recovery during production incidents and coordinate technical resolver teams
- Investigate complex application and infrastructure issues using logs, SQL, and Unix/Linux tools.
- Troubleshoot production issues across distributed application environments
- Respond to queries related to application behaviour and service usage
- Communicate incident impact, recovery progress and resolution plans to stakeholders
- Lead and contribute to Root Cause Analysis (RCA) and post-incident activities
- Identify recurring operational issues and implement permanent improvements
- Improve monitoring, alerting and operational processes to increase service resilience
- Create and maintain operational documentation, runbooks and knowledge materials
- Design and implement automation using Bash, PowerShell or Python
- Develop internal operational tools to reduce manual effort and improve team efficiency
- Use AI productivity tools such as Copilot or Claude to support automation, documentation and operational analysis
- Participate in change reviews and ensure operational readiness
- Support deployments, patching, platform maintenance, and Disaster Recovery activities
- Identify operational risks and recommend appropriate mitigation
- Share technical knowledge and support less experienced team members
