IT Operations Engineer
120 - 135 PLN/ godz.B2B
SeniorFull-time·B2B
#412875·Dodano 5 dni temu·3
Źródło: nofluffjobs.comTech Stack / Keywords
JenkinsCI/CDSplunkSysdigPrometheusGrafanaKubernetesAnsibleDockerDocker ComposeBashPython
Firma i stanowisko
Ework Group is looking for an engineer experienced in IT operations and application support focused on production environment maintenance and incident management. The role involves working within Agile teams and using various tools to maintain stable and consistent production and pre-production environments.
Wymagania
- Minimum 5 years of experience in IT operations, application support (2nd and 3rd line support), or similar roles maintaining production environments.
- Experience in comprehensive incident management from reporting/alarm through RCA and implementation of preventive measures.
- At least 2 years of working with ITIL framework (incident, problem, and change management).
- Experience working in Agile environments collaborating with development teams.
- Proficient English language skills (minimum C1) to communicate technical issues to IT specialists and non-technical stakeholders.
- Strong knowledge of Jenkins for creating and maintaining pipelines and handling deployment issues.
- Strong understanding of CI/CD processes, including automation and continual improvements.
- Excellent analytical and problem-solving skills to diagnose complex production issues independently.
- Experience with log analysis and alert monitoring tools such as Splunk and Sysdig.
- Knowledge of observability tools: Prometheus, Grafana (dashboard reading and alert tuning).
- Experience working with services running on Kubernetes (pod status checks, log analysis, service restarts; without cluster administration).
- Strong knowledge of Ansible for controlled configuration change deployments.
- Good knowledge of Docker and Docker Compose.
- Basic scripting skills in Bash and Python for automation of repetitive tasks and data reconciliation.
Obowiązki
- Manage incident processes including Root Cause Analysis (RCA) for production incidents, problem diagnosis, resolution, and preventive actions.
- Continuously monitor service health to detect anomalies and prevent incidents.
- Design, implement, and maintain CI/CD pipelines using GitHub Actions and related tools for automation of build, test, security scanning, and deployment processes.
- Maintain stability and consistency of Pre-Production and Production environments without building them from scratch.
- Document operational procedures, known issues, and resolutions to build team knowledge base.
- Collaborate closely with development and platform teams to analyze and resolve issues, refine operational requirements, and ensure information flow between production and development environments.
Ework Group
87 aktywnych ofert