Network Reliability Engineer
200 - 250 PLN/ godz.B2B
NieokreśloneFull-time·B2B
#451280·Dodano 3 miesiące temu·0
Źródło: MargoTech Stack / Keywords
GoPythonBashLinuxUbuntuDebianGPUHPCVLANTCP/IP
Wymagania
- Experience with Go or Python
- Strong scripting skills (Bash, Python)
- Hands-on experience with Linux systems (Ubuntu/Debian)
- Preferred hands-on experience with GPU & HPC infrastructure
- Knowledge of networking (VLAN/LAN, TCP/IP, DNS, BGP, load-balancing, IPv6)
- Familiarity with monitoring and logging tools (Prometheus, Grafana, Elastic)
- Comfortable with Infrastructure-as-Code (Ansible, Salt, AWX)
- Experience managing relational databases (MariaDB)
- Understanding of CI/CD pipelines (GitLab)
- Comfortable with English (written and spoken)
Obowiązki
- Build a large AI infrastructure with monitoring, diagnosis, and remediation of production incidents
- Troubleshoot high-impact production issues in collaboration with other engineering teams
- Participate in an on-call rotation to handle incidents and ensure service continuity
- Implement and maintain observability solutions to monitor AI infrastructure and application health
- Contribute to AI infrastructure lifecycle management across different environments and countries
- Promote and apply best practices in terms of stability, resiliency, scalability, and security
- Maintain clear technical documentation for tools and procedures
- Contribute to system and tool evolution based on production feedback
- Collaborate closely with development teams to ensure infrastructure readiness
- Participate in team rituals and knowledge-sharing initiatives
Margo
30 aktywnych ofert