Site Reliability Engineer (DevSecOps), Production Engineering - ThousandEyes
Brak informacji o wynagrodzeniu
MidFull-time
#452847·Dodano 4 dni temu·1
Źródło: CiscoTech Stack / Keywords
KubernetesPythonGoAWSPrometheusOpenTelemetryArgoCDSASTDASTLinux
Firma i stanowisko
Cisco ThousandEyes is a leading Digital Experience Assurance platform integrated with Cisco's Networking, Security, Collaboration, and Observability portfolios. The team focuses on AI-powered visibility into network and application performance through large-scale, cloud-based telemetry data.
Wymagania
- Hands-on experience deploying, operating, and troubleshooting containerized workloads in production Kubernetes environments.
- Professional experience administering Linux/Unix systems including process management, file systems, and networking protocols (TCP/IP, DNS, HTTP).
- Experience developing automation tooling, operational scripts, or backend services using Python or Go.
- Practical experience building hardened container images and integrating automated security scanning tools (SAST, DAST, or container vulnerability scanners) into CI/CD pipelines.
Nice to have:
- Familiarity with best practices for operating large-scale, highly available enterprise platforms.
- Over 3 years of experience in a related role.
- Excellent communication and documentation skills.
- Strong sense of ownership, drive, and attention to detail.
Obowiązki
- Collaborate with software engineers to optimize architecture and services for availability, latency, performance, and reliability using cloud-native tools.
- Design and implement scalable operations tooling to support platform growth and scaling across multiple regions.
- Design, deploy, and maintain AWS cloud-native services that are elastic and resilient to failure.
- Participate in and improve 24x7 incident response and on-call rotation.
- Use and expand CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to increase platform reliability.
- Automate production operations for continuous platform operation.
- Develop automation solutions for deployment, scale testing, graceful failure, and chaos testing.
- Stay updated on industry best practices for scalability and reliability.
- Identify and solve obstacles hindering operational excellence.
- Generalize and standardize solutions and processes across a microservice-based multi-region platform.
- Leverage scale testing and additional environments to improve system reliability.
- Manage infrastructure emphasizing operations/infrastructure/everything as code.
Inne informacje
This role requires working in a hybrid model with 2 days a week onsite in one of the Engineering hubs.
Cisco
68 aktywnych ofert