Senior Site Reliability Engineer

28k - 31k PLN/ mies.B2B
SeniorFull-time·B2B
#399424·Dodano 11 dni temu·4
Źródło: justjoin.it
Aplikuj teraz

Tech Stack / Keywords

GrafanaDockerTerraformKubernetesAnsibleDevOpsAzure

Firma i stanowisko

Hard Rock Digital is a team focused on becoming the best online sportsbook, casino, and social gaming company in the world. Rooted in the kindred spirits of the Seminole Tribe of Florida, the new Hard Rock Digital taps a brand known globally as a leader in gaming, entertainment, and hospitality, expanding into the digital space.

Wymagania

Core SRE & Infrastructure:

  • Degree in Computer Science or equivalent experience.
  • 5+ years in SRE, DevOps, or infrastructure roles managing high-availability systems.
  • 3+ years managing production Kubernetes clusters with expertise in architecture, networking, storage, and security.
  • Experience with cluster autoscaling (Karpenter), upgrades, multi-cluster management.
  • Proficiency with kubectl, Helm, Kubernetes operators.
  • Advanced expertise with the Grafana observability stack.
  • Proficiency in PromQL and experience with Loki.
  • Experience managing Java applications including JVM tuning.
  • Cloud expertise (AWS preferred; GCP, Azure valued).
  • Familiar with Infrastructure as Code tools (Terraform, Terragrunt, Ansible).
  • Proficient with ArgoCD for GitOps.
  • Strong scripting in Python, Bash, or Go with CI/CD experience.
  • Experience with on-call rotations, incident response, root cause analysis.

AI, Automation & Agentic Systems:

  • 1+ years building or operating AI/LLM-powered tools or agentic workflows.
  • Ability to design agentic systems using tool calling, RAG, multi-step reasoning.
  • Experience integrating LLM APIs (e.g., Anthropic Claude, OpenAI).
  • Familiarity with orchestration frameworks (LangChain, LangGraph, CrewAI, n8n, Temporal).
  • Understanding prompt engineering best practices.
  • Familiarity with AI-assisted coding tools (Claude Code, Codex, Cursor).
  • Experience with MCP servers for AI tool exposure.
  • Awareness of AI safety and human-in-the-loop design.

Preferred / Bonus:

  • Experience with vector databases (Pinecone, Weaviate, pgvector) for RAG.
  • Experience with LLM evaluation frameworks (e.g., Galileo, LangSmith).
  • Contributions to open-source AI/ML or SRE tooling.
  • Background in data engineering or ML pipelines.

Soft Skills:

  • Strong communication skills.
  • Proactive problem-solver focusing on automation.
  • Ability to mentor junior team members.
  • Positive attitude and openness to feedback.

Obowiązki

Application Reliability & Performance:

  • Ensure availability, reliability, and performance of high-traffic Java applications.
  • Troubleshoot complex issues across production and non-production environments.
  • Participate in performance testing and monitoring.
  • Optimize with JVM tuning, resource utilization, and horizontal scaling.

Monitoring, Observability & AIOps:

  • Deploy and manage the Grafana stack (Grafana, Prometheus, Loki, Mimir, Alloy).
  • Implement observability strategies.
  • Create dashboards, alerts, and log queries.
  • Integrate AI/ML models for anomaly detection and alert correlation.

AI & Agentic Workflow Engineering:

  • Design and operate agentic AI workflows automating alert triage and incident summarization.
  • Develop tool-calling LLM agents interacting with APIs like Kubernetes, Grafana, Jira, Slack, PagerDuty.
  • Build and maintain MCP servers exposing systems as AI tool surfaces.
  • Evaluate and operationalize LLM frameworks (e.g., LangChain, LangGraph, CrewAI, n8n).
  • Implement guardrails and feedback loops for AI safety and accuracy.
  • Promote AI-assisted development and operations practices.

Incident Management & Root Cause Analysis:

  • Support incident response, conduct post-mortems.
  • Use AI tools to accelerate incident timelines and pattern discovery.
  • Document lessons learned.

Automation & Toil Reduction:

  • Identify repetitive workflows for AI-augmented automation.
  • Build self-service tools and chatbot interfaces using natural language.
  • Measure toil reduction metrics.

Collaboration & Cross-functional Support:

  • Work with developers, architects, and ML engineers.
  • Collaborate with DevOps and NOC teams.
  • Communicate SRE practices and AI capabilities to stakeholders.
  • Provide feedback on application performance and observability.

Inne informacje

Please be informed that the data controller is Hard Rock Digital (hereinafter "controller"). You have the right to request access to your personal data, their rectification, erasure or restriction of processing, the right to object to processing, as well as the right to data portability and to lodge a complaint to the supervisory authority. Personal data will be processed for the purpose of the recruitment process. Provision of data to the extent resulting from the Act of 26 June 1974 Labour Code is mandatory. In the remaining scope, providing data is voluntary. Refusal to provide mandatory data may result in the impossibility to carry out the recruitment process. The Administrator processes mandatory data on the basis of a legal obligation incumbent upon him/her, while with regard to additional data, the basis for processing is consent. Personal data will be processed until the recruitment procedure is completed and for the period of the possibility of asserting potential claims, and in the case of consent to participate in future recruitment procedures - until the withdrawal of such consent. Consent to the processing of personal data can be withdrawn at any time. The recipient of the data is the Just Join IT service and other entities to whom we have entrusted the processing of data in connection with recruitment.

Hard Rock Digital

Hard Rock Digital

Pracodawca

Aplikuj teraz