Senior Site Reliability Engineer

28k - 31k PLN/ mies.B2B
SeniorFull-time·B2B
#430510·Dodano dziś·0
Źródło: justjoin.it
Aplikuj teraz

Tech Stack / Keywords

GrafanaDockerTerraformKubernetesAnsibleDevOpsAzure

Firma i stanowisko

Hard Rock Digital is a team focused on becoming the best online sportsbook, casino, and social gaming company in the world, rooted in the Seminole Tribe of Florida and bringing a well-known entertainment brand to the digital space.

Wymagania

Core SRE & Infrastructure:

  • Degree in Computer Science or equivalent experience.
  • 5+ years in SRE, DevOps, or similar roles managing large-scale production systems.
  • 3+ years managing Kubernetes clusters including architecture, networking, storage, security.
  • Experience with cluster autoscaling (Karpenter), upgrades, multi-cluster management.
  • Proficiency with kubectl, Helm, Kubernetes operators.
  • Advanced expertise with Grafana stack, PromQL, and Loki.
  • Hands-on experience with Java applications and JVM tuning.
  • Cloud platform expertise (AWS preferred; GCP or Azure valued).
  • Familiar with Terraform/Terragrunt or Ansible.
  • Proficiency with ArgoCD and scripting (Python, Bash, Go).
  • Experience with on-call, incident response, and root cause analysis.

AI, Automation & Agentic Systems:

  • 1+ years practical experience building/operating AI/LLM-powered tools or workflows.
  • Ability to design tool calling or multi-step reasoning agentic systems.
  • Experience integrating LLM APIs (Anthropic Claude, OpenAI, etc.).
  • Familiarity with agentic orchestration frameworks (LangChain, LangGraph, CrewAI, etc.).
  • Understanding of prompt engineering best practices.
  • Familiar with AI-assisted coding tools.
  • Experience with MCP servers.
  • Awareness of AI safety and human-in-the-loop design.

Preferred / Bonus:

  • Experience with vector databases (Pinecone, Weaviate, pgvector).
  • Experience with LLM evaluation frameworks (Galileo, LangSmith).
  • Contributions to open-source AI/ML or SRE tooling.
  • Background in data engineering or ML pipelines.

Soft Skills:

  • Strong communication skills for diverse audiences.
  • Proactive problem-solving and automation focus.
  • Ability to mentor juniors in SRE and AI-driven approaches.
  • Positive attitude and openness to feedback.

Obowiązki

Application Reliability & Performance:

  • Ensure availability, reliability, and performance of high-traffic Java-based applications in distributed environments.
  • Troubleshoot and resolve complex issues in production and non-production.
  • Participate in performance testing and monitoring.
  • Optimize Java application performance focusing on JVM tuning and scaling.

Monitoring, Observability & AIOps:

  • Deploy and manage the Grafana stack including Prometheus and Loki.
  • Implement observability strategies.
  • Create and maintain dashboards, alerts, and log queries.
  • Integrate AI/ML models for anomaly detection and alerting.

AI & Agentic Workflow Engineering:

  • Design and operate AI workflows automating alert triage, root cause analysis, and incident summarization.
  • Develop LLM agents interacting with APIs like Kubernetes, Grafana, Jira, Slack, PagerDuty.
  • Build and maintain MCP servers for AI tool exposure.
  • Evaluate and operationalize LLM frameworks and orchestration platforms.
  • Implement safety guardrails and feedback loops for AI outputs.
  • Promote AI-assisted development and operations.

Incident Management & Root Cause Analysis:

  • Support incident response and conduct post-mortems.
  • Use AI tools to accelerate incident handling and pattern detection.
  • Document and share lessons learned.

Automation & Toil Reduction:

  • Identify and automate repetitive operational workflows.
  • Build self-service tools and chatbots for system queries and procedures.
  • Measure and report toil reduction.

Collaboration & Cross-functional Support:

  • Collaborate with developers, architects, and ML engineers.
  • Work with DevOps and NOC teams.
  • Communicate SRE and AI capabilities to stakeholders.
  • Provide feedback on application performance and observability.

Benefity

  • Opportunity to experiment with cutting-edge AI tools.
  • Leadership support to deploy AI in production.
  • Autonomy to reduce operational toil via intelligent systems.

Inne informacje

Please be informed that the data controller is Hard Rock Digital ("controller"). Candidates have rights to access, rectify, erase, restrict processing, object, and portability of personal data. Data is processed for recruitment purposes, provision of mandatory data as per Labour Code. Consent can be withdrawn anytime. Data recipients include Just Join IT and entities processing data for recruitment.

Hard Rock Digital

Hard Rock Digital

Pracodawca

Aplikuj teraz