Inference Stack Engineer

Brak informacji o wynagrodzeniu
SeniorFull-time
#411203·Dodano wczoraj·0
Źródło: EER Poland
Aplikuj teraz

Tech Stack / Keywords

AIPythonArchitectureC++PyTorchTensorFlow

Firma i stanowisko

EER Poland is hiring for a role focused on developing a next-generation AI inference stack designed for high-performance execution on modern and custom compute architectures.

Wymagania

  • Strong software engineering background with C++ and Python
  • Experience with performance-critical systems or compiler-related work
  • Understanding of AI model execution, especially transformers and LLMs
  • Familiarity with compute graphs, tensor operations, or execution frameworks
  • Ability to analyze complex systems end-to-end from model to runtime to hardware
  • Experience working with large codebases and system-level debugging
  • Strong communication skills and ability to work in cross-functional teams

Nice to have:

  • Experience with compiler frameworks such as LLVM, MLIR, Triton, TVM, XLA
  • Experience contributing to deep learning frameworks like PyTorch, TensorFlow, JAX
  • Understanding of GPU or accelerator execution models
  • Experience with kernel optimization or operator-level performance tuning
  • Knowledge of distributed inference systems (e.g., NCCL, RPC-based serving)
  • Familiarity with hardware-aware optimizations including memory hierarchy, vectorization, and scheduling

Obowiązki

  • Design and build components of an AI inference stack, from high-level model representation to low-level execution
  • Develop and extend a Python-based DSL for expressing AI workloads and kernels
  • Work on compiler infrastructure including:
    • IR design and transformation pipelines
    • graph lowering and optimization passes
    • backend code generation for target execution environments
  • Optimize model execution for latency, throughput, memory efficiency, and numerical stability
  • Contribute to runtime systems responsible for model execution and scheduling
  • Profile and analyze inference workloads to identify system bottlenecks
  • Collaborate closely with hardware and systems engineers on execution efficiency
  • Influence architecture decisions for next-generation AI execution platforms

Benefity

  • Work on the core execution layer of modern AI systems
  • Direct impact on inference performance of large-scale AI workloads
  • Collaboration with experts in compilers, systems, and AI infrastructure
  • Highly technical environment with strong engineering autonomy
  • Opportunity to shape the architecture of a next-generation inference stack
  • Competitive compensation and flexible working model
EER Poland

EER Poland

21 aktywnych ofert

Zobacz wszystkie oferty
Aplikuj teraz