Inference Stack Engineer
Brak informacji o wynagrodzeniu
SeniorFull-time
#411203·Dodano wczoraj·0
Źródło: EER PolandTech Stack / Keywords
AIPythonArchitectureC++PyTorchTensorFlow
Firma i stanowisko
EER Poland is hiring for a role focused on developing a next-generation AI inference stack designed for high-performance execution on modern and custom compute architectures.
Wymagania
- Strong software engineering background with C++ and Python
- Experience with performance-critical systems or compiler-related work
- Understanding of AI model execution, especially transformers and LLMs
- Familiarity with compute graphs, tensor operations, or execution frameworks
- Ability to analyze complex systems end-to-end from model to runtime to hardware
- Experience working with large codebases and system-level debugging
- Strong communication skills and ability to work in cross-functional teams
Nice to have:
- Experience with compiler frameworks such as LLVM, MLIR, Triton, TVM, XLA
- Experience contributing to deep learning frameworks like PyTorch, TensorFlow, JAX
- Understanding of GPU or accelerator execution models
- Experience with kernel optimization or operator-level performance tuning
- Knowledge of distributed inference systems (e.g., NCCL, RPC-based serving)
- Familiarity with hardware-aware optimizations including memory hierarchy, vectorization, and scheduling
Obowiązki
- Design and build components of an AI inference stack, from high-level model representation to low-level execution
- Develop and extend a Python-based DSL for expressing AI workloads and kernels
- Work on compiler infrastructure including:
- IR design and transformation pipelines
- graph lowering and optimization passes
- backend code generation for target execution environments
- Optimize model execution for latency, throughput, memory efficiency, and numerical stability
- Contribute to runtime systems responsible for model execution and scheduling
- Profile and analyze inference workloads to identify system bottlenecks
- Collaborate closely with hardware and systems engineers on execution efficiency
- Influence architecture decisions for next-generation AI execution platforms
Benefity
- Work on the core execution layer of modern AI systems
- Direct impact on inference performance of large-scale AI workloads
- Collaboration with experts in compilers, systems, and AI infrastructure
- Highly technical environment with strong engineering autonomy
- Opportunity to shape the architecture of a next-generation inference stack
- Competitive compensation and flexible working model
EER Poland
21 aktywnych ofert