Inference & Performance Engineer

Brak informacji o wynagrodzeniu
MidFull-time
#416319·Dodano 7 dni temu·6
Źródło: nofluffjobs.com
Aplikuj teraz

Tech Stack / Keywords

PythonMLLLMC++CUDAGPUKubernetes

Wymagania

  • Strong Python and solid general software engineering fundamentals
  • Hands-on experience deploying at least one ML/LLM model to production inference (cloud serving or edge)
  • Working knowledge of at least one inference/serving framework: vLLM, Triton, TensorRT-LLM, ONNX Runtime, llama.cpp, ggml, TGI or similar
  • Practical understanding of core optimization techniques: quantization, batching, caching, graph- or kernel-level optimization
  • Comfortable reasoning about latency/throughput/memory trade-offs
  • Solid grasp of deep learning fundamentals and transformer architectures

Nice to have:

  • Production C++ experience, especially for edge/on-device or runtime-level work
  • CUDA / GPU kernel programming exposure
  • Direct experience with llama.cpp, ggml, TensorRT-LLM, SGLang, FlashInfer, or similar low-level inference engines
  • Kubernetes / cloud infrastructure experience for GPU workloads
  • Experience with diffusion models

Obowiązki

  • Deploy and optimize ML/LLM models for production inference across cloud GPU and edge/on-device targets
  • Work with inference/serving frameworks — vLLM, Triton Inference Server, TensorRT-LLM, ONNX Runtime, or llama.cpp/ggml depending on the project's stack
  • Apply optimization techniques: quantization, pruning/distillation, operator fusion, graph/kernel-level compilation, KV-cache and batching strategies
  • Profile and tune runtime performance including latency, throughput, memory footprint, startup time, and stability under long-running sessions
  • Build and maintain inference infrastructure: containerized deployment, GPU scheduling (Kubernetes), autoscaling, observability, benchmarking pipelines
  • On select engagements, work directly in C++ inference runtimes (e.g., llama.cpp/ggml-style engines), including custom CUDA kernel work for edge and on-device deployment
  • Partner with research/ML engineers to take models from prototype to production, and with client engineering teams on integration
Innowise

Innowise

67 aktywnych ofert

Zobacz wszystkie oferty
Aplikuj teraz