Inference & Performance Engineer
Brak informacji o wynagrodzeniu
MidFull-time
#416319·Dodano 7 dni temu·6
Źródło: nofluffjobs.comTech Stack / Keywords
PythonMLLLMC++CUDAGPUKubernetes
Wymagania
- Strong Python and solid general software engineering fundamentals
- Hands-on experience deploying at least one ML/LLM model to production inference (cloud serving or edge)
- Working knowledge of at least one inference/serving framework: vLLM, Triton, TensorRT-LLM, ONNX Runtime, llama.cpp, ggml, TGI or similar
- Practical understanding of core optimization techniques: quantization, batching, caching, graph- or kernel-level optimization
- Comfortable reasoning about latency/throughput/memory trade-offs
- Solid grasp of deep learning fundamentals and transformer architectures
Nice to have:
- Production C++ experience, especially for edge/on-device or runtime-level work
- CUDA / GPU kernel programming exposure
- Direct experience with llama.cpp, ggml, TensorRT-LLM, SGLang, FlashInfer, or similar low-level inference engines
- Kubernetes / cloud infrastructure experience for GPU workloads
- Experience with diffusion models
Obowiązki
- Deploy and optimize ML/LLM models for production inference across cloud GPU and edge/on-device targets
- Work with inference/serving frameworks — vLLM, Triton Inference Server, TensorRT-LLM, ONNX Runtime, or llama.cpp/ggml depending on the project's stack
- Apply optimization techniques: quantization, pruning/distillation, operator fusion, graph/kernel-level compilation, KV-cache and batching strategies
- Profile and tune runtime performance including latency, throughput, memory footprint, startup time, and stability under long-running sessions
- Build and maintain inference infrastructure: containerized deployment, GPU scheduling (Kubernetes), autoscaling, observability, benchmarking pipelines
- On select engagements, work directly in C++ inference runtimes (e.g., llama.cpp/ggml-style engines), including custom CUDA kernel work for edge and on-device deployment
- Partner with research/ML engineers to take models from prototype to production, and with client engineering teams on integration
Innowise
67 aktywnych ofert