Nebius
Nebius
New

Senior Applied Scientist, Efficient LLM Inference & Model Optimization

Brak informacji o wynagrodzeniu
SeniorFull-time
#448258·Dodano wczoraj·1
Źródło: Nebius
Aplikuj teraz

Tech Stack / Keywords

LLMVLMPyTorchCUDATritonPythonQuantizationQATMoENVIDIA

Firma i stanowisko

Nebius is building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, focusing on large-scale GPU orchestration, inference optimization, compute, storage, networking, and applied AI. The company is listed on Nasdaq (NBIS), headquartered in Amsterdam, and has R&D hubs across Europe, the UK, North America, and Israel with a team of over 1,500 engineers.

Wymagania

  • PhD in computer science, machine learning, ML systems, computer systems, computer architecture, electrical engineering, applied math, or a closely related field.
  • Strong publication record or equivalent research artifacts in ML, ML systems, efficient inference, model compression, quantization, distillation, serving systems, or related areas.
  • Strong hands-on coding ability in Python and PyTorch, with ability to move from idea to experiment to prototype quickly.
  • Deep understanding of LLMs, VLMs, transformer inference, decoding algorithms, model compression, quantization, and production-serving tradeoffs.
  • Strong experimental design skills including ablations, baselines, metrics, statistical reasoning, and failure analysis.
  • Excellent written and verbal communication.

Nice to have:

  • First-author publications in NeurIPS, ICML, ICLR, MLSys, ACL, EMNLP, ASPLOS, OSDI, SOSP, ISCA, HPCA, or comparable venues.
  • Experience deploying ML models or inference optimizations in production.
  • Experience with vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, FlashAttention, FlashInfer, Triton, CUDA, or PyTorch internals.
  • Experience with post-training, SFT, DPO, RLHF, RLAIF, preference optimization, or synthetic data generation when connected to inference quality or efficiency.
  • Open-source research artifacts, widely used benchmarks, high-quality technical blogs, or invited talks in efficient AI systems.

Obowiązki

  • Own focused research projects from hypothesis through experiment, ablation, prototype, and production handoff.
  • Prepare internal reports, technical blogs, or papers when the work achieves external credibility.
  • Partner directly with MLEs to ensure research prototypes become usable production components.
  • Define and execute research programs in efficient LLM and VLM inference with measurable production impact.
  • Invent, evaluate, and productionize methods for quantization, QAT, distillation, speculative decoding, KV-cache reuse, KV-cache compression, long-context inference, MoE routing, and model/runtime co-optimization.
  • Build high-quality prototypes in PyTorch, Triton, CUDA-adjacent tooling, or inference-serving frameworks, then collaborate with platform engineers to productionize them.
  • Design rigorous evaluation methodology covering quality, latency, throughput, numerical stability, memory footprint, tail latency, and cost per token.
  • Publish papers, technical reports, blog posts, and open-source artifacts.
  • Collaborate with MLE, GPU kernel, backend infrastructure, product, and customer teams to select high-leverage research projects.
  • Mentor engineers and scientists on experimental design, scientific rigor, and model/system tradeoffs.

Inne informacje

Applicants must be authorized to work in the country in which they apply and provide proof of employment eligibility as a condition of hire. Nebius is an equal opportunity employer committed to providing equal employment opportunities without discrimination.

Nebius

Nebius

Pracodawca

Aplikuj teraz