Founding Inference Engineer — FPGA + GPU Disaggregated Serving

5000 - 15k USD/ mies.B2B
C-Level / ManagerFull-time·B2B
#390377·Dodano 15 dni temu·22
Źródło: justjoin.it
Aplikuj teraz

Tech Stack / Keywords

CUDAFPGAGPUvllm/sglang

Firma i stanowisko

AlloyCompute.ai is an early-stage startup building the next generation of LLM inference infrastructure on FPGAs. They run disaggregated inference pipelines splitting workloads across FPGA and GPU, achieving sub-200ms end-to-end latency. The company is funded and moving fast, offering a founding-engineer role with meaningful equity and core architecture ownership.

Wymagania

  • Deep hands-on experience with CUDA kernel development and GPU performance engineering (Nsight, occupancy tuning, memory hierarchy optimization)
  • Real production experience with modern inference engines such as vLLM, SGLang, TensorRT-LLM, or equivalent internals-level work
  • Strong FPGA skills: Verilog/SystemVerilog HDL, timing closure; HLS familiarity is a plus
  • Experience with AMD/Xilinx or Intel/Altera toolchains
  • Solid grasp of network latency engineering: RDMA, kernel bypass (DPDK, io_uring), NIC-level optimization, and distributed serving topologies
  • Understanding of LLM serving internals: continuous batching, paged/radix attention, KV-cache management, quantization (FP8/INT4)
  • Ability to handle ambiguity, bare-metal debugging, and shipping without safety nets

Nice to have:

  • Experience with disaggregated or heterogeneous serving architectures (e.g., Mooncake-style prefill/decode separation)
  • Experience with P4/SmartNIC or network-attached accelerators
  • Contributions to open-source inference projects

Obowiązki

  • Architect and build disaggregated inference pipelines spanning FPGA and GPU, including prefill/decode splitting and KV-cache transfer across devices
  • Write and optimize custom CUDA kernels for attention, GEMM, quantization, and sampling paths
  • Extend and integrate open-source inference engines (vLLM, SGLang) with the FPGA backend, including scheduling, batching, paged attention, and speculative decoding
  • Design and implement FPGA dataflow in Verilog/HDL: systolic arrays, memory controllers, on-chip interconnect, and PCIe/Ethernet interfaces
  • Attack network and interconnect latency using RDMA/RoCE, NIC offload, kernel-bypass networking, and direct FPGA-to-GPU transfer
  • Profile, benchmark, and optimize the stack for tokens/sec, time-to-first-token, tail latency, and cost per million tokens

Benefity

  • Founding-level equity with meaningful ownership
  • Direct influence over architecture, roadmap, and hiring
  • Access to latest FPGA, GPU, and high-speed networking hardware for experimentation
  • Competitive salary
  • Flexible location
Elastyczne godziny

Inne informacje

Informujemy, że administratorem danych jest Gravete Sp. z o.o. z siedzibą w Poznaniu , ul. Krysiewicza 9(dalej jako "administrator"). Masz prawo do żądania dostępu do swoich danych osobowych, ich sprostowania, usunięcia lub ograniczenia przetwarzania, prawo do wniesienia sprzeciwu wobec przetwarzania, a także prawo do przenoszenia danych oraz wniesienia skargi do organu nadzorczego. Dane osobowe przetwarzane będą w celu realizacji procesu rekrutacji. Podanie danych w zakresie wynikającym z ustawy z dnia 26 czerwca 1974 r. Kodeks pracy jest obowiązkowe. W pozostałym zakresie podanie danych jest dobrowolne. Odmowa podania danych obowiązkowych może skutkować brakiem możliwości przeprowadzenia procesu rekrutacji. Administrator przetwarza dane obowiązkowe na podstawie ciążącego na nim obowiązku prawnego, zaś w zakresie danych dodatkowych podstawą przetwarzania jest zgoda. Dane osobowe będą przetwarzane do czasu zakończenia postępowania rekrutacyjnego i przez okres możliwości dochodzenia ewentualnych roszczeń, a w przypadku wyrażenia zgody na udział w przyszłych postępowaniach rekrutacyjnych - do czasu wycofania tej zgody. Zgoda na przetwarzanie danych osobowych może zostać wycofana w dowolnym momencie.

alloycompute.ai

alloycompute.ai

Pracodawca

Aplikuj teraz