Founding Inference Engineer — FPGA + GPU Disaggregated Serving
Tech Stack / Keywords
Firma i stanowisko
AlloyCompute.ai is an early-stage startup building the next generation of LLM inference infrastructure on FPGAs. They run disaggregated inference pipelines splitting workloads across FPGA and GPU, achieving sub-200ms end-to-end latency. The company is funded and moving fast, offering a founding-engineer role with meaningful equity and core architecture ownership.
Wymagania
- Deep hands-on experience with CUDA kernel development and GPU performance engineering (Nsight, occupancy tuning, memory hierarchy optimization)
- Real production experience with modern inference engines such as vLLM, SGLang, TensorRT-LLM, or equivalent internals-level work
- Strong FPGA skills: Verilog/SystemVerilog HDL, timing closure; HLS familiarity is a plus
- Experience with AMD/Xilinx or Intel/Altera toolchains
- Solid grasp of network latency engineering: RDMA, kernel bypass (DPDK, io_uring), NIC-level optimization, and distributed serving topologies
- Understanding of LLM serving internals: continuous batching, paged/radix attention, KV-cache management, quantization (FP8/INT4)
- Ability to handle ambiguity, bare-metal debugging, and shipping without safety nets
Nice to have:
- Experience with disaggregated or heterogeneous serving architectures (e.g., Mooncake-style prefill/decode separation)
- Experience with P4/SmartNIC or network-attached accelerators
- Contributions to open-source inference projects
Obowiązki
- Architect and build disaggregated inference pipelines spanning FPGA and GPU, including prefill/decode splitting and KV-cache transfer across devices
- Write and optimize custom CUDA kernels for attention, GEMM, quantization, and sampling paths
- Extend and integrate open-source inference engines (vLLM, SGLang) with the FPGA backend, including scheduling, batching, paged attention, and speculative decoding
- Design and implement FPGA dataflow in Verilog/HDL: systolic arrays, memory controllers, on-chip interconnect, and PCIe/Ethernet interfaces
- Attack network and interconnect latency using RDMA/RoCE, NIC offload, kernel-bypass networking, and direct FPGA-to-GPU transfer
- Profile, benchmark, and optimize the stack for tokens/sec, time-to-first-token, tail latency, and cost per million tokens
Benefity
- Founding-level equity with meaningful ownership
- Direct influence over architecture, roadmap, and hiring
- Access to latest FPGA, GPU, and high-speed networking hardware for experimentation
- Competitive salary
- Flexible location
Inne informacje
Informujemy, że administratorem danych jest Gravete Sp. z o.o. z siedzibą w Poznaniu , ul. Krysiewicza 9(dalej jako "administrator"). Masz prawo do żądania dostępu do swoich danych osobowych, ich sprostowania, usunięcia lub ograniczenia przetwarzania, prawo do wniesienia sprzeciwu wobec przetwarzania, a także prawo do przenoszenia danych oraz wniesienia skargi do organu nadzorczego. Dane osobowe przetwarzane będą w celu realizacji procesu rekrutacji. Podanie danych w zakresie wynikającym z ustawy z dnia 26 czerwca 1974 r. Kodeks pracy jest obowiązkowe. W pozostałym zakresie podanie danych jest dobrowolne. Odmowa podania danych obowiązkowych może skutkować brakiem możliwości przeprowadzenia procesu rekrutacji. Administrator przetwarza dane obowiązkowe na podstawie ciążącego na nim obowiązku prawnego, zaś w zakresie danych dodatkowych podstawą przetwarzania jest zgoda. Dane osobowe będą przetwarzane do czasu zakończenia postępowania rekrutacyjnego i przez okres możliwości dochodzenia ewentualnych roszczeń, a w przypadku wyrażenia zgody na udział w przyszłych postępowaniach rekrutacyjnych - do czasu wycofania tej zgody. Zgoda na przetwarzanie danych osobowych może zostać wycofana w dowolnym momencie.
alloycompute.ai
Pracodawca