Sr. Software Engineer, Runtime
TLDR
Build asynchronous runtime pipelines, RDMA communication primitives, and embedded firmware that optimize NPU inference performance.
About the Job
Designs and implements the low-level runtime stack that drives FuriosaAI's NPU hardware to its theoretical limits — from device driver interfaces and DMA-based I/O to kernel execution scheduling, multi-node inference, and embedded firmware.
Responsibilities
Develops the low-level runtime responsible for DMA-based I/O operations and kernel execution scheduling, maximizing inference throughput while minimizing end-to-end latency.
Builds and optimizes asynchronous execution pipelines that orchestrate data movement and compute across the NPU hardware.
Enables multi-node inference by implementing foundational communication primitives, including RDMA-based data transfer for low-latency, high-bandwidth inter-node operations.
Develops embedded firmware (PERT) that runs on the NPU's integrated ARM core, managing on-device scheduling, synchronization, and hardware resource control.
Profiles and tunes system-level performance across the full runtime stack — from firmware to user-space — to eliminate bottlenecks in real-world inference workloads.
Minimum Qualifications
BS degree in Computer Science, Engineering, or a related field, or equivalent practical experience
3+ years of relevant industry experience or equivalent practical experience in systems programming using Rust, C, or C++
Solid understanding of computer architecture fundamentals, including memory hierarchy, cache coherency, operating systems, DMA, interrupts, and MMIO
Strong communication skills, with the ability to gather requirements and drive technical alignment across teams
Preferred Qualifications
Deep expertise in low-latency runtime systems, embedded firmware development, or high-performance I/O — especially in the context of accelerator hardware.
Experience designing and implementing low-latency asynchronous execution models and scheduling systems.
Experience with DMA engines, scatter-gather I/O, or other zero-copy data transfer mechanisms.
Experience developing embedded firmware for ARM-based processors (bare-metal or lightweight RTOS environments).
Familiarity with RDMA technologies and high-performance networking for distributed or multi-node systems.
Experience with CUDA low-level runtime internals such as CUDA Graphs, stream-based execution, and asynchronous kernel launch optimization.
Experience with kernel-level performance optimizations (e.g., Linux kernel modules, eBPF, perf, ftrace).
Understanding of deep learning inference workloads and their hardware execution characteristics.
Experience with profiling and performance tuning of system software on accelerator or SoC platforms.
Contact
recruit@furiosa.ai
FuriosaAI develops high-performance, energy-efficient AI semiconductor solutions designed to drive the next generation of AI inference. By focusing on disruptive technology and rapid expansion into global markets, we aim to establish ourselves as a top player in the AI infrastructure landscape.
- Founded
- Founded 2017
- Employees
- 51-200 employees
- Industry
- semiconductors
- Funding stage
- Series C+
- Total raised
- $270M raised