SmartNICs Today

Ultra-low Latency Inference at the Network Edge with Xelera


Listen Later

05-07-26: Today, I caught up with an old friend, Ron Renwick, and the team from Xelera Felix Winterstein, the CEO, and Andrea Suardi, Head of Acceleration, to discuss their new Silva offering and how they are bringing AI inference to the network edge. Recently, the Xelera team completed the STAC-ML™ Markets (Inference) benchmark audit on a stack that includes a STAC-ML™ Pack for Xelera Silva with AMD Alveo™ V80 on an HPE Proliant DL385 Gen10 Plus v2 server. For context, here are some highlights from this report:

  • For the small (GBT_A) and medium (GBT_B) models, 99th percentile latencies were <= 1.95µs for all Numbers of Model Instances (NMI) tested, with worst-case instance throughput > 560K inferences per second at the highest NMIs tested 
  • For the large (GBT_C) model, the 99th percentile latency was 2.88µs, with worst-case instance throughput of 379K inferences per second  
  • The maximum latency was <= 12.3µs across all models and NMI tested
  • Table of Contents

    • 00:00 Hello & Welcome
    • 00:09 Ron Background
    • 01:04 Felix Background
    • 01:40 Andrea Background
    • 02:30 Xelera Origin Story
    • 06:00 FPGAs are very sticky once you start working with them
    • 07:33 What is Xelera, what do they do?
    • 07:55 It’s not about FPGAs
    • 08:30 It’s network acceleration for the data center
    • 08:40 We are seeing a massive buildout in data center capacity
    • 09:05 Number one KPI (Key Performance Indicator) is compute capacity
    • 10:31 All that processing is starting to reach its limits
    • 10:50 In cybersecurity and networking, we see this lack of computing becoming painful
    • 11:10 The thing every infrastructure buyer or architect needs to consider
    • 11:24 What are the products that Xelera offers, and who do they address?
    • 11:45 First is a classic DPU softNIC
    • 12:05 Strong footprint in Cyber Security vertical
    • 12:26 Second product is AI Acceleration
    • 12:48 We bring Andrea back in to talk about his STAC Research Event Talk
    • 13:40 When you attend STAC events, you’re in a room with really deeply technical people
    • 14:00 All these people are here to solve one problem: optimize their execution stack
    • 14:57 Tail latency spikes are discussed, and the impact on trading
    • 15:17 Ron drills into this tail latency issue a bit more
    • 16:12 The cost of latency varies from firm to firm
    • 16:45 Too slow, and you are a price taker, not a price maker
    • 17:13 It’s important to win more than 50.001% of the time
    • 17:24 What we’re selling today is the ability to trade both faster and smarter
    • 19:40 Gradient boosting 50 us running in the CPU, went down to 5us running on the FPGA
    • 20:10 This became Silvia, an agent who runs fast.
    • 20:30 What other markets can Silva be applied to, for example, security
    • 22:00 Silva can look at each packet at the edge, for Ransomware or DDoS, pattern detection
    • 23:00 Silva is the engine, the model can vary depending on the use case, time
    • 23:30 What numbers can you provide? 
    • 24:00 Two main Silva modes, the first is Offload, inference via API with one 2M nodes
    • 25:10 Now 30us to run gradient boosting on CPU, if you use Silva, it’s 1us on FPGA
    • 25:40 LSTN with a 1M parameter model is 1ms. If you offload to an FPGA, you get 3us
    • 26:10 There is a cost using the PCIe bus, 500-700ns depending on packet size
    • 26:40 The second Silva mode is Inline mode, and this runs entirely on the FPGA
    • 27:10 Are there specific use cases or environments driving deployment environment
    • 28:15 The lowest latency requires optimize the stack; cloud won’t work
    • 29:00 We rewrote the CPU kernel to tune the cache and have it work with Gradient Boosting
    • 30:00 You can go from 500us for a standard framework 15us if you rewrite the kernel
    • 30:20 Silvia makes using an FPGA easier 
    • 30:30 SmartNICs can remove the 30% CPU overhead for processing packets
    • 31:20 Offload allows you to get your hands on the data before any packet touches the CPU
    • 32:25 In HFT, Silva can save a tremendous amount of latency by avoiding the PCIe bus
    • 33:10 When you are talking to potential customers about Silva, who is expressing interest?
    • 33:30 All three, C-level, architects, and engineers
    • 33:40 The C-Suite is often the starting point.
    • 34:35 From a Silva perspective, who do you talk with initially?
    • 34:50 The quant team initially approaches us from trading side, sometimes, infrastructure
    • 35:30 Where do you see Xelera going over the next few years?
    • 36:00 Cybersecurity, network security
    • 36:10 We are announcing on the podcast today support for 400 Gbps
    • 36:40 There is an extension of the Silvia product into network security
    • 37:40 On Silvia, we have just launched the CPU-only version for not only HFT, but also others
    • 38:00 Embedded AI use case running on the FPGA, like Gradient boosting tree
    • 39:00 Moving to LLM in general, the market for Inference is moving
    • 39:15 Cost of executing LLMs is exploding
    • 39:40 What is the one thing people should remember about Xelera?
    • 39:35 Key points is that there are plenty of SmartNIC/DPU use cases, gaps in architectures
    • 40:00 SmartNIC and DPU use cases are diverse, focusing on specifics has proven successful
    • 40:30 One thing to remember about Xelera is that we are offering a full solution
    • 40:40 You don’t have to become an FPGA expert or developer to use this technology 
    • 41:00 Closing statements
    • ...more
      View all episodesView all episodes
      Download on the App Store

      SmartNICs TodayBy Scott Schweitzer