05-07-26: Today, I caught up with an old friend, Ron Renwick, and the team from Xelera Felix Winterstein, the CEO, and Andrea Suardi, Head of Acceleration, to discuss their new Silva offering and how they are bringing AI inference to the network edge. Recently, the Xelera team completed the STAC-ML™ Markets (Inference) benchmark audit on a stack that includes a STAC-ML™ Pack for Xelera Silva with AMD Alveo™ V80 on an HPE Proliant DL385 Gen10 Plus v2 server. For context, here are some highlights from this report:
For the small (GBT_A) and medium (GBT_B) models, 99th percentile latencies were <= 1.95µs for all Numbers of Model Instances (NMI) tested, with worst-case instance throughput > 560K inferences per second at the highest NMIs tested For the large (GBT_C) model, the 99th percentile latency was 2.88µs, with worst-case instance throughput of 379K inferences per second The maximum latency was <= 12.3µs across all models and NMI tested00:00 Hello & Welcome00:09 Ron Background01:04 Felix Background01:40 Andrea Background02:30 Xelera Origin Story06:00 FPGAs are very sticky once you start working with them07:33 What is Xelera, what do they do?07:55 It’s not about FPGAs08:30 It’s network acceleration for the data center08:40 We are seeing a massive buildout in data center capacity09:05 Number one KPI (Key Performance Indicator) is compute capacity10:31 All that processing is starting to reach its limits10:50 In cybersecurity and networking, we see this lack of computing becoming painful11:10 The thing every infrastructure buyer or architect needs to consider11:24 What are the products that Xelera offers, and who do they address?11:45 First is a classic DPU softNIC12:05 Strong footprint in Cyber Security vertical12:26 Second product is AI Acceleration12:48 We bring Andrea back in to talk about his STAC Research Event Talk13:40 When you attend STAC events, you’re in a room with really deeply technical people14:00 All these people are here to solve one problem: optimize their execution stack14:57 Tail latency spikes are discussed, and the impact on trading15:17 Ron drills into this tail latency issue a bit more16:12 The cost of latency varies from firm to firm16:45 Too slow, and you are a price taker, not a price maker17:13 It’s important to win more than 50.001% of the time17:24 What we’re selling today is the ability to trade both faster and smarter19:40 Gradient boosting 50 us running in the CPU, went down to 5us running on the FPGA20:10 This became Silvia, an agent who runs fast.20:30 What other markets can Silva be applied to, for example, security22:00 Silva can look at each packet at the edge, for Ransomware or DDoS, pattern detection23:00 Silva is the engine, the model can vary depending on the use case, time23:30 What numbers can you provide? 24:00 Two main Silva modes, the first is Offload, inference via API with one 2M nodes25:10 Now 30us to run gradient boosting on CPU, if you use Silva, it’s 1us on FPGA25:40 LSTN with a 1M parameter model is 1ms. If you offload to an FPGA, you get 3us26:10 There is a cost using the PCIe bus, 500-700ns depending on packet size26:40 The second Silva mode is Inline mode, and this runs entirely on the FPGA27:10 Are there specific use cases or environments driving deployment environment28:15 The lowest latency requires optimize the stack; cloud won’t work29:00 We rewrote the CPU kernel to tune the cache and have it work with Gradient Boosting30:00 You can go from 500us for a standard framework 15us if you rewrite the kernel30:20 Silvia makes using an FPGA easier 30:30 SmartNICs can remove the 30% CPU overhead for processing packets31:20 Offload allows you to get your hands on the data before any packet touches the CPU32:25 In HFT, Silva can save a tremendous amount of latency by avoiding the PCIe bus33:10 When you are talking to potential customers about Silva, who is expressing interest?33:30 All three, C-level, architects, and engineers33:40 The C-Suite is often the starting point.34:35 From a Silva perspective, who do you talk with initially?34:50 The quant team initially approaches us from trading side, sometimes, infrastructure35:30 Where do you see Xelera going over the next few years?36:00 Cybersecurity, network security36:10 We are announcing on the podcast today support for 400 Gbps36:40 There is an extension of the Silvia product into network security37:40 On Silvia, we have just launched the CPU-only version for not only HFT, but also others38:00 Embedded AI use case running on the FPGA, like Gradient boosting tree39:00 Moving to LLM in general, the market for Inference is moving39:15 Cost of executing LLMs is exploding39:40 What is the one thing people should remember about Xelera?39:35 Key points is that there are plenty of SmartNIC/DPU use cases, gaps in architectures40:00 SmartNIC and DPU use cases are diverse, focusing on specifics has proven successful40:30 One thing to remember about Xelera is that we are offering a full solution40:40 You don’t have to become an FPGA expert or developer to use this technology 41:00 Closing statements