AI Post Transformers

Automating CNN Mapping on Embedded FPGAs


Listen Later

This episode explores how fpgaConvNet turns CNN inference into a Synchronous Dataflow problem so embedded FPGA accelerators can be designed with analyzable schedules, buffers, and resource tradeoffs instead of ad hoc hardware tuning. It explains why CNN deployment on robots, drones, and cars is constrained as much by data movement, latency, and power as by raw arithmetic, and why FPGAs can outperform embedded GPUs when the hardware is tailored carefully to a model’s structure. The discussion highlights the paper’s central claim that formalizing the mapping problem enables automated design-space exploration across very different CNN topologies, rather than just optimizing a single benchmark. It also examines where that approach is strong and where listeners should be skeptical, including whether the reported GPU speedups are fair and how well the clean SDF abstraction survives real hardware implementation.
Sources:
1. fpgaConvNet: A Toolflow for Mapping Diverse Convolutional Neural Networks on Embedded FPGAs — Stylianos I. Venieris, Christos-Savvas Bouganis, 2017
http://arxiv.org/abs/1711.08740
2. Static Scheduling of Synchronous Data Flow Programs for Digital Signal Processing — Edward A. Lee, David G. Messerschmitt, 1987
https://scholar.google.com/scholar?q=Static+Scheduling+of+Synchronous+Data+Flow+Programs+for+Digital+Signal+Processing
3. Synchronous Data Flow — Edward A. Lee, David G. Messerschmitt, 1987
https://scholar.google.com/scholar?q=Synchronous+Data+Flow
4. Scenario-aware dataflow: modeling, analysis and implementation of dynamic applications — Sander Stuijk, Marc C. W. Geilen, Bart D. Theelen, Twan Basten, 2011
https://scholar.google.com/scholar?q=Scenario-aware+dataflow%3A+modeling%2C+analysis+and+implementation+of+dynamic+applications
5. fpgaConvNet: A Toolflow for Mapping Diverse Convolutional Neural Networks on Embedded FPGAs — Stylianos I. Venieris, Christos-Savvas Bouganis, 2017
https://scholar.google.com/scholar?q=fpgaConvNet%3A+A+Toolflow+for+Mapping+Diverse+Convolutional+Neural+Networks+on+Embedded+FPGAs
6. DNNWeaver: From High-Level Deep Network Models to FPGA Acceleration — Hyoukjun Sharma, Jongse Park, Emmanuel Amaro, Bradley Thwaites, Praneeth Kotha, Anmol Gupta, Joon Kyung Kim, Asit Mishra, and Hsien-Hsin S. Lee, 2016
https://scholar.google.com/scholar?q=DNNWeaver%3A+From+High-Level+Deep+Network+Models+to+FPGA+Acceleration
7. Caffeine: Towards Uniformed Representation and Acceleration for Deep Convolutional Neural Networks — Chen Zhang, Peng Li, Guangyu Sun, Yijin Guan, Bingjun Xiao, and Jason Cong, 2016
https://scholar.google.com/scholar?q=Caffeine%3A+Towards+Uniformed+Representation+and+Acceleration+for+Deep+Convolutional+Neural+Networks
8. FINN: A Framework for Fast, Scalable Binarized Neural Network Inference — Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella, Michaela Blott, Philip Leong, Magnus Jahre, and Kees Vissers, 2017
https://scholar.google.com/scholar?q=FINN%3A+A+Framework+for+Fast%2C+Scalable+Binarized+Neural+Network+Inference
9. Optimizing FPGA-based Accelerator Design for Deep Convolutional Neural Networks — Yu Wang, Jiajun Xu, Yanzhi Wang, and Huazhong Yang, 2015
https://scholar.google.com/scholar?q=Optimizing+FPGA-based+Accelerator+Design+for+Deep+Convolutional+Neural+Networks
10. Efficient Processing of Deep Neural Networks: A Tutorial and Survey — Vivienne Sze, Yu-Hsin Chen, Tien-Ju Yang, and Joel S. Emer, 2017
https://scholar.google.com/scholar?q=Efficient+Processing+of+Deep+Neural+Networks%3A+A+Tutorial+and+Survey
11. ViTA: A Vision Transformer Inference Accelerator for Edge Applications — Shashank Nag, Gourav Datta, Souvik Kundu, Nitin Chandrachoodan, Peter A. Beerel, 2023
https://scholar.google.com/scholar?q=ViTA%3A+A+Vision+Transformer+Inference+Accelerator+for+Edge+Applications
12. ME-ViT: A Single-Load Memory-Efficient FPGA Accelerator for Vision Transformers — Kyle Marino, Pengmiao Zhang, Viktor K. Prasanna, 2024
https://scholar.google.com/scholar?q=ME-ViT%3A+A+Single-Load+Memory-Efficient+FPGA+Accelerator+for+Vision+Transformers
13. DRViT: A Dynamic Redundancy-Aware Vision Transformer Accelerator via Algorithm and Architecture Co-Design on FPGA — Xiangfeng Sun, Yuanting Zhang, Qinyu Wang, Xiaofeng Zou, et al., 2025
https://scholar.google.com/scholar?q=DRViT%3A+A+Dynamic+Redundancy-Aware+Vision+Transformer+Accelerator+via+Algorithm+and+Architecture+Co-Design+on+FPGA
14. Realisation of Early-Exit Dynamic Neural Networks on Reconfigurable Hardware — Anastasios Dimitriou, Lei Xun, Jonathon Hare, Geoff V. Merrett, 2024
https://scholar.google.com/scholar?q=Realisation+of+Early-Exit+Dynamic+Neural+Networks+on+Reconfigurable+Hardware
15. Compute-In-Memory on FPGAs for Deep Learning: A Review — Aman Arora, 2025
https://scholar.google.com/scholar?q=Compute-In-Memory+on+FPGAs+for+Deep+Learning%3A+A+Review
16. A Heterogeneous System With Computing in Memory Processing Elements to Accelerate CNN Inference — Jinkai Wang, Youxiang Chen, Zekun Wang, Zhengkun Gu, et al., 2025
https://scholar.google.com/scholar?q=A+Heterogeneous+System+With+Computing+in+Memory+Processing+Elements+to+Accelerate+CNN+Inference
17. AI Post Transformers: FPGA Neural Network Accelerators for Space — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-26-fpga-neural-network-accelerators-for-spa-3087ae.mp3
Interactive Visualization: Automating CNN Mapping on Embedded FPGAs
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof