AI Post Transformers

Benchmarking PEFT Techniques for Large Language Models


Listen Later

Hal Turing and Dr. Ada Shannon examine an empirical study of parameter-efficient fine-tuning for large language models, centered on a practical question: when does a small task-specific update beat retraining the entire model? Using FLAN-T5-XL as the test bed, they frame PEFT as a transfer-learning strategy that freezes most of the transformer while learning a compact adaptation layer, whether through LoRA’s low-rank weight updates, adapter-style modules, IA3 scaling vectors, BitFit bias updates, or learned soft prompts. The discussion keeps returning to the real systems tradeoff: quality matters, but so do training speed, storage cost, and the burden of maintaining separate model copies for many downstream tasks.
The episode walks through the benchmark design in detail rather than treating PEFT as a vague category. The paper compares full tuning, LoRA, IA3, prompt tuning, and BitFit on the same backbone across classification tasks like AG News and CoLA, generation tasks like E2E and SAMSum, and data budgets of roughly 100, 1,000, and 10,000 examples. The hosts emphasize why those controls matter: same model, same stopping rule, and fixed method settings make it easier to see where each technique actually helps, while also limiting how far the results should be generalized to other architectures, especially decoder-only chat models.
They then dig into the paper’s uneven but useful results. In low-resource settings, LoRA and BitFit frequently outperform full tuning, with LoRA posting a notably stronger CoLA score and BitFit leading on AG News, E2E, and SAMSum, while prompt tuning performs strikingly poorly on the generation benchmarks under this setup. In medium-resource settings, IA3, LoRA, and BitFit remain competitive, but at higher data scales full tuning starts reclaiming ground on some tasks even as LoRA and IA3 still win specific cases. The takeaway is not that one PEFT method universally dominates, but that the strengths and weaknesses of each approach shift with task type, data regime, and the exact adaptation recipe.
Sources:
1. Empirical Analysis of the Strengths and Weaknesses of PEFT Techniques for LLMs — George Pu, Anirudh Jain, Jihan Yin, Russell Kaplan, 2023
http://arxiv.org/abs/2304.14999
2. Parameter-Efficient Transfer Learning for NLP — Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, et al., 2019
https://arxiv.org/abs/1902.00751
3. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, et al., 2021
https://arxiv.org/abs/2106.09685
4. Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning — Vladislav Lialin, Vijeta Deshpande, Xiaowei Yao, Anna Rumshisky, 2023
https://arxiv.org/abs/2303.15647
5. Empirical Analysis of the Strengths and Weaknesses of PEFT Techniques for LLMs — George Pu, Anirudh Jain, Jihan Yin, Russell Kaplan, 2023
https://arxiv.org/abs/2304.14999
6. Prefix-Tuning: Optimizing Continuous Prompts for Generation — Xiang Lisa Li, Percy Liang, 2021
https://arxiv.org/abs/2101.00190
7. The Power of Scale for Parameter-Efficient Prompt Tuning — Brian Lester, Rami Al-Rfou, Noah Constant, 2021
https://arxiv.org/abs/2104.08691
8. P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks — Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Lam Tam, Zhengxiao Du, Zhilin Yang, Jie Tang, 2021
https://arxiv.org/abs/2110.07602
9. SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer — Tu Vu, Brian Lester, Noah Constant, Rami Al-Rfou, Daniel Cer, 2021
https://arxiv.org/abs/2110.07904
10. Revisiting Parameter-Efficient Tuning: Are We Really There Yet? — Guanzheng Chen, Fangyu Liu, Zaiqiao Meng, Shangsong Liang, 2022
https://scholar.google.com/scholar?q=Revisiting+Parameter-Efficient+Tuning%3A+Are+We+Really+There+Yet%3F
11. Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning — Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, Colin Raffel, 2022
https://scholar.google.com/scholar?q=Few-Shot+Parameter-Efficient+Fine-Tuning+is+Better+and+Cheaper+than+In-Context+Learning
12. On the Effectiveness of Adapter-based Tuning for Pretrained Language Model Adaptation — Ruidan He, Linlin Liu, Hai Ye, Qingyu Tan, Bosheng Ding, Liying Cheng, Jia-Wei Low, Lidong Bing, Luo Si, 2021
https://scholar.google.com/scholar?q=On+the+Effectiveness+of+Adapter-based+Tuning+for+Pretrained+Language+Model+Adaptation
13. Scaling Instruction-Finetuned Language Models — Hyung Won Chung et al., 2022
https://scholar.google.com/scholar?q=Scaling+Instruction-Finetuned+Language+Models
14. When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method — Biao Zhang, Zhongtao Liu, Colin Cherry, Orhan Firat, 2024
https://arxiv.org/abs/2402.17193
15. Parameter-Efficient Fine-Tuning Design Spaces — Jiaao Chen, Aston Zhang, Xingjian Shi, Mu Li, Alex Smola, Diyi Yang, 2023
https://arxiv.org/abs/2301.01821
16. AI Post Transformers: ForkKV for Multi-LoRA Agent Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-forkkv-for-multi-lora-agent-serving-ccafa4.mp3
17. AI Post Transformers: Mooncake for KV Cache-Centric LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-mooncake-for-kv-cache-centric-llm-servin-1086d0.mp3
18. AI Post Transformers: ZeRO-Offload: Democratizing Billion-Scale Model Training — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/zero-offload-democratizing-billion-scale-model-training/
Interactive Visualization: Benchmarking PEFT Techniques for Large Language Models
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof