
Sign up to save your podcasts
Or


In this episode, we delve into the paper "TransAct: Transformer-based Realtime User Action Model for Recommendation at Pinterest" . This research introduces TransAct, a novel Transformer-based model designed to enhance Pinterest's recommendation system by capturing users' short-term preferences through their real-time activities.
Research Paper Link - arxiv.org+4arxiv.org+4export.arxiv.org+4
🔹 What’s Inside?
Tune in to learn how TransAct balances real-time responsiveness with efficiency in large-scale AI-driven personalization. 🚀
In today’s episode, we’re diving into the fascinating world of model merging—a technique that allows multiple AI models to be combined, often enhancing their capabilities without the need for costly retraining. Our focus? A groundbreaking paper titled "Do Merged Models Copy or Compose? Evaluating the Transfer of Capabilities in Model Merging" by researchers exploring the inner workings of this emerging technique.
We'll be discussing:
🔹 What is model merging? Why it's gaining traction in AI research.
🔹 Do merged models simply copy knowledge, or can they create something new?
🔹 How does merging affect generalization, robustness, and performance?
🔹 Real-world implications—from adapting models across different domains to fine-tuning AI with fewer resources.
In this episode, we delve into the transformative impact of Generative Models on modern Recommender Systems (RS), as detailed in the comprehensive survey titled "A Review of Modern Recommender Systems Using Generative Models (Gen-RecSys)". This multidisciplinary study explores how traditional RS, which primarily relied on user-item rating histories, are evolving through the integration of advanced generative techniques.
Key Discussion Points:
Join us as we explore these advancements, shedding light on the future directions of recommender systems in the era of generative AI.
Research paper: https://arxiv.org/pdf/2502.04677
Authors: Gregory Dexter, Shao Tang, Ata Fatahi Baarzi, Qingquan Song, Tejas Dharamsi, and Aman Gupta
Introduction
In this episode, we explore the challenge of efficiently deploying large language models (LLMs) in online settings, where strict latency constraints—such as time-to-first-token (TTFT) and time-per-output-token (TPOT)—must be met. As demand for AI-generated content grows, optimizing inference performance becomes a critical bottleneck.
Key Topics Covered
Conclusion
This research highlights the need for advanced scheduling strategies to improve LLM efficiency in real-world applications. Tune in to learn how k-LPM is pushing the boundaries of AI inference optimization!
In this episode, we explore Meta's ACH system, a novel mutation-guided test generation approach that leverages LLMs (Large Language Models) to enhance software robustness. Unlike traditional mutation testing, which generates numerous random faults, ACH focuses on identifying undetected faults related to specific concerns, such as privacy vulnerabilities.
🔍 Key Highlights:
🔎 Why It Matters: ACH represents a paradigm shift in mutation testing, using AI to pinpoint real-world vulnerabilities instead of generating irrelevant noise. This approach not only improves software reliability but also streamlines engineering workflows by focusing on actionable test cases.
🔗 Reference Paper: 📄 Meta’s ACH System for Mutation-Guided LLM-Based Test Generation – Read here
📢 Tune in as we break down how ACH is redefining software testing, enhancing privacy safeguards, and paving the way for AI-driven quality assurance! 🚀
Ranking and recommendation systems are the foundation for numerous online experiences, ranging from search results to personalized content delivery. These systems have evolved into complex, multilayered architectures that leverage vast datasets and often incorporate thousands of predictive models. The maintenance and enhancement of these models is a labor intensive process that requires extensive feature engineering. This approach not only exacerbates technical debt but also hampers innovation in extending these systems to emerging problem domains. In this report, we present our research to address these challenges by utilizing a large foundation model with a textual interface for ranking and recommendation tasks. We illustrate several key advantages of our approach: (1) a single model can manage multiple predictive tasks involved in ranking and recommendation, (2) decoder models with textual interface due to their comprehension of reasoning capabilities, can generalize to new recommendation surfaces and out-of-domain problems, and (3) by employing natural language interfaces for task definitions and verbalizing member behaviors and their social connections, we eliminate the need for feature engineering and the maintenance of complex directed acyclic graphs of model dependencies. We introduce our research pre-production model, 360Brew V1.0, a 150B parameter, decoder-only model that has been trained and fine-tuned on LinkedIn's data and tasks. This model is capable of solving over 30 predictive tasks across various segments of the LinkedIn platform, achieving performance levels comparable to or exceeding those of current production systems based on offline metrics, without task-specific fine-tuning. Notably, each of these tasks is conventionally addressed by dedicated models that have been developed and maintained over multiple years by teams of a similar or larger size than our own.
From the publisher's feed
Welcome to Paper Bytes, where we distill cutting-edge research papers into bite-sized, engaging audio episodes! Our mission is to bring complex innovations to life, making them accessible to…