Weaviate Podcast

Weaviate Podcast

By WeaviateTechnology
Download on the App Store

Weaviate Podcast episodes

  • Patrick Lewis on Retrieval-Augmented Generation - Weaviate Podcast #76!

    Hey everyone, I am SUPER excited to present our 76th Weaviate Podcast featuring Patrick Lewis, an NLP Research Scientist at Cohere! Patrick has had an absolutely massive impact on Natural Language Processing with AI and Deep Learning! Especially notable for the current climate in AI and Weaviate is that Patrick is the lead author of the original "Retrieval-Augmented Generation" paper!! Patrick has contributed to many other profoundly impactful papers in the space as well such as DPR, Atlas, Task-Aware Retrieval with Instruction, and many many others! This was such an illuminating conversation, here is a quick overview of the chapters in the podcast!

    1. Origin of RAG - Patrick explains the build-up that lead to the RAG paper, AskJeeves, IBM Watson, conceptual shift to retrieve-read in mainstream connectionist approaches to AI.
    2. Atlas - Atlas shows that a much smaller LLM when paired with Retrieval-Augmentation can still achieve competitive few-shot and zero-shot task performance. This is super impactful because this few-shot and zero-shot capability has been a massive evangelist for AI broadly, and the fact that smaller Retrieval-Augmented models can do this is massive for the economically unlocking these applications.
    Teasing apart some architectural details of RAG:
    3. Fusion In-Decoder - Interesting encoder-decoder transformer design in which each document + the query is encoded separately, then concatenated and passed to the LM.
    4. End-to-End RAG - How to think about jointly training an embedding model and an LLM augmented with retrieval?
    5. Query Routers - How to route queries from say SQL or Vector DBs? (More nuance on this later with Multi-Index Retrieval)
    6. ConcurrentQA - Super interesting work on the privacy of multi-index routers. For example, if you ask "Who is the father of our new CEO" - this may reveal the private information of the new CEO with the public query of their father.
    7. Multi-Index Retrieval
    8. New APIs for LLMs
    9. Self-Instructed Gorillas
    10. Task-Aware Retrieval with Instructions
    11. Editing Text, EditEval and PEER
    12. What future direction excites you the most?
    Links:
    Learn more about Patrick Lewis: https://www.patricklewis.io/
    Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks: https://arxiv.org/abs/2005.11401
    Atlas: https://arxiv.org/pdf/2208.03299.pdf
    Fusion In-Decoder: https://arxiv.org/pdf/2007.01282.pdf
    Chapters
    0:00 Welcome Patrick Lewis!
    0:36 Origin of RAG
    5:20 Atlas
    10:43 Fusion In-Decoder
    17:50 End-to-End RAG
    27:05 Query Routers
    32:05 ConcurrentQA
    37:30 Multi-Index Retrieval
    40:05 New APIs for LLMs
    41:50 Self-Instructed Gorillas
    44:35 Task-Aware Retrieval with Instructions
    52:00 Editing Text, EditEval and PEER
    55:35 What future direction excites you the most?

    59 min
  • Tanmay Chopra on Emissary - Weaviate Podcast #75!

    Hey everyone! Thank you so much for watching the 75th Weaviate Podcast featuring Tanmay Chopra! The podcast details Tanmay's incredible career in Machine Learning from Tik Tok to Neeva and now building his own startup, Emissary! Tanmay shared some amazing insights into Search AI such as how to process Temporal Queries, how to think about diversity in Retrieval, and Query Recommendation products! We then dove into the opportunity Tanmay sees in fine-tuning LLMs and knowledge distillation that motivated Tanmay to build Emissary! I thought Tanmay's analogy of GPT-4 to 3D printers was really interesting, tons of great nuggets in here! I really hope you enjoy the podcast, as always more than happy to answer any questions or discuss any ideas with you related to the content in the podcast!

    Chapters
    0:00 Welcome Tanmay!
    0:23 Early Career Story
    2:02 Tik Tok
    4:10 Neeva
    8:45 Temporal Queries
    11:40 Retrieval Diversity
    17:22 Query Recommendation
    23:20 Emissary, starting a company!
    30:20 A Simple API for Custom Models
    35:42 GPT-4 = 3D Printer?

    51 min
  • Simba Khadder on FeatureForm - Weaviate Podcast #74!

    Hey everyone! Thank you so much for watching the 74th Weaviate Podcast feature Simba Khadder, the CEO and Co-Founder of FeatureForm! To begin, "features" broadly describe the inputs to machine learning models that they use to produce outputs, or predictions. Feature stores orchestrate the construction of features, whether that be transformations for tabular machine learning models such as XGBoost, to chunking for vector embedding inference, and now features for LLM inference in RAG. Right out of the gate, Simba really opened my eyes to the role that feature engineering plays in RAG. Further touching on this at the very end under the "Exciting future for RAG with Features" chapter, Simba further describes how we can use more advanced features to provide better context to LLMs. In addition to these insights on RAG, there are so many nuggets in the podcast, Simba is a world class professional when it comes to building distributed systems, production scale recommendation systems, and more! I learned so much from chatting with Simba, I hope you enjoy listening to the podcast! As always we are more than happy to answer any questions or discuss any ideas you have about the content in the podcast!

    FeatureForm: https://www.featureform.com/
    Highly Recommend!! Simba Khadder at the CMU DB Seminar series: https://www.youtube.com/watch?v=ZsWa6XiBc-U
    FeatureForm and Weaviate demo! https://docs.featureform.com/providers/weaviate
    Chapters
    0:00 Simba Khadder
    0:35 RAG and Feature Stores
    4:30 Experience building Recommendation Systems
    9:47 The End-to-End Feature Lifecycle
    15:08 Virtual Feature Store Orchestration
    26:45 RAG Evaluation
    31:27 Feature Engineering
    34:15 LLM Tuning and Features
    39:55 Streaming Features
    51:15 Data Drift Detection
    54:20 Exciting future for RAG with Features

    57 min
  • Charles Packer on MemGPT - Weaviate Podcast #73!

    Hey everyone! I am SUPER excited to publish our 73rd Weaviate Podcast with Charles Packer, the lead author of MemGPT at UC Berkeley! MemGPT presents the "Operating System for LLMs", an incredibly exciting idea to explicitly prompt the LLM with the information that it has a limited context window and give it memory management tools to behave accordingly! This was such a fun discussion with Charles diving into all things related to the paper! I hope you enjoy the podcast!!

    Check out MemGPT here! https://memgpt.ai/
    Chapters
    0:00 Welcome Charles!
    0:27 LLM Operating System
    4:47 Memory Management Tools
    6:50 Interrupts in LLM Applications
    10:15 LLM Tools
    17:45 Self-Instruct Data Creation
    20:50 Cost of Experiments
    24:28 Explicit Context Annotation
    29:40 Recall vs. Archival Storage
    33:12 Page Replacement Inspiration
    38:00 Creativity in AI
    43:40 Evolutionary Perspective
    46:18 Inspiring Future Directions
    48:45 Multi-Threaded LLM Processing

    52 min
  • Madelon Hulsebos on Tabular Machine Learning - Weaviate Podcast #72!

    Hey everyone! Thank you so much for watching the 72nd episode of the Weaviate Podcast with Madelon Hulsebos!! Madelon is one of the world's experts on Machine Learning with Tables and Tabular-Structured Data, this was such an eye-opening conversation! We discussed all sorts of topics from the relationship of tabular data and embeddings, to searching through tables, semantic joins, more complex Text-to-SQL, using machine learning for query execution, using tabular data in search and recommendation reranking, and many more! This was easily one of the most knowledge packed episodes of the Weaviate podcast so far, please don't hesitate to leave any questions or ideas you have related to the content discussed!

    You can learn more about Madelon's incredible research career and publications / talks here: https://www.madelonhulsebos.com/! Papers such as GitTables are listed here!
    Another nice nugget form the podcast - Madelon introduced me to the BIRD-SQL benchmark which really expanded my understanding of Text-to-SQL (https://arxiv.org/pdf/2305.03111.pdf.
    Chapters
    0:00 Welcome Madelon!
    0:58 Tabular Data and Embeddings
    3:10 Tabular Representation Learning
    5:48 Semantic Type Detection
    9:50 Pandas as an LLM Tool
    11:52 Table-Based Question Answering and Text-to-SQL
    19:35 Joins with Machine Learning
    21:38 Query Execution with Machine Learning
    22:45 Graph Neural Networks
    24:07 XGBoost
    28:28 Merging Tables
    32:10 Fact Representation
    35:50 GPT-4V and Tables
    39:00 Metadata in Embeddings
    42:45 Table Retrieval in Weaviate
    46:25 Exciting future directions!!

    50 min
  • Vibs Abhishek on Alltius AI - Weaviate Podcast #71!

    Hey everyone! Thank you so much for watching the 71st Weaviate Podcast with Vibs Abhishek! Vibs is the CEO and Founder of Alltius AI, as well as a professor at UC Irvine business school! In order to tame the somewhat chaotic emerging landscape of RAG and LLM applications, Alltius has settled on 3 core pillars of Knowledge, Skills, and Deployment Channels! Vibs further explained how he sees the distinction between Assistants and Agents and many more topics important to Enterprise deployment of RAG applications such as reducing hallucinations and employing classifiers to route skills and knowledge sources! I learned so much from this conversation, I hope you enjoy the podcast!

    Alltius KNO Plus Demo Video: https://www.loom.com/share/fcfe516b75ea4f069b1a8d6a3510fa4c?sid=5f43317f-c20b-4dd9-91d3-2cde993fd91f
    Chapters
    0:00 Welcome Vibs
    0:22 Background
    2:30 Alltius’ UI for Assistants
    7:15 The Knowledge Pillar
    12:05 SQL Router and Intent Management
    14:10 Classifying a Pipeline / Skill
    17:30 Flexibility of Zero-Shot versus Fine-Tuning
    21:00 The Channels Pillar
    23:00 Connecting the Warehouse / Lakehouse
    24:50 Assistant versus Agent
    28:30 MemGPT
    31:25 Offline LLM Research
    35:50 Multi-Agent Role-Playing Assistants
    39:25 From Clicks to Conversations
    44:10 CEO / Professor and Evolution of the Field

    56 min
  • MemGPT Explained!

    Thank you so much for watching our paper summary video on MemGPT! MemGPT is a super exciting new work bridging together concepts in how Operating Systems manage memory and LLMs!

    Links:
    Paper: https://arxiv.org/pdf/2310.08560.pdf
    Andrej Karpathy on Operating Systems and LLMs: https://twitter.com/karpathy/status/1707437820045062561
    Run LLM Podcast with Charles Packer: https://www.youtube.com/watch?v=4aOLxPdx1Dg
    SciPhi: https://github.com/SciPhi-AI/sciphi/tree/main
    Our perspectives on Database Agents that WRITE to Vector Databases: https://weaviate.io/blog/generative-feedback-loops-with-llms
    Chapters
    0:00 Introduction to MemGPT
    2:45 MemGPT Architecture
    6:15 Operating System for LLMs
    11:48 Types of Context and Storage
    15:42 Control Flow
    18:00 Experiments
    22:04 Future Work
    24:46 Personal Takeaways
    30:34 Thank you for watching!

    31 min
  • Kevin Cohen on Neum AI - Weaviate Podcast #70!

    Hey everyone! Thank you so much for watching the 70th episode of the Weaviate podcast with Neum AI CTO and Co-Founder Kevin Cohen! I first met Kevin when he was debugging an issue with his distributed node utilization and have since learned so much from him about how he sees the space of Data Ingestion, also commonly referenced as ETL for LLMs! There are so many interesting parts to this from the general flow of data connectors, chunkers and metadata extractors, embedding inference, and the last leg of the mile of importing the vectors to a Vector DB such as Weaviate! I really loved how Kevin broke down the distributed messaging queue and system design for orchestrating data ingestion at massive scale such as dealing with failures and optimizing the infrastructure as code setup. We also discussed things like new use cases with quadrillion scale vector indexes and the role of knowledge graphs in all this! I really hope you enjoy the podcast, please check out this amazing article below from Neum AI!

    https://medium.com/@neum_ai/retrieval-augmented-generation-at-scale-building-a-distributed-system-for-synchronizing-and-eaa29162521
    Chapters
    0:00 Check this out!
    1:18 Welcome Kevin!
    1:58 Founding Neum AI
    6:55 Data Ingestion, End-to-End Overview
    9:10 Chunking and Metadata Extraction
    14:20 Embedding Cache
    16:57 Distributed Messaging Queues
    22:15 Embeddings Cache ELI5
    25:30 Customizing Weaviate Kubernetes
    38:10 Multi-Tenancy and Resource Allocation
    39:20 Billion-Scale Vector Search
    45:05 Knowledge Graphs
    52:10 Y Combinator Experience

    56 min
  • Charles Pierse on Tactic Generate - Weaviate Podcast #69!

    Hey everyone! Thank you so much for watching the 69th episode of the Weaviate Podcast featuring Charles Pierse from Tactic! Tactic has recently launched their new Tactic Generate project, an incredible UI for conducting research across multiple documents. I think there is a massive opportunity to pair these prompts and LLM workflows with User Interfaces and take more of a holistic User Experience perspective. Tactic Generate has done an incredible job of that, please take a look from the link below! I had such a fun conversation catching up with Charles (Charles was our 2nd Weaviate Podcast guest!), I hope you enjoy the podcast!

    Tactic Generate: https://tactic.fyi/generative-insights/
    Chapters
    0:00 Tactic Generate
    1:40 Welcome Charles!
    2:38 Charles’ work at Tactic
    4:40 LLMs comparing documents
    9:10 LLM Chaining
    17:30 Discovering LLM Chains
    20:28 Moats in ML Products
    28:48 Fine-Tuning vs. RAG
    34:30 Fine-Tuning Search Models
    39:45 Skepticism on RLHF
    41:52 Gorilla, Integrations, and CRM
    45:40 Query Routers
    47:55 CRM and Tree-of-Thoughts
    55:54 Graph Embeddings
    1:02:20 Llama CPP / GGML
    1:04:28 What are you looking forward to most in AI?

    1 hr 9 min
  • Weights and Biases on Fine-Tuning LLMs - Weaviate Podcast #68!

    Hey everyone! Thank you so much for watching the 68th episode of the Weaviate Podcast! We are super excited to welcome Morgan McGuire, Darek Kleczek, and Thomas Capelle! This was such a fun discussion beginning with generally how see the space of fine-tuning from why you would want to do it, to the available tooling, intersection with RAG and more!

    Check out W&B Prompts! https://wandb.ai/site/prompts
    Check out the W&B Tiny Llama Report! https://wandb.ai/capecape/llamac/reports/Training-Tiny-Llamas-for-Fun-and-Science--Vmlldzo1MDM2MDg0
    Chapters
    0:00 Tiny Llamas!
    1:53 Welcome!
    2:22 LLM Fine-Tuning
    5:25 Tooling for Fine-Tuning
    7:55 Why Fine-Tune?
    9:55 RAG vs. Fine-Tuning
    12:25 Knowledge Distillation
    14:40 Gorilla LLMs
    18:25 Open-Source LLMs
    22:48 Jonathan Frankle on W&B
    23:45 Data Quality for LLM Training
    25:55 W&B for Data Versioning
    27:25 Curriculum Learning
    29:28 GPU Rich and Data Quality
    30:30 Vector DBs and Data Quality
    32:50 Tuning Training with Weights & Biases
    35:47 Training Reports
    42:28 HF Collections and W&B Sweeps
    44:50 Exciting Directions for AI

    53 min

About Weaviate Podcast

From the publisher's feed

Join Connor Shorten as he interviews machine learning experts and explores Weaviate use cases from users and customers.