Weaviate Podcast

Weaviate Podcast

By WeaviateTechnology
Download on the App Store

Weaviate Podcast episodes

  • XMC.dspy with Karel D'Oosterlinck - Weaviate Podcast #87!

    Hey everyone! Thank you so much for watching the 87th episode of the Weaviate Podcast! I am SUPER excited to welcome Karel D'Oosterlinck! Karel is the creator of IReRa (Infer-Retrieve-Rank)! IReRa is one of the most impressive systems that have been built for Extreme Multi-Label Classification, leveraging the emerging paradigm of DSPy compilation! This podcast dives into all things IReRa, XMC, DSPy compilation, and applications in Biomedical NLP and Recommendation! I hope you find this useful!

    1 hr 9 min
  • Open-Source AI with Vinod Valloppillil and Bob van Luijt - Weaviate Podcast #86!

    Hey everyone! We are super excited to publish this podcast with Vinod Valloppillil and Bob van Luijt on Open-Source AI and future directions for RAG! The podcast begins by discussing Vinod's "Halloween Documents", a series of internal strategy writings at Microsoft related to the open-source software movement! The conversation continues to discuss the current state of Open-Source in AI. One of the major points Bob has been making about the business of AI models is that the models themselves are *stateless*, akin to an MP3 file. Vinod pushes back a bit on this definition and jointly it is then settled that these models neither fall into the pure stateful or stateless bucket, rather a "pre-baked" bucket -- presenting completely new opportunities to build business around software. The conversation then continues to discuss the particular details of how people are building RAG systems and many directions for how that may evolve!

    56 min
  • DSPy and ColBERT with Omar Khattab! - Weaviate Podcast #85

    Hey everyone! I am beyond excited to present our interview with Omar Khattab from Stanford University! Omar is one of the world's leading scientists on AI and NLP. I highly recommend you check out Omar's remarkable list of publications linked below! This interview completely transformed my understanding of building RAG and LLM applications! I believe that DSPy will be one of the most impactful software project in LLM development because of the abstractions around *program optimization*. Here is my TLDR of this concept of LLM programs and program optimization with DSPy, I of course encourage you to view the podcast and listen to Omar's explanation haha.

    RAG is one of the most popular LLM programs we have seen. RAG typically consists of two components of retrieve and then generate. Within the generate component we have a prompt like "please ground your answer based on the search results {search_results}". DSPy gives us a framework to optimize this prompt, bootstrap few-shot examples, or even fine-tune the model if needed. This works by compiling the program based on some evaluation criteria we give DSPy. Now let's say we add a query re-writer that takes the query and writes a new query before sending it to the retrieval system, and a reranker that takes the search results and re-orders them before handing them to the answer generator. Now we have 4 components of query writer, retrieve, rerank, answer. The 3 components of query writer, rerank, and answer all have a prompt that can be optimized with DSPy to enhance the description of the task or add examples! This optimization is done with DSPy's Teleprompters.
    There are a few other really interesting components to DSPy as well -- such as the formatting of prompts with the docstrings and Signature abstraction, which in my view is quite similar to instructor or LMQL. DSPy also comes with built-in prompts like Chain-of-Thought that offer a really quick way to add this reasoning step and follow a structured output format. I am having so much fun learning about DSPy and I highly recommend you join me in viewing the GitHub repository linked below (with new examples!!):
    Omar also discusses ColBERT and late interaction retrieval! Omar describes how this achieves the contextualized attention of cross encoders but in a much more scalable system with the maximum similarity between vectors! Stay tuned for more updates from Weaviate as we are diving into multi vector representations to hopefully support systems like this soon!


    Chapters

    0:00 Weaviate at NeurIPS 2023!

    0:38 Omar Khattab

    0:57 What is the state of AI?

    2:35 DSPy

    10:37 Pipelines

    14:24 Prompt Tuning and Optimization

    18:12 Models for Specific Tasks

    21:44 LLM Compiler

    23:32 Colbert or ColBERT?

    24:02 ColBERT

    32 min
  • Subjectivity in AI with Dan Shipper: AI-Native Databases #4

    Hey everyone! Thank you so much for watching the fourth and final episode of the AI-Native Database series with Dan Shipper! This was another epic one! Dan has had an absolutely remarkable career creating and selling a company and now co-founding and working as the CEO of Every! Every is an incredibly future-looking business focused on content online, both with an amazing newsletter, community of writers and thinkers, an AI-note taking app, and more! I think Dan brings a very unique perspective to the series, as well as the Weaviate podcast broadly, because of his experience with writers and understanding how writers are going to use these new technologies! We heavily discussed the role of personality or subjectivity in AI, amongst many other topics! I really hope you enjoy the podcast, as always we are more than happy to answer any questions or discuss any ideas you have about the content in the podcast!

    Read writings from Dan Shipper on Every: https://every.to/@danshipper
    Chapters
    0:00 AI-Native Databases
    0:58 Welcome Dan Shipper!
    1:37 GPT-4 is a Reasoning Engine
    8:40 Subjectivity in LLMs
    12:14 AI in Note Taking
    16:38 The opinions of LLMs
    25:50 Cookbooks for you
    31:16 Overdrive in LLMs
    34:50 Tweaking the voice of AI
    40:45 Multi-Agent Personalities

    43 min
  • Humans and AI with John Maeda: AI-Native Databases #3

    Hey everyone! Thank you so much for watching the 3rd episode of the AI-Native Database series featuring John Maeda and Bob van Luijt! This one dives into how humans perceive AI, from Anthroaormorphization to Doomsday scenario thinking and how important understanding how AI actually work is to the engineering of these systems. Bob and John discuss the evolution of the design in tech report, 3 categories of design, and many others! I hope you enjoy the podcast! As always, we are more than happy to answer any questions or discuss any ideas you have about the content in the podcast!

    Links:
    Design in Tech Report: https://designintech.report/
    3 Kinds of Design: https://qz.com/1585165/john-maeda-on-the-importance-of-computational-design
    Microsoft Semantic Kernel: https://github.com/microsoft/semantic-kernel
    Chapters
    0:00 AI-Native Databases
    0:58 Welcome John Maeda!
    1:35 Design in Tech Report
    4:07 Anthropomorphizing AI
    15:30 3 Types of Design
    19:30 The ChatGPT Shift
    22:58 Explaining Technology
    32:54 Impact of AI on the Creative Industries
    39:00 Semantic Kernel

    41 min
  • Structure in Data with Paul Groth: AI-Native Databases #2

    Hey everyone! Thank you so much for watching the second episode of AI-Native Databases with Paul Groth! This was another epic one, diving deep into the role of structure in our data! Beginning with Knowledge Graphs and LLMs, there are two perspectives: LLMs for Knowledge Graphs (using LLMs to extract relationships or predict missing links) and then Knowledge Graph for LLMs (to provide factual information in RAG). There is another intersection that sits in the middle of both LLMs for KGs and KGs for LLMs, which is using LLMs to query Knowledge Graphs, e.g. Text-to-Cypher/SPARQL/... From there I think the conversation evolves in a really fascinating way exploring the ability to structure data on-the-fly. Paul says "Unstructured data is now becoming a peer to structured data"! I think in addition to RAG, Generative Search is another underrated use case -- where we use LLMs to summarize search results or parse out the structure. Super interesting ideas, I hope you enjoy the podcast -- as always more than happy to answer any questions or discuss any ideas you have about the content in the podcast!

    Learn more about Professor Groth's research here: https://scholar.google.com/citations?...
    Knowledge Engineering using Large Language Models: https://arxiv.org/pdf/2310.00637.pdf
    How Much Knowledge Can You Pack into the Parameters of a Language Model? https://arxiv.org/abs/2002.08910
    Chapters
    0:00 AI-Native Databases!
    0:58 Welcome Paul!
    1:25 Bob’s overview of the series
    2:30 How do we build great datasets?
    4:28 Defining Knowledge Graphs
    7:15 LLM as a Knowledge Graph
    15:18 Adding CRUD Support to Models
    28:10 Database of Model Weights
    32:50 Structuring Data On-the-Fly

    46 min
  • Self-Driving Databases with Andy Pavlo: AI-Native Databases #1

    Hey everyone! Thank you so much for watching the first episode of AI-Native Databases with Andy Pavlo! This was an epic one! We began by explaining the "Self-Driving Database" and all the opportunities to optimize DBs with AI and ML at both the low-level, as well as how we query and interact with them. We also discussed new opportunities with DBs + LLMs, such as bringing the data to the model (such as ROME, MEMIT, GRACE), in addition to bringing the model to the data (such as RAG). We also discuss the subjective "opinion" of these models and many more!

    I hope you enjoy the podcast! As always we are more than happy to answer any questions or discuss any ideas you have about the content in the podcast! This one means a lot to me. Andy Pavlo's CMU DB course was one of the most impactful resources in my personal education, and I love the vision for the future outlined by OtterTune! It was amazing to see Etienne Dilocker featured in the ML for DBs, DBs for ML series at CMU. I am so grateful to Andy for joining the Weaviate Podcast!
    Links:
    CMU Database Group on YouTube: https://www.youtube.com/@CMUDatabaseGroup/videos
    Self-Driving Database Management Systems - Pavlo et al. - https://db.cs.cmu.edu/papers/2017/p42-pavlo-cidr17.pdf
    Database of Databases: https://dbdb.io/
    Generative Feedback Loops: https://weaviate.io/blog/generative-feedback-loops-with-llms
    Weaviate Gorilla: https://weaviate.io/blog/weaviate-gorilla-part-1
    Chapters
    0:00 AI-Native Databases
    0:58 Welcome Andy
    1:58 Bob’s overview of the series
    3:20 Self-Driving Databases
    8:18 Why isn’t there just 1 Database?
    12:46 Collaboration of Models and Databases
    20:05 LLM Schema Tuning
    23:44 The Opinion of the System
    28:20 PyTorchDB - Moving the Data to the Model
    33:30 Database APIs
    38:15 Learning to operate Databases
    42:54 Vector DBs and the DB Hype Cycle
    51:38 SQL in Weaviate?
    1:07:40 The Future of DBs
    1:14:00 Thank you Andy!

    1 hr 15 min
  • Weaviate 1.23 Release Podcast with Etienne Dilocker!

    Hey everyone! Thank you so much for watching the Weaviate 1.23 Release Podcast with Weaviate Co-Founder and CTO Etienne Dilocker! Weaviate 1.23 is a massive step forward for managing multi-tenancy with vector databases. For most RAG and Vector DB applications, you will have an uneven distribution in the # of vectors per user. Some users have 10k docs, others 10M+! Weaviate now offers a flat index with binary quantization to efficiently balance when you need an HNSW graph for the 10M doc users and when brute force is all you need for the 10k doc users!

    Weaviate also comes with some other "self-driving database" features like lazy shard loading for faster startup times with multi-tenancy and automatic resource limiting with the GOMEMLIMIT and other details Etienne shares in the podcast!
    I am also beyond excited to present our new integration with Anyscale (@anyscalecompute)! Anyscale has amazing pricing for serving and fine-tuning popular open-source LLMs. At the time of this release we are now integrating the Llama 70B/13B/7B, Mistral 7B, and Code Llama 34B into Weaviate -- but we expect much further development with adding support for fine-tuned models, the super cool new function calling models Anyscale announced yesterday. and other model such as Diffusion and multimodal models!
    Chapters
    0:00 Weaviate 1.23
    1:08 Lazy Shard Loading
    8:20 Flat Index + BQ
    33:15 Default Segments for PQ
    38:55 AutoPQ
    42:20 Auto Resource Limiting
    46:04 Node Endpoint Update
    47:25 Generative Anyscale
    Links:
    Etienne Dilocker on Native Multi-Tenancy at the AI Conference in SF:
    https://www.youtube.com/watch?v=KT2RFMTJKGs
    Etienne Dilocker in the CMU DB Series:
    https://www.youtube.com/watch?v=4sLJapXEPd4
    Self-Driving Databases by Andy Pavlo: https://www.cs.cmu.edu/~pavlo/blog/2018/04/what-is-a-self-driving-database-management-system.html

    56 min
  • Rudy Lai on Tactic Generate - Weaviate Podcast #78!

    Hey everyone! Thank you so much for watching the 78th episode of the Weaviate podcast featuring Rudy Lai, the founder and CEO of Tactic Generate! Tactic Generate has developed a user experience around applying LLMs in parallel to multiple documents, or even folders / collections / databases. Rudy discussed the user research that lead the company to this direction and how he sees the opportunities in building AI products with new LLM and Vector Database technologies! I hope you enjoy the podcast, as always more than happy to answer any questions or discuss any ideas you have about the content in the podcast!

    Learn more about Tactic Generate here: https://tactic.fyi/generative-insights/
    Weaviate Podcast #69 with Charles Pierse: https://www.youtube.com/watch?v=L_nyz1xs9AU
    Chapters
    0:00 Welcome Rudy!
    0:48 Story of Tactic Generate
    7:45 Finding Common Workflows
    19:30 Multiple Document RAG UIs
    26:14 Parallel LLM Execution
    32:40 Aggregating Parallel LLM Analysis
    38:25 Pretty Reports
    44:28 Research Agents

    57 min
  • RAGAS with Jithin James, Shahul Es, and Erika Cardenas - Weaviate Podcast #77!

    Hey everyone, thank you so much for watching the 77th Weaviate Podcast on RAGAS, featuring Jithin James, Shahul ES, and Erika Cardenas! RAGAS is one of the hottest rising startups in Retrieval-Augmented Generation! RAGAS began it's journey with the RAGAS score, a matrix of evaluations for generation and retrieval. Generation evaluated on Faithfulness (is the response grounded in the context) as well as Relevancy (is the response useful). Retrieval is then evaluated on Precision (How many of the search results are relevant to the question?) and Recall (How many of the relevant search results are captured in the retrieved results?). Now, the super novel thing about this is that an LLM is used to determine these metrics. So we circumvent painstaking manual labeling effort with the RAGAS score! This podcast dives into the development of the RAGAS score as well as how RAG application builders should think about the knobs to tune for optimizing their RAGAS score: embedding models, chunking strategies, hybrid search tuning, rerankers, ... ?!? We also discussed tons of exciting directions for the future such as fine-tuning smaller LLMs for these metrics, agents that use tuning APIs, and long context RAG!

    Check out the docs here for getting started with RAGAS! https://docs.ragas.io/en/latest/getstarted/index.html#get-started
    Chapters
    0:00 Welcome Jithin and Shahul!
    0:44 Welcome Erika!
    0:56 RAGAS, Founding Story
    2:38 Weaviate + RAGAS integration plans
    4:44 RAG Knobs to Tune
    25:50 RAG Experiment Tracking
    34:52 LangSmith and RAGAS
    38:55 LLM Evaluation
    40:25 RAGAS Agents
    44:00 Long Context RAG Evaluation

    50 min

About Weaviate Podcast

From the publisher's feed

Join Connor Shorten as he interviews machine learning experts and explores Weaviate use cases from users and customers.