Data Engineering Podcast

ArangoDB: Fast, Scalable, and Multi-Model Data Storage with Jan Steeman and Jan Stücke - Episode 34


Listen Later

Summary

Using a multi-model database in your applications can greatly reduce the amount of infrastructure and complexity required. ArangoDB is a storage engine that supports documents, dey/value, and graph data formats, as well as being fast and scalable. In this episode Jan Steeman and Jan Stücke explain where Arango fits in the crowded database market, how it works under the hood, and how you can start working with it today.

Preamble
  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
  • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
  • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
  • Your host is Tobias Macey and today I’m interviewing Jan Stücke and Jan Steeman about ArangoDB, a multi-model distributed database for graph, document, and key/value storage.
  • Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you give a high level description of what ArangoDB is and the motivation for creating it?
      • What is the story behind the name?

      • How is ArangoDB constructed?

        • How does the underlying engine store the data to allow for the different ways of viewing it?

        • What are some of the benefits of multi-model data storage?

          • When does it become problematic?

          • For users who are accustomed to a relational engine, how do they need to adjust their approach to data modeling when working with Arango?

          • How does it compare to OrientDB?

          • What are the options for scaling a running system?

            • What are the limitations in terms of network architecture or data volumes?

            • One of the unique aspects of ArangoDB is the Foxx framework for embedding microservices in the data layer. What benefits does that provide over a three tier architecture?

              • What mechanisms do you have in place to prevent data breaches from security vulnerabilities in the Foxx code?
              • What are some of the most interesting or surprising uses of this functionality that you have seen?

              • What are some of the most challenging technical and business aspects of building and promoting ArangoDB?

              • What do you have planned for the future of ArangoDB?

              • Contact Info
                • Jan Steemann
                  • jsteemann on GitHub
                  • @steemann on Twitter

                  • Parting Question
                    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                    • Links

                      • ArangoDB
                      • Köln
                      • Multi-model Database
                      • Graph Algorithms
                      • Apache 2
                      • C++
                      • ArangoDB Foxx
                      • Raft Protocol
                      • Target Partners
                      • RocksDB
                      • AQL (ArangoDB Query Language)
                      • OrientDB
                      • PostGreSQL
                      • OrientDB Studio
                      • Google Spanner
                      • 3-Tier Architecture
                      • Thomson-Reuters
                      • Arango Search
                      • Dell EMC
                      • Google S2 Index
                      • ArangoDB Geographic Functionality
                      • JSON Schema
                      • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                        Support Data Engineering Podcast

                        ...more
                        View all episodesView all episodes
                        Download on the App Store

                        Data Engineering PodcastBy Tobias Macey

                        • 4.6
                        • 4.6
                        • 4.6
                        • 4.6
                        • 4.6

                        4.6

                        135 ratings


                        More shows like Data Engineering Podcast

                        View all
                        Software Engineering Radio - the podcast for professional software developers by se-radio@computer.org

                        Software Engineering Radio - the podcast for professional software developers

                        272 Listeners

                        The Changelog: Software Development, Open Source by Changelog Media

                        The Changelog: Software Development, Open Source

                        283 Listeners

                        The Cloudcast by Massive Studios

                        The Cloudcast

                        153 Listeners

                        Thoughtworks Technology Podcast by Thoughtworks

                        Thoughtworks Technology Podcast

                        41 Listeners

                        Data Skeptic by Kyle Polich

                        Data Skeptic

                        483 Listeners

                        Talk Python To Me by Michael Kennedy

                        Talk Python To Me

                        592 Listeners

                        Software Engineering Daily by Software Engineering Daily

                        Software Engineering Daily

                        624 Listeners

                        The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) by Sam Charrington

                        The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

                        444 Listeners

                        Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

                        Super Data Science: ML & AI Podcast with Jon Krohn

                        298 Listeners

                        Python Bytes by Michael Kennedy and Brian Okken

                        Python Bytes

                        213 Listeners

                        DataFramed by DataCamp

                        DataFramed

                        266 Listeners

                        Practical AI by Practical AI LLC

                        Practical AI

                        190 Listeners

                        The Stack Overflow Podcast by The Stack Overflow Podcast

                        The Stack Overflow Podcast

                        64 Listeners

                        The Real Python Podcast by Real Python

                        The Real Python Podcast

                        140 Listeners

                        Latent Space: The AI Engineer Podcast by swyx + Alessio

                        Latent Space: The AI Engineer Podcast

                        77 Listeners