Data Engineering Podcast

Stateful, Distributed Stream Processing on Flink with Fabian Hueske - Episode 57


Listen Later

Summary

Modern applications and data platforms aspire to process events and data in real time at scale and with low latency. Apache Flink is a true stream processing engine with an impressive set of capabilities for stateful computation at scale. In this episode Fabian Hueske, one of the original authors, explains how Flink is architected, how it is being used to power some of the world’s largest businesses, where it sits in the lanscape of stream processing tools, and how you can start using it today.

Preamble
  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
  • When you’re ready to build your next pipeline, or want to test out the projects you hear about on the show, you’ll need somewhere to deploy it, so check out Linode. With 200Gbit private networking, scalable shared block storage, and a 40Gbit public network, you’ve got everything you need to run a fast, reliable, and bullet-proof data platform. If you need global distribution, they’ve got that covered too with world-wide datacenters including new ones in Toronto and Mumbai. Go to dataengineeringpodcast.com/linode today to get a $20 credit and launch a new server in under a minute.
  • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the mailing list, read the show notes, and get in touch.
  • Join the community in the new Zulip chat workspace at dataengineeringpodcast.com/chat
  • Your host is Tobias Macey and today I’m interviewing Fabian Hueske, co-author of the upcoming O’Reilly book Stream Processing With Apache Flink, about his work on Apache Flink, the stateful streaming engine
  • Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you start by describing what Flink is and how the project got started?
    • What are some of the primary ways that Flink is used?
    • How does Flink compare to other streaming engines such as Spark, Kafka, Pulsar, and Storm?
      • What are some use cases that Flink is uniquely qualified to handle?

      • Where does Flink fit into the current data landscape?

      • How is Flink architected?

        • How has that architecture evolved?
        • Are there any aspects of the current design that you would do differently if you started over today?

        • How does scaling work in a Flink deployment?

          • What are the scaling limits?
          • What are some of the failure modes that users should be aware of?

          • How is the statefulness of a cluster managed?

            • What are the mechanisms for managing conflicts?
            • What are the limiting factors for the volume of state that can be practically handled in a cluster and for a given purpose?
            • Can state be shared across processes or tasks within a Flink cluster?

            • What are the comparative challenges of working with bounded vs unbounded streams of data?

            • How do you handle out of order events in Flink, especially as the delay for a given event increases?

            • For someone who is using Flink in their environment, what are the primary means of interacting with and developing on top of it?

            • What are some of the most challenging or complicated aspects of building and maintaining Flink?

            • What are some of the most interesting or unexpected ways that you have seen Flink used?

            • What are some of the improvements or new features that are planned for the future of Flink?

            • What are some features or use cases that you are explicitly not planning to support?

            • For people who participate in the training sessions that you offer through Data Artisans, what are some of the concepts that they are challenged by?

              • What do they find most interesting or exciting?

              • Contact Info
                • LinkedIn
                • @fhueske on Twitter
                • fhueske on GitHub
                • Parting Question
                  • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                  • Links
                    • Flink
                    • Data Artisans
                    • IBM
                    • DB2
                    • Technische Universität Berlin
                    • Hadoop
                    • Relational Database
                    • Google Cloud Dataflow
                    • Spark
                    • Cascading
                    • Java
                    • RocksDB
                    • Flink Checkpoints
                    • Flink Savepoints
                    • Kafka
                    • Pulsar
                    • Storm
                    • Scala
                    • LINQ (Language INtegrated Query)
                    • SQL
                    • Backpressure
                    • Watermarks
                    • HDFS
                    • S3
                    • Avro
                    • JSON
                    • Hive Metastore
                    • Dell EMC
                    • Pravega
                    • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                      Support Data Engineering Podcast

                      ...more
                      View all episodesView all episodes
                      Download on the App Store

                      Data Engineering PodcastBy Tobias Macey

                      • 4.5
                      • 4.5
                      • 4.5
                      • 4.5
                      • 4.5

                      4.5

                      142 ratings


                      More shows like Data Engineering Podcast

                      View all
                      The Changelog: Software Development, Open Source by Changelog Media

                      The Changelog: Software Development, Open Source

                      289 Listeners

                      Software Engineering Daily by Software Engineering Daily

                      Software Engineering Daily

                      623 Listeners

                      Talk Python To Me by Michael Kennedy

                      Talk Python To Me

                      583 Listeners

                      Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

                      Super Data Science: ML & AI Podcast with Jon Krohn

                      302 Listeners

                      NVIDIA AI Podcast by NVIDIA

                      NVIDIA AI Podcast

                      342 Listeners

                      Practical AI by Practical AI LLC

                      Practical AI

                      203 Listeners

                      AWS Podcast by Amazon Web Services

                      AWS Podcast

                      205 Listeners

                      Last Week in AI by Skynet Today

                      Last Week in AI

                      305 Listeners

                      Dwarkesh Podcast by Dwarkesh Patel

                      Dwarkesh Podcast

                      521 Listeners

                      The Data Engineering Show by The Firebolt Data Bros

                      The Data Engineering Show

                      8 Listeners

                      No Priors: Artificial Intelligence | Technology | Startups by Conviction

                      No Priors: Artificial Intelligence | Technology | Startups

                      130 Listeners

                      Latent Space: The AI Engineer Podcast by swyx + Alessio

                      Latent Space: The AI Engineer Podcast

                      92 Listeners

                      This Day in AI Podcast by Michael Sharkey, Chris Sharkey

                      This Day in AI Podcast

                      228 Listeners

                      The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

                      The AI Daily Brief: Artificial Intelligence News and Analysis

                      631 Listeners

                      AI + a16z by a16z

                      AI + a16z

                      36 Listeners