Data Engineering Podcast

Building A Real Time Event Data Warehouse For Sentry


Listen Later

Summary

The team at Sentry has built a platform for anyone in the world to send software errors and events. As they scaled the volume of customers and data they began running into the limitations of their initial architecture. To address the needs of their business and continue to improve their capabilities they settled on Clickhouse as the new storage and query layer to power their business. In this episode James Cunningham and Ted Kaemming describe the process of rearchitecting a production system, what they learned in the process, and some useful tips for anyone else evaluating Clickhouse.

Announcements
  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
  • When you’re ready to build your next pipeline, or want to test out the projects you hear about on the show, you’ll need somewhere to deploy it, so check out our friends at Linode. With 200Gbit private networking, scalable shared block storage, and a 40Gbit public network, you’ve got everything you need to run a fast, reliable, and bullet-proof data platform. If you need global distribution, they’ve got that covered too with world-wide datacenters including new ones in Toronto and Mumbai. And for your machine learning workloads, they just announced dedicated CPU instances. Go to dataengineeringpodcast.com/linode today to get a $20 credit and launch a new server in under a minute. And don’t forget to thank them for their continued support of this show!
  • You listen to this show to learn and stay up to date with what’s happening in databases, streaming platforms, big data, and everything else you need to know about modern data management. For even more opportunities to meet, listen, and learn from your peers you don’t want to miss out on this year’s conference season. We have partnered with organizations such as O’Reilly Media, Dataversity, Corinium Global Intelligence, Alluxio, and Data Council. Go to dataengineeringpodcast.com/conferences to learn more about these and other events, and take advantage of our partner discounts to save money when you register today.
  • Your host is Tobias Macey and today I’m interviewing Ted Kaemming and James Cunningham about Snuba, the new open source search service at Sentry implemented on top of Clickhouse
  • Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you start by describing the internal and user-facing issues that you were facing at Sentry with the existing search capabilities?
      • What did the previous system look like?
      • What was your design criteria for building a new platform?
        • What was your initial list of possible system components and what was your evaluation process that resulted in your selection of Clickhouse?
        • Can you describe the system architecture of Snuba and some of the ways that it differs from your initial ideas of how it would work?
          • What have been some of the sharp edges of Clickhouse that you have had to engineer around?
          • How have you found the operational aspects of Clickhouse?
          • How did you manage the introduction of this new piece of infrastructure to a business that was already handling massive amounts of real-time data?
          • What are some of the downstream benefits of using Clickhouse for managing event data at Sentry?
          • For someone who is interested in using Snuba for their own purposes, how flexible is it for different domain contexts?
          • What are some of the other data challenges that you are currently facing at Sentry?
            • What is your next highest priority for evolving or rebuilding to address technical or business challenges?
            • Contact Info
              • James
                • @JTCunning on Twitter
                • JTCunning on GitHub
                • Ted
                  • tkaemming on GitHub
                  • Website
                  • @tkaemming on Twitter
                  • Parting Question
                    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                    • Closing Announcements
                      • Thank you for listening! Don’t forget to check out our other show, Podcast.__init__ to learn about the Python language, its community, and the innovative ways it is being used.
                      • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                      • If you’ve learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                      • To help other people find the show please leave a review on iTunes and tell your friends and co-workers
                      • Join the community in the new Zulip chat workspace at dataengineeringpodcast.com/chat
                      • Links
                        • Sentry
                          • Podcast.__init__ Episode
                          • Snuba
                            • Blog Post
                            • Clickhouse
                              • Podcast Episode
                              • Disqus
                              • Urban Airship
                              • HBase
                              • Google Bigtable
                              • PostgreSQL
                              • Redis
                              • HyperLogLog
                              • Riak
                              • Celery
                              • RabbitMQ
                              • Apache Spark
                              • Presto
                              • Cassandra
                              • Apache Kudu
                              • Apache Pinot
                              • Apache Druid
                              • Flask
                              • Apache Kafka
                              • Cassandra Tombstone
                              • Sentry Blog
                              • XML
                              • Change Data Capture
                              • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                Support Data Engineering Podcast

                                ...more
                                View all episodesView all episodes
                                Download on the App Store

                                Data Engineering PodcastBy Tobias Macey

                                • 4.6
                                • 4.6
                                • 4.6
                                • 4.6
                                • 4.6

                                4.6

                                135 ratings


                                More shows like Data Engineering Podcast

                                View all
                                Software Engineering Radio - the podcast for professional software developers by se-radio@computer.org

                                Software Engineering Radio - the podcast for professional software developers

                                272 Listeners

                                The Changelog: Software Development, Open Source by Changelog Media

                                The Changelog: Software Development, Open Source

                                283 Listeners

                                The Cloudcast by Massive Studios

                                The Cloudcast

                                152 Listeners

                                Thoughtworks Technology Podcast by Thoughtworks

                                Thoughtworks Technology Podcast

                                41 Listeners

                                Data Skeptic by Kyle Polich

                                Data Skeptic

                                482 Listeners

                                Talk Python To Me by Michael Kennedy

                                Talk Python To Me

                                592 Listeners

                                Software Engineering Daily by Software Engineering Daily

                                Software Engineering Daily

                                625 Listeners

                                The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) by Sam Charrington

                                The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

                                443 Listeners

                                Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

                                Super Data Science: ML & AI Podcast with Jon Krohn

                                296 Listeners

                                Python Bytes by Michael Kennedy and Brian Okken

                                Python Bytes

                                213 Listeners

                                DataFramed by DataCamp

                                DataFramed

                                266 Listeners

                                Practical AI by Practical AI LLC

                                Practical AI

                                189 Listeners

                                The Stack Overflow Podcast by The Stack Overflow Podcast

                                The Stack Overflow Podcast

                                64 Listeners

                                The Real Python Podcast by Real Python

                                The Real Python Podcast

                                140 Listeners

                                Latent Space: The AI Engineer Podcast by swyx + Alessio

                                Latent Space: The AI Engineer Podcast

                                77 Listeners