Data Engineering Podcast

Safely Test Your Applications And Analytics With Production Quality Data Using Tonic AI


Listen Later

Summary

The most interesting and challenging bugs always happen in production, but recreating them is a constant challenge due to differences in the data that you are working with. Building your own scripts to replicate data from production is time consuming and error-prone. Tonic is a platform designed to solve the problem of having reliable, production-like data available for developing and testing your software, analytics, and machine learning projects. In this episode Adam Kamor explores the factors that make this such a complex problem to solve, the approach that he and his team have taken to turn it into a reliable product, and how you can start using it to replace your own collection of scripts.

Announcements
  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
  • Truly leveraging and benefiting from streaming data is hard - the data stack is costly, difficult to use and still has limitations. Materialize breaks down those barriers with a true cloud-native streaming database - not simply a database that connects to streaming systems. With a PostgreSQL-compatible interface, you can now work with real-time data using ANSI SQL including the ability to perform multi-way complex joins, which support stream-to-stream, stream-to-table, table-to-table, and more, all in standard SQL. Go to dataengineeringpodcast.com/materialize today and sign up for early access to get started. If you like what you see and want to help make it better, they're hiring across all functions!
  • Data and analytics leaders, 2023 is your year to sharpen your leadership skills, refine your strategies and lead with purpose. Join your peers at Gartner Data & Analytics Summit, March 20 – 22 in Orlando, FL for 3 days of expert guidance, peer networking and collaboration. Listeners can save $375 off standard rates with code GARTNERDA. Go to dataengineeringpodcast.com/gartnerda today to find out more.
  • Your host is Tobias Macey and today I'm interviewing Adam Kamor about Tonic, a service for generating data sets that are safe for development, analytics, and machine learning
  • Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what Tonic is and the story behind it?
    • What are the core problems that you are trying to solve?
    • What are some of the ways that fake or obfuscated data is used in development and analytics workflows?
    • challenges of reliably subsetting data
      • impact of ORMs and bad habits developers get into with database modeling
      • Can you describe how Tonic is implemented?
        • What are the units of composition that you are building to allow for evolution and expansion of your product?
        • How have the design and goals of the platform evolved since you started working on it?
        • Can you describe some of the different workflows that customers build on top of your various tools
        • What are the most interesting, innovative, or unexpected ways that you have seen Tonic used?
        • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Tonic?
        • When is Tonic the wrong choice?
        • What do you have planned for the future of Tonic?
        • Contact Info
          • LinkedIn
          • @AdamKamor on Twitter
          • Parting Question
            • From your perspective, what is the biggest gap in the tooling or technology for data management today?
            • Closing Announcements
              • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
              • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
              • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
              • To help other people find the show please leave a review on Apple Podcasts and tell your friends and co-workers
              • Links
                • Tonic
                  • Djinn
                  • Django
                  • Ruby on Rails
                  • C#
                  • Entity Framework
                  • PostgreSQL
                  • MySQL
                  • Oracle DB
                  • MongoDB
                  • Parquet
                  • Databricks
                  • Mockaroo
                  • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                    Sponsored By:

                    • Materialize: ![Materialize](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/NuMEahiy.png)
                    Looking for the simplest way to get the freshest data possible to your teams? Because let's face it: if real-time were easy, everyone would be using it. Look no further than Materialize, the streaming database you already know how to use.
                    Materialize’s PostgreSQL-compatible interface lets users leverage the tools they already use, with unsurpassed simplicity enabled by full ANSI SQL support. Delivered as a single platform with the separation of storage and compute, strict-serializability, active replication, horizontal scalability and workload isolation — Materialize is now the fastest way to build products with streaming data, drastically reducing the time, expertise, cost and maintenance traditionally associated with implementation of real-time features.
                    Sign up now for early access to Materialize and get started with the power of streaming data with the same simplicity and low implementation cost as batch cloud data warehouses.
                    Go to [materialize.com](https://materialize.com/register/?utm_source=depodcast&utm_medium=paid&utm_campaign=early-access)
                  • Gartner: ![Gartner](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/4ODnKDqa.jpg)
                  • The evolving business landscape continues to create challenges and opportunities for data and analytics (D&A) leaders — shifting away from focusing solely on tools and technology to decision making as a business competency. D&A teams are now in a better position than ever to help lead this change within the organization.
                    Harnessing the full power of D&A today requires D&A leaders to guide their teams with purpose and scale their scope beyond organizational silos as companies push to transform and accelerate their data-driven strategies. Gartner Data & Analytics Summit 2023 addresses the most significant challenges D&A leaders face while navigating disruption and building the adaptable, innovative organizations this shifting environment demands.
                    Go to [dataengineeringpodcast.com/gartnerda](https://www.dataengineeringpodcast.com/gartnerda) Listeners can save $375 off standard rates with code GARTNERDA Promo Code: GartnerDA

                    Support Data Engineering Podcast

                    ...more
                    View all episodesView all episodes
                    Download on the App Store

                    Data Engineering PodcastBy Tobias Macey

                    • 4.6
                    • 4.6
                    • 4.6
                    • 4.6
                    • 4.6

                    4.6

                    134 ratings


                    More shows like Data Engineering Podcast

                    View all
                    Software Engineering Radio - the podcast for professional software developers by se-radio@computer.org

                    Software Engineering Radio - the podcast for professional software developers

                    262 Listeners

                    The Changelog: Software Development, Open Source by Changelog Media

                    The Changelog: Software Development, Open Source

                    285 Listeners

                    The Cloudcast by Massive Studios

                    The Cloudcast

                    153 Listeners

                    Thoughtworks Technology Podcast by Thoughtworks

                    Thoughtworks Technology Podcast

                    43 Listeners

                    Data Skeptic by Kyle Polich

                    Data Skeptic

                    474 Listeners

                    Talk Python To Me by Michael Kennedy

                    Talk Python To Me

                    585 Listeners

                    Software Engineering Daily by Software Engineering Daily

                    Software Engineering Daily

                    630 Listeners

                    AWS Podcast by Amazon Web Services

                    AWS Podcast

                    200 Listeners

                    Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

                    Super Data Science: ML & AI Podcast with Jon Krohn

                    295 Listeners

                    Python Bytes by Michael Kennedy and Brian Okken

                    Python Bytes

                    212 Listeners

                    DataFramed by DataCamp

                    DataFramed

                    267 Listeners

                    Practical AI by Practical AI LLC

                    Practical AI

                    196 Listeners

                    The Stack Overflow Podcast by The Stack Overflow Podcast

                    The Stack Overflow Podcast

                    63 Listeners

                    The Real Python Podcast by Real Python

                    The Real Python Podcast

                    136 Listeners

                    Latent Space: The AI Engineer Podcast by swyx + Alessio

                    Latent Space: The AI Engineer Podcast

                    64 Listeners