The Data Standard

The Data Standard

By The Data StandardBusinessNewsTechnologyTech News
Download on the App Store

The Data Standard episodes

  • Data Standard Audio Experience Tim Eller talks about his new company
    Data Standard Audio Experience Host Darren Kaplan sits down with Tim Eller to talk about his new company dcyd.  dcyd is the most intuitive and informative ML model performance monitoring SaaS. Think NewRelic for AI. dcyd watches the watchmen.https://www.dcyd.io/For more information about The Data Standard please visit www.datastandard.ioThe Data Standard LinkedIn: https://www.linkedin.com/company/the-data-standard/
    17 min
  • DataStandard.io Audio Experience with Marc Rind
    In this episode, Darren Kaplan sits down Marc Rind, VP Software Development and Data Analytics at Fiserv to discuss synchronous collaboration tools.For more information about the Data Standard visit our website at www.datastandard.ioThe Data Standard LinkedIn: https://www.linkedin.com/company/the-data-standard/
    14 min
  • Data Standard Audio Experience Alison McCauley Talks Blockchain
    Alison McCauley best-selling author of Unblocked sits down with The Data Standard to talk about the business and cultural impacts of Blockchain.Host Darren KaplanFor more information about The Data Standard visit www.datastandard.ioThe Data Standard LinkedIn: https://www.linkedin.com/company/the-data-standard/For more information Alison McCauley and her best-selling book Unblocked visit: https://www.alisonmccauley.io/
    11 min
  • Tips From a Kaggle Expert to Amplify Your Brand and Data Science Work
    Show Notes:In this episode host, Darren Kaplan talks with Kaggle expert and soon to become Ph.D. Max Halford about how to amplify your data science work and research by sharing it on Medium, Github, and Kaggle competitions. Check out Max's opensource library for online machine learning https://github.com/creme-ml/cremeAnd for more information about The Data Standard visit us at www.datastandard.ioThe Data Standard LinkedIn: https://www.linkedin.com/company/the-data-standard/
    15 min
  • Zoom, Hangouts and Teams End to End Encryption and Why It Is a Big Deal
    In this episode host Darren Kaplan sits down with Cyber Security Data Scientist and Author Charles Givre to talk about End-to-End encryption and the new API tool he is building.As many of around the world work from Video collaboration tools like Zoom, Google Hangouts, and Microsoft Teams are the new normal. What is the business impact of end-to-end encryption and why does it matter? Rest APIs are also hot right now. Charles shares a new platform he built for the enterprise. For more information about his tool check out his video. https://www.youtube.com/watch?v=oEOhFWm3D9A&t=739sHost Darren Kaplan https://www.linkedin.com/in/darrenkaplan500/Guest: Charles Givre https://www.linkedin.com/in/cgivre/For more information about The Data Standard visit http://www.datastandard.ioThe Data Standard LinkedIn: https://www.linkedin.com/company/the-data-standard/
    17 min
  • Tips to Build a Business Case to Get Your Data Science Initiative Funded
    Host Darren Kaplan sits down with The Good Factor CEO and SEO Ali Bouhouch to talk about how to get your data science initiative funded. The Good Data Factory is a team of Data and Algorithm professionals bent on changing the way AI, Data Science and Machine Learning are used to transform business and society at large for a better, more sustainable and more human experience. www.thegoodfactory.comFor more information about The Data Standard go to https://www.datastandard.ioThe Data Standard LinkedIn: https://www.linkedin.com/company/the-data-standard/
    30 min
  • Bitcoin, Blockchain and Ready Player One
    In this podcast, we cover:- A quick intro to blockchain technology and cryptocurrency- How this relates to virtual worlds and digital economies- Why Ready Player One defined an economy that would need blockchain technology- Trends shaping the world today- Rev (revstarts.com), an online accelerator helping founders tackle meaningful opportunities from day 1.The Data Standard LinkedIn: https://www.linkedin.com/company/the-data-standard/
    19 min
  • Josh Odmark, Pandio.com CTO and Co-Founder talks about Presto by Facebook

    “You can think of Presto almost like a database, but it’s more of an abstraction layer for a large group of databases” – Josh Odmark, CTO & Founder of Pandio. 

    In this episode of The Data Standard, Darren Kaplan is joined by the CTO & Founder of Pandio, Josh Odmark. Josh is a full-stack engineer expert that worked with PHP, Python, Ruby, JavaScript, and SQL. 

    As an engineer working in many different technologies, it’s interesting to hear what Josh has to say about Presto and how he uses it for public data around museum information. Josh joins us to talk about how he uses Presto by Facebook and gives us a quick demo of his approach and the benefits that he thinks are most valuable. 

    Presto is a distributed SQL query engine. It’s an open-source Facebook technology and can be adjusted to anyone’s needs. It gives users a completely different approach to querying data. Traditionally, data is copied or moved into some warehouse, but Presto doesn’t do that; it lets you query data in place.  

    But Presto requires access to a flat file or database. To show this as an example, Odmark downloaded data from the Museum of Modern Art from their open-access database. There’s no need for any kind of preparation, simply download the file and open it. 

    Even though this file requires a bit of ETL, it’s not an issue since Presto lets users run SQL against these kinds of data sets and open up many options. The second data set that he used was from The Metropolitan Museum of Art. Both of these files are typical CSVs. 

    Using the AWS s3 Presto one-click install, you can instantly run Presto within AWS. This allows users to run SQL commands against different datasets that have been added to the AWS. Presto’s simplicity is a thing of beauty because users only have to set up the schema and point it to the desired files.

    With the “Show Table” command, the user can easily see all of the datasets added visually displayed as tables. These tables have the traditional data frame, including columns and data types. Josh used Presto to set up the table view of the datasets from these two museums in New York. 

    Even though both museums have different data structures, the platform takes raw files and puts them up in a similar manner and runs SQL queries against them. Odmark showed this on an example where he ran an SQL query against both of the museum datasets to join the two tables based on artist names. 

    Even though both museums have a different way of storing and structuring their data, Presto can run the query effectively. Still, certain conditionals need to be added, but this is what Presto is about. It gives full SQL capabilities, making it easy to do some ETL actions that will enable you to do your queries.

    Presto takes a couple of seconds to run a query against two datasets around 300MB in size. It goes through all the rows and columns to find all of the artist names and display several types of results: 

        Which artists are displayed or working in both museums; 

        Which artwork is from a single artist is in one museum, and which pieces of the same artist are in the other museum;

    Presto has a query planner that adjusts the query that has been added. This way, it allows the same query to run effectively against different clusters of data. Despite these complex processes going on in the background, the execution is really fast. 

    Even though he used two files to simulate databases, Presto lets users add multiple databases and files and works in the same manner. Presto can connect many different types of data but also connect things within it. 

    It can query data no matter where it is stored, including services like Cassandra or Hive. It works both with proprietary data stores and relational databases. Simply put, it has the unique ability to combine different data from various sources using queries. 

    This opens up many opportunities for organizations to do essential analytics quickly that can give valuable answers. At the same time, it’s designed for those analysts, developers, or engineers that need quick solutions. 

    The results are displayed from a couple of seconds to a couple of minutes, depending on how large the datasets are and how many are there. It’s the best of both worlds when it comes to analytics platforms. 

    Presto is both quick and free. Even though it already offers some significant benefits, we can expect to see even greater things in the future. As more people use this open-source platform, we will likely see new upgrades and changes that perfect its functionalities and expand its potential.

    Make sure to check out the full podcast with Josh Odmark and Darren Kaplan at The Data Standard website. 

    12 min

About The Data Standard

From the publisher's feed

Listen to some of the top data professionals in the world speak on the latest innovations in the data field. Learn how machine learning and artificial intelligence are impacting the tech industries.