Data Hurdles

Data Hurdles

By Michael Burke and Chris DetzelBusinessTechnology
Download on the App Store

Data Hurdles episodes

  • EU's AI Act: A Journey from Open Source Tech to High-Stakes Policy

    When Christopher Detzel and Michael Burke sat down for their podcast episode, they had an in-depth conversation about the potential impact of the European Union's (EU) AI Act on open-source artificial intelligence (AI) technologies like large language models (LLMs). The conversation offers crucial insights into the implications of AI regulation, privacy concerns, and the future of the tech industry.

    Starting off on a lighter note, Detzel and Burke exchanged weekend plans, creating an informal atmosphere for their podcast discussion. Soon, the conversation delved into more serious matters—the EU AI Act and its potential ramifications on the open-source AI ecosystem.

    The main point of their conversation was centered on the fact that the EU AI Act targets US open software, including LLMs. The potential disruptive impact of this Act on the global AI landscape, particularly around the open-source movement, was of significant concern. Privacy issues around AI models and the Act's intention to control and safeguard user privacy by regulating the use and deployment of AI was another important topic that came up.

    One of the critical challenges that Burke pointed out is the potential threat to privacy that large language models could pose. According to him, the possibility that LLMs store information input into them and the lack of clarity on the sources of data these models are trained on, are matters of concern. Burke stressed that organizations and governments alike share this worry, particularly in relation to the accuracy and reliability of the information being processed by these models. He further highlighted the severe implications for users sharing sensitive or private information with AI systems unknowingly or without understanding the potential uses of their data.

    29 min
  • AI: The Dawn of a New Era - How Localized Language Models are Shaping the Tech Landscape

    In a recent podcast episode, Michael Burke and Christopher Detzel delve into the rapidly evolving world of large language models (LLMs), discussing their potential impacts on technology and society. The conversation explores the development and application of these models, touching on topics such as localized language models, IoT, democratization of AI, and potential future applications.


    Localized Language Models and IoT


    Localized language models, which can run locally on a device without an internet connection, are gaining traction in the tech world. The ability to provide AI-related services and solutions without significant data or domain expertise presents new opportunities for innovation. Michael Burke shares his experience of using a localized large language model offline during a flight, demonstrating the potential for these models to function independently of internet connectivity.

    Localized LLMs have the potential to revolutionize the Internet of Things (IoT) space by giving IoT devices the ability to understand and interpret the world around them in real-time without needing an internet connection. This capability could enable AI capabilities in areas where it was previously not possible.


    Democratization of AI


    The democratization of AI has made it possible for startups and smaller companies to access the same computational power and data resources that were previously exclusive to tech giants. This democratization fosters innovation, with new companies emerging to solve complex problems using AI.

    As AI models continue to improve, they will be able to hold more questions in their memory, leading to better contextual understanding and more accurate responses. AI models with larger parameters can answer more specific and complex questions, though more computational power is needed to run these models.

    Model Cards and Transformers

    The podcast also discusses the concept of "model cards," which are documents that provide key information about a machine learning model, increasing transparency. They also touch on the emergence of new technologies that provide better traceability and accountability for models.


    Transformers in machine learning are designed to understand and recognize relationships and connections between words and concepts. These models use a self-attention mechanism to understand different ways to ask the same question, improving their ability to understand and respond to queries.

    Future Applications


    Potential future applications of machine learning models include their use in the stock market to understand perception at a global level and make real-time decisions based on this understanding.

    Michael Burke equates the functioning of large language models like OpenAI's GPT-4 to programming languages, which are continuously maintained and updated. Users can fine-tune these AI models for their specific use cases, and they can even translate text between different languages.

    Impact on Jobs and Society

    The impact of AI and machine learning could be greater than previous technological shifts, like the advent of social media platforms or the smartphone revolution. While some areas might experience drastic changes overnight, others might still be decades away from true innovation. Despite the uncertainty, these models have already made a significant impact and opened a new pocket of innovation and potential.


    Localized large language models are shaping the future of AI and technology, with implications for industries and society as a whole. As the democratization of AI continues, the potential for groundbreaking innovations grows. While there are challenges to overcome, the rapid pace of progress in this field suggests that these models could soon become an integral part of our daily lives.

    30 min
  • Entity Resolution Enhanced with LLMs: Insights from Detzel and Burke

    Chris Detzel and Michael Burke discussed the role of large language models (LLMs) in entity resolution, a process that identifies and links records referring to the same real-world entity. LLMs can improve accuracy and efficiency while addressing challenges like data quality and transparency.

    Key Points:
    LLMs enhance entity resolution by understanding context, processing unstructured data, and improving matching processes.

    Ethical considerations, including privacy and bias, are essential when using machine learning in entity resolution.

    Best practices include establishing clear goals, assessing data quality, and choosing suitable algorithms.

    Effectiveness can be measured by having a human in the loop and maintaining feedback between data consumers and entity resolution managers.

    Data quality is vital for success, and machine learning can monitor and ensure accuracy and consistency.

    Real-world applications of machine learning and entity resolution include fraud detection and construction project management.

    24 min
  • Data Quality: The Key to Effective Business Decisions

    The importance of data quality in business decisions and best practices for managing it effectively. It defines data quality as accurate, reliable, and relevant information for intended use cases. The importance of governance and ownership in data management is also explained through a waterworks system analogy. 

    The need for cleansing, standardization, and enrichment to improve data quality. It also covers best practices for managing data quality, such as identifying relevant metrics, designing monitoring strategies, and tailoring metrics to stakeholders' needs. The role of emerging technologies, such as machine learning, in improving data quality and ethical considerations around data quality are also discussed. Focusing on data accuracy, consistency, completeness, and integrity is crucial for informed decision-making and business growth.

    32 min
  • DataArmor Analysis: Dissecting Cybersecurity Breaches and Best Practices


    In the recent Data Hurdles podcast episode, hosts Michael Burke and Chris Detzel interview Kristof Holm from DataBlend, discussing the 3CX data breach orchestrated by North Korean hackers. The blog explores the key aspects of the breach, the response by the company, and the importance of proper security practices and communication in protecting businesses and individuals from cyber threats.

    Key Sections:
    The Breach and Its Impact: A detailed account of the 3CX breach, the Lazarus group's involvement, and the potential risks posed by such attacks.

    3CX's Response: A critical analysis of the company's initial response, emphasizing the need for robust internal security processes and communication plans.

    Protecting Businesses and Individuals: A comprehensive discussion of measures to safeguard customers and businesses, including due diligence, open communication, basic security hygiene, and additional support services.

    Limiting the Value of Attacks: A strategic approach to discouraging cyber attacks by making it more challenging for hackers to access sensitive data and implementing strong security measures.

    Conclusion: A summary emphasizing the significance of effective security practices and communication in addressing the ever-evolving landscape of cyber risks, urging businesses and individuals to take necessary precautions for enhanced protection.

    31 min
  • Data Literacy for Better Decision-Making

    In a conversation between Chris Detzel and Michael Burke, the importance of data literacy in making informed decisions across various aspects of life was emphasized. Data literacy helps individuals gain a competitive advantage by understanding and interpreting complex information. Applications of data-driven decision-making include diet, exercise, personal finance, and more. 

    By learning tools like Excel or Google Sheets, individuals can become more data literate and make better choices in their lives. Embracing transparency, accountability, and data-driven decision-making can lead to improved financial, physical, and mental well-being.

    24 min
  • Data Pipelines: Transforming Raw Data into Actionable Insights

    This podcast episode talks about data pipelines, which are used to move data from one place to another and transform it into a more usable form. The podcast compares data pipelines to water pipelines, where raw data is like dirty water that needs to be cleaned and enriched. 

    The podcast covers topics such as batch and real-time pipelines, serverless computing, ETL, and the challenges of building and managing data pipelines. Michael and Chris also discuss the importance of data quality, involving the right people, and understanding business objectives. They emphasize the need to view data pipelines as a product that requires ongoing maintenance and support, and they provide tips for managing data pipelines in organizations.

    Check out Data Pipelines explained here in this video: https://youtu.be/6kEGUCrBEU0

    41 min
  • Reinforcement Learning in Machine Learning: Real-World Applications

    This Data Hurdles podcast episode discusses reinforcement learning in machine learning. The hosts define reinforcement learning as the process of decision making where the model learns an optimal behavior in an environment obtained by a reward. They use the analogy of a child learning how to engage with fire to explain this concept. The hosts also highlight some real-life examples of reinforcement learning being used in various fields, including gaming, robotics, marketing, healthcare, and finance. 

    They note that while reinforcement learning can be challenging to implement and sensitive to the choice of reward function, with careful design and tuning, it can lead to powerful and adaptable AI systems. The conversation also covers the Mario case as an interesting example of reinforcement learning in a controlled environment.

    22 min
  • Data Security Challenges: Insights from a CISO in the Integration Platform Industry

    In this episode of the Data Hurdles podcast, Chris Detzel and Michael Burke interviewed Kristof Holm, CISO of a small integration platform as a service company called DataBlend. The discussion focused on the role of a Chief Information Security Officer (CISO) and the challenges that CEOs face in managing data and machine learning.

    Kristof emphasized the importance of balancing the trade-offs between security and accessibility, while keeping up with evolving regulations and compliance standards. Michael and Kristof discussed the challenge of sharing information about a company's system with security professionals without compromising intellectual property, and the importance of establishing nondisclosure agreements.

    The conversation also covered traditional approaches to security, including the castle walls and layers of an onion analogy, as well as more modern approaches such as the perimeter-free zone and Zero Trust. Kristof noted that their environment heavily relies on AWS, which allows for easy adoption of new technologies.

    Overall, the episode provides valuable insights into the role of a CISO and the challenges and opportunities of managing data in today's digital landscape.

    23 min
  • Data-Driven Customer Experience: A Guide to Understanding and Utilizing Customer Data

    We explore the concept of the data-driven consumer experience and how companies are using customer data to improve their products and services. We discuss the potential benefits and drawbacks of collecting customer data and how businesses can act responsibly with this information. We also explore how companies can measure the success of a data-driven customer experience initiative and the tools available to consolidate and analyze customer data. Ultimately, we conclude that while data can be messy, a complete understanding of customers and their actions can provide valuable insights for businesses.

    31 min

About Data Hurdles

From the publisher's feed

Data Hurdles is a podcast that brings the stories of data professionals to life, showcasing the challenges, triumphs, and insights from those shaping the future of data. Hosted by Michael Burke and Chris Detzel, this podcast dives into the real-world experiences of data experts as they navigate topics like data quality, security, AI, data literacy, and machine learning.