Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

Download on the App Store

Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov episodes

  • Using Event-Driven Design with Apache Kafka Streaming Applications ft. Bobby Calderwood

    What is event modeling and how does it differ from standard data modeling?

    In this episode of Streaming Audio, Bobby Calderwood, founder of Evident Systems and creator of oNote observes that at the dawn of the computer age, due to the fact that memory and computing power were expensive, people began to move away from time-and-narrative-oriented record-keeping systems (in the manner of a ship's log or a financial ledger) to systems based on aggregation. Such data-model systems, still dominant today, only retain the current state generated from their inputs, with the inputs themselves going lost. A converse approach to the reductive data-model system is the event-model system, which is enabled by tools like Apache Kafka®, and which effectively saves every bit of activity that the system generates. The event model actually marks a return, in a sense, to the earlier, narrative-like recording methods.

    To further illustrate, Bobby uses a chess example to show the distinction between the data model and the event model. In a chess context, the event modeling system would retain each move in the game from beginning to end, such that any moment in the game could be derived by replaying the sequence of moves. Conversely, chess based on the data model would save only the current state of the game, destructively mutating the data structure to reflect it. 

    The event model maintains an immutable log of all of a system's activity, which means that teams downstream from the transactions team have access to all of the system's data, not just the end transactions, and they can analyze the data as they wish in order to make their own conclusions. Thus there can be several read models over the same body of events. Bobby has found that non-programming stakeholding teams tend to intuitively comprehend the event model better than other data paradigms, given its natural narrative form.    

    Transitioning from the data model to the event model, however, can be challenging. Bobby’s oNote—event modeling platform aims to help by providing a digital canvas that allows a system to be visually redesigned according to the event model. oNote generates Avro schema based on its models, and also uses Avro to generate runtime code.

    EPISODE LINKS

    • Event Sourcing and Event Storage with Apache Kafka
    • oNote
    • Event Modeling
    • Toward a Functional Programming Analogy for Microservices
    • Event-Driven Architecture - Common Mistakes and Valuable Lessons ft. Simon Aubury
    • Watch the video version of this podcast
    • Coding in Motion Workshop: Build a Streaming App
    • Kris Jenkins’ Twitter
    • Join the Confluent Community
    • Learn more with Kafka tutorials, resources, and guides at Confluent Developer
    • Live demo: Intro to Event-Driven Microservices with Confluent
    • Use PODCAST100 to get an additional $100 of free Confluent Cloud usage (details)

    SEASON 2
    Hosted by Tim Berglund, Adi Polak and Viktor Gamov
    Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
    Music by Coastal Kites 
    Artwork by Phil Vo 

    •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
    • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
    • 👍 If you enjoyed this, please leave us a rating. 
    • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
    52 min
  • Monitoring Extreme-Scale Apache Kafka Using eBPF at New Relic

    New Relic runs one of the larger Apache Kafka® installations in the world, ingesting circa 125 petabytes a month, or approximately three billion data points per minute. Anton Rodriguez is the architect of the system, responsible for hundreds of clusters and thousands of clients, some of them implemented in non-standard technologies. In addition to the large volume of servers, he works with many teams, which must all work together when issues arise.

    Monitoring New Relic's large Kafka installation is critical and of course challenging, even for a company that itself specializes in monitoring. Specific obstacles include determining when rebalances are happening, identifying particularly old consumers, measuring consumer lag, and finding a way to observe all producing and consuming applications.

    One way that New Relic has improved the monitoring of its architecture is by directly consuming metrics from the Linux kernel using its new eBPF technology, which lets programs run inside the kernel without changing source code or adding additional modules (the open-source tool Pixie enables access to eBPF in a Kafka context). eBPF is very low impact, so doesn’t affect services, and it allows New Relic to see what’s happening at the network level—and to take action as necessary.

    EPISODE LINKS

    • Monitoring Kafka Without Instrumentation Using eBPF
    • What Is eBPF and Why Does It Matter for Observability?
    • Kafka Monitoring
    • Kafka Summit: Monitoring Kafka Without Instrumentation Using eBPF
    • Watch the video version of this podcast
    • Kris Jenkins’ Twitter
    • Streaming Audio Playlist 
    • Join the Confluent Community
    • Learn more with Kafka tutorials, resources, and guides at Confluent Developer
    • Live demo: Intro to Event-Driven Microservices with Confluent
    • Use PODCAST100 to get an additional $100 of free Confluent Cloud usage (details)   

    SEASON 2
    Hosted by Tim Berglund, Adi Polak and Viktor Gamov
    Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
    Music by Coastal Kites 
    Artwork by Phil Vo 

    •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
    • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
    • 👍 If you enjoyed this, please leave us a rating. 
    • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
    39 min
  • Confluent Platform 7.1: New Features + Updates

    Confluent Platform 7.1 expands upon its already innovative features, adding improvements in key areas that benefit data consistency, allow for increased speed and scale, and enhance resilience and reliability.

    Previously, the Confluent Platform 7.0 release introduced Cluster Linking, which enables you to bridge on-premises and cloud clusters, among other configurations. Maintaining data quality standards across multiple environments can be challenging though. To assist with this problem, CP 7.1 adds Schema Linking, which lets you share consistent schemas across your clusters—synced in real time.

    Confluent for Kubernetes lets you build your own private-cloud Apache Kafka® service. Now you can enhance the global resilience of your architecture by employing to multiple regions. With the new release you can also configure custom volumes attached to Confluent deployments and you can declaratively define and manage the new Schema Links. As of this release, Confluent for Kubernetes now supports the full feature set of the Confluent Platform. 

    Tiered Storage was released in Confluent Platform 6.0, and it offers immense benefits for a cluster by allowing the offloading of older topic data out of the broker and into slower, long-term object storage. The reduced amount of local data makes maintenance, scaling out, recovery from failure, and adding brokers all much quicker. CP 7.1 adds compatibility for object storage using Nutanix, NetApp, MinIO, and Dell, integrations that have been put through rigorous performance and quality testing.

    Health+ was introduced in CP 6.2—offers intelligent cloud-based alerting and monitoring tools in a dashboard. New as of CP 7.1, you can choose to be alerted when anomalies in broker latency are detected, when there is an issue with your connectors linking Kafka and external systems, as well as when a ksqlDB query will interfere with a continuous, real-time processing stream. 

    Shipping with CP 7.1 is ksqlDB 0.23, which adds support for pull queries against streams as opposed to only against tables—a milestone development that greatly helps when debugging since a subset of messages within a topic can now be inspected. ksqlDB 0.23 also supports custom schema selection, which lets you choose a specific schema ID when you create a new stream or table, rather than use the latest registered schema. A number of additional smaller enhancements are also included in the release.

    EPISODE LINKS

    • Download Confluent Platform 7.1
    • Check out the release notes
    • Read the Confluent Platform 7.1 blog post
    • Watch the video version of this podcast
    • Join the Confluent Community
    • Learn more with Kafka tutorials, resources, and guides at Confluent Developer
    • Live demo: Intro to Event-Driven Microservices with Confluent
    • Use PODCAST100 to get $100 of free Confluent Cloud usage (details)

    SEASON 2
    Hosted by Tim Berglund, Adi Polak and Viktor Gamov
    Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
    Music by Coastal Kites 
    Artwork by Phil Vo 

    •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
    • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
    • 👍 If you enjoyed this, please leave us a rating. 
    • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
    11 min
  • Scaling an Apache Kafka Based Architecture at Therapie Clinic

    Scaling Apache Kafka® can be tricky, let alone scaling a team. When he was first hired, Domenico Fioravanti of Therapie Clinic was given the challenging task of assembling a sizable tech team from scratch, while simultaneously building a scalable and decoupled architecture from the ground up. In addition, he wanted to deliver value to the company from day one. One way that Domenico ultimately accomplished these goals was by focusing on managed solutions in order to avoid large investments in engineering know-how. Another way was to deliver quickly to production by using the existing knowledge of his team.

    Domenico's biggest initial priority was to make a real-time reporting dashboard that collated data generated by third-party systems, such as call centers and front-of-house software solutions that managed bookings and transactions. (Before Domenico's arrival, all reporting had been done by aggregating data from different sources through an expensive, manual, error-prone, and slow process—which tended to result in late and incomplete insights.)

    Establishing an initial stack with AWS and a BI/analytics tool only took a month and required minimal DevOps resources, but Domenico's team ended up wanting to leverage their efforts to free up third-party data for more than just the reporting/data insights use case.

    So they began considering Apache Kafka® as a central repository for their data. For Kafka itself, they investigated Amazon MSK vs. Confluent, carefully weighing setup and time costs, maintenance costs, limitations, security, availability, risks, migration costs, Kafka updates frequency, observability, and errors and troubleshooting needs.

    Domenico's team settled on Confluent Cloud and built the following stack:

    • AWS AppSync, a managed GraphQL layer to interact with and abstract third-party APIs (data sources)
    • AWS Lambdas for extracting data and producing to Kafka topics
    • Kafka topics for the raw as well as transformed data
    • Kafka Streams for data transformation
    • Kafka Redshift sink connector for loading data
    • ​​AWS Redshift as the destination cloud data warehouse 
    • Looker for business intelligence and big data analytics 

    This stack allowed the company's data to be consumed by multiple teams in a scalable way. Eventually, DynamoDB was added and by the end of a year, along with a scalable architecture, Domenico had successfully grown his staff to 45 members on six teams.

    EPISODE LINKS

    • Confluent’s Data Streaming Platform Can Save Over $2.5M vs. Self-Managing Apache Kafka
    • Accelerate Your Cloud Data Warehouse Migration and Modernization with Confluent
    • Watch the video version of this podcast
    • Kris Jenkins' Twitter
    • Streaming Audio Playlist 
    • Join the Confluent Community
    • Learn more with Kafka tutorials, resources, and guides at Confluent Developer
    • Live demo: Intro to Event-Driven Microservices with Confluent
    • Use PODCAST100 to get an additional $100 of free Confluent Cloud usage (details)  

    SEASON 2
    Hosted by Tim Berglund, Adi Polak and Viktor Gamov
    Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
    Music by Coastal Kites 
    Artwork by Phil Vo 

    •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
    • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
    • 👍 If you enjoyed this, please leave us a rating. 
    • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
    1 hr 11 min
  • Bridging Frontend and Backend with GraphQL and Apache Kafka ft. Gerard Klijs

    What is GraphQL? And how can you combine GraphQL with Apache Kafka® to query data in real time?

    With over 10 years of experience as a backend engineer, Gerard Klijs is a Confluent Community Catalyst, a contributor to several GraphQL libraries, and also a creator and maintainer of a Rust library to use Confluent Schema Registry with Java client. In this episode, he explains why you want to use Kafka with GraphQL and how they work together to bridge the gap between backend and frontend to make data more easily accessible in the frontend.  

    As an alternative to REST, GraphQL is an open source programming language developed by Meta, which lets you pull data from multiple data sources via a single API call. GraphQL lets you migrate and deprecate data easily. For example, if you have a `name` field, which you later decided to replace by `firstName` and `lastName`, you can group the field names together and monitor the server for query requests. If there are no additional query requests for the deprecated field, then it can be removed from the server.

    Usually, GraphQL is used in the frontend with a server implemented in Node.js, while Kafka is often used as an integration layer between backend components. When it comes to connecting Kafka with GraphQL, the use cases might not seem as vast at first glance, but Gerard thinks that it is due to unfamiliarity and misconceptions on how the two can work together. For example, some may think Kafka is merely a message bus and GraphQL is for graph databases.

    Gerard also talks about the backend for frontend (BFF) pattern as well as tips on working with GraphQL. 

    EPISODE LINKS

    • Getting Started with GraphQL and Apache Kafka
    • Kafka and GraphQL: Misconceptions and Connections
    • Gerard Klijs Github
    • Watch the video version of this podcast
    • Kris Jenkins Twitter
    • Streaming Audio Playlist 
    • Join the Confluent Community
    • Learn more with Kafka tutorials, resources, and guides at Confluent Developer
    • Live demo: Intro to Event-Driven Microservices with Confluent
    • Use PODCAST100 to get an additional $100 of free Confluent Cloud usage (details)  

    SEASON 2
    Hosted by Tim Berglund, Adi Polak and Viktor Gamov
    Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
    Music by Coastal Kites 
    Artwork by Phil Vo 

    •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
    • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
    • 👍 If you enjoyed this, please leave us a rating. 
    • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
    24 min
  • Building Real-Time Data Governance at Scale with Apache Kafka ft. Tushar Thole

    Data availability, usability, integrity, and security are words that we sometimes hear a lot. But what do they actually look like when put into practice? That’s where data governance comes in. This becomes especially tricky when working with real-time data architectures.

    Tushar Thole (Senior Manager, Engineering, Trust & Security, Confluent) focuses on delivering features for software-defined storage, software-defined networking (SD-WAN), security, and cloud-native domains. In this episode, he shares the importance of real-time data governance and the product portfolio—Stream Governance, which his team has been building to fostering the collaboration and knowledge sharing necessary to become an event-centric business while remaining compliant within an ever-evolving landscape of data regulations. 

    With the increase of data volume, variety, and velocity, data governance is mandatory for trustworthy, usable, accurate, and accessible data across organizations, especially with distributed data in motion. 

    When it comes to choosing a tool to govern real-time distributed data, there is often a paradox of choice. Some tools are built for handling data at rest, while open source alternatives lack features and are not managed services that can be integrated with the Apache Kafka® ecosystem natively. 

    To solve governance use cases by delivering high-quality data assets, Tushar and his team have been taking Confluent Schema Registry, considered the de facto metadata management standard for the ecosystem, to the next level. This approach to governance allows organizations to scale Kafka operations for real-time observability with security and quality. 

    The fully managed, cloud-native Stream Governance framework is based on three key workflows: 

    • Stream catalog: Search and discover data in a self-service fashion
    • Stream lineage: Understand the complex data relationships with interactive, end-to-end maps of event streams 
    • Stream quality: Deliver trusted, high-quality event streams to the organization 

    Tushar also shares use cases around data governance and sheds light on the Stream Governance roadmap. 

    EPISODE LINKS

    • Stream Governance – How it Works
    • Data Mess to Data Mesh | Jay Kreps
    • Demo: Stream Governance
    • Data Governance for Real Time Data
    • Watch the video version of this podcast
    • Kris Jenkins Twitter
    • Streaming Audio Playlist 
    • Join the Confluent Community
    • Learn more with Kafka tutorials, resources, and guides at Confluent Developer
    • Live demo: Intro to Event-Driven Microservices with Confluent
    • Use PODCAST100 to get an additional $100 of free Confluent Cloud usage (details)  

    SEASON 2
    Hosted by Tim Berglund, Adi Polak and Viktor Gamov
    Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
    Music by Coastal Kites 
    Artwork by Phil Vo 

    •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
    • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
    • 👍 If you enjoyed this, please leave us a rating. 
    • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
    43 min
  • Handling 2 Million Apache Kafka Messages Per Second at Honeycomb

    How many messages can Apache Kafka® process per second? At Honeycomb, it's easily over one million messages. 
     
    In this episode,  get a taste of how Honeycomb uses Kafka on massive scale. Liz Fong-Jones (Principal Developer Advocate, Honeycomb) explains how Honeycomb manages Kafka-based telemetry ingestion pipelines and scales Kafka clusters. 

    And what is Honeycomb? Honeycomb is an observability platform that helps you visualize, analyze, and improve cloud application quality and performance. Their data volume has grown by a factor of 10 throughout the pandemic, while the total cost of ownership has only gone up by 20%. 

    But how, you ask? As a developer advocate for site reliability engineering (SRE) and observability, Liz works alongside the platform engineering team on optimizing infrastructure for reliability and cost. Two years ago, the team was facing the prospect of growing from 20 Kafka brokers to 200 Kafka brokers as data volume increased. The challenge was to scale and shuffle data between the number of brokers while maintaining cost efficiency.

    The Honeycomb engineering team has experimented with using sc1 or st1 EBS hard disks to store the majority of longer-term archives and keep only the latest hours of data on NVMe instance storage. However, this approach to cost reduction was not ideal, which resulted in needing to keep data that is older than 24 hours on SSD. The team began to explore and adopt Zstandard compression to decrease bandwidth and disk size; however, the clusters were still struggling to keep up. 

    When Confluent Platform 6.0 rolled out Tiered Storage, the team saw it as a feature to help them break away from being storage bound. Before bringing the feature into production, the team did a proof of concept, which helped them gain confidence as they watched Kafka tolerate broker death and reduce latencies in fetching historical data. Tiered Storage now shrinks their clusters significantly so that they can hold on to local NVMe SSD and the tiered data is only stored once on Amazon S3, rather than consuming SSD on all replicas. In combination with the AWS Im4gn instance, Tiered Storage allows the team to scale for long-term growth. 

    Honeycomb also saved 87% on the cost per megabyte of Kafka throughput by optimizing their Kafka clusters.

    EPISODE LINKS

    • Tiered Storage
    • Introducing Confluent Platform 6.0
    • Scaling Kafka at Honeycomb
    • Watch the video version of this podcast
    • Kris Jenkins Twitter
    • Streaming Audio Playlist 
    • Join the Confluent Community
    • Learn more with Kafka tutorials, resources, and guides at Confluent Developer
    • Live demo: Intro to Event-Driven Microservices with Confluent
    • Use PODCAST100 to get an additional $100 of free Confluent Cloud usage (details)  

    SEASON 2
    Hosted by Tim Berglund, Adi Polak and Viktor Gamov
    Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
    Music by Coastal Kites 
    Artwork by Phil Vo 

    •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
    • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
    • 👍 If you enjoyed this, please leave us a rating. 
    • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
    42 min
  • Why Data Mesh? ft. Ben Stopford

    With experience in data infrastructure and distributed data technologies, author of the book “Designing Event-Driven Systems” Ben Stopford (Lead Technologist, Office of the CTO, Confluent) explains the data mesh paradigm, differences between traditional data warehouses and microservices, as well as how you can get started with data mesh.   

    Unlike standard data architecture, data mesh is about moving data away from a monolithic data warehouse into distributed data systems. Doing so will allow data to be available as a product—this is also one of the four principles of data mesh: 

    1. Data ownership by domain
    2. Data as a product
    3. Data available everywhere for self-service
    4. Data governed wherever it is

    These four principles are technology agnostic, which means that they don’t restrict you to a programming language, Apache Kafka®, or other databases. Data mesh is all about building point-to-point architecture that lets you evolve and accommodate real-time data needs with governance tools.

    Fundamentally, data mesh is more than a technological shift. It’s a mindset shift that requires cultural adaptation of product thinking—treating data as a product instead of data as an asset or resource. Data mesh invests ownership of data by the people who create it with requirements that ensure quality and governance. Because data mesh consists of a map of interconnections, it’s important to have governance tools in place to identify data sources and provide data discovery capabilities. 

    There are many ways to implement data mesh, event streaming being one of them. You can ingest data sets from across organizations and sources into your own data system. Then you can use stream processing to trigger an application response to the data set. By representing each data product as a data stream, you can tag it with sub-elements and secondary dimensions to enable data searchability. If you are using a managed service like Confluent Cloud for data mesh, you can visualize how data flows inside the mesh through a stream lineage graph. 

    Ben also discusses the importance of keeping data architecture as simple as you can to avoid derivatives of data products.

    EPISODE LINKS

    • Data Mesh 101 course
    • Data Mesh 101 with Live Walkthrough Exercise
    • Introduction and Guide to Data Mesh
    • The Definitive Guide to Building a Data Mesh with Event Streams
    • What is Data Mesh, and How Does it Work? ft. Zhamak Dehghani
    • Designing Event-Driven Systems
    • Watch the video version of this podcast
    • Join the Confluent Community
    • Learn more with Kafka tutorials, resources, and guides at Confluent Developer
    • Live demo: Intro to Event-Driven Microservices with Confluent
    • Use PODCAST100 to get an additional $100 of free Confluent Cloud usage (details)

    SEASON 2
    Hosted by Tim Berglund, Adi Polak and Viktor Gamov
    Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
    Music by Coastal Kites 
    Artwork by Phil Vo 

    •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
    • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
    • 👍 If you enjoyed this, please leave us a rating. 
    • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
    45 min
  • Serverless Stream Processing with Apache Kafka ft. Bill Bejeck

    What is serverless?

    Having worked as a software engineer for over 15 years and as a regular contributor to Kafka Streams, Bill Bejeck (Integration Architect, Confluent) is an Apache Kafka® committer and author of “Kafka Streams in Action.” In today’s episode, he explains what serverless and the architectural concepts behind it are. 

    To clarify, serverless doesn’t mean you can run an application without a server—there are still servers in the architecture, but they are abstracted away from your application development. In other words, you can focus on building and running applications and services without any concerns over infrastructure management. 

    Using a cloud provider such as Amazon Web Services (AWS) enables you to allocate machine resources on demand while handling provisioning, maintenance, and scaling of the server infrastructure. 

    There are a few important terms to know when implementing serverless functions with event stream processors: 

    • Functions as a service (FaaS)
    • Stateless stream processing
    • Stateful stream processing

    Serverless commonly falls into the FaaS cloud computing service category—for example, AWS Lambda is the classic definition of a FaaS offering. You have a greater degree of control to run a discrete chunk of code in response to certain events, and it lets you write code to solve a specific issue or use case. 

    Stateless processing is simpler in comparison to stateful processing, which is more complex as it involves keeping the state of an event stream and needs a key-value store. ksqlDB allows you to perform both stateless and stateful processing, but its strength lies in stateful processing to answer complex questions while AWS Lambda is better suited for stateless processing tasks. 

    By integrating ksqlDB with AWS Lambda together, they deliver serverless event streaming and analytics at scale.

    EPISODE LINKS

    • What is Serverless?
    • Serverless Stream Processing with Apache Kafka, AWS Lambda, and ksqlDB
    • Stateful Serverless Architectures with ksqlDB and AWS Lambda 
    • Serverless GitHub repository
    • Kafka Streams in Action
    • Watch the video version of this podcast
    • Join the Confluent Community
    • Learn more with Kafka tutorials, resources, and guides at Confluent Developer
    • Live demo: Intro to Event-Driven Microservices with Confluent
    • Use PODCAST100 to get an additional $100 of free Confluent Cloud usage (details)

    SEASON 2
    Hosted by Tim Berglund, Adi Polak and Viktor Gamov
    Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
    Music by Coastal Kites 
    Artwork by Phil Vo 

    •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
    • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
    • 👍 If you enjoyed this, please leave us a rating. 
    • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
    43 min
  • The Evolution of Apache Kafka: From In-House Infrastructure to Managed Cloud Service ft. Jay Kreps

    When it comes to Apache Kafka®, there’s no one better to tell the story than Jay Kreps (Co-Founder and CEO, Confluent), one of the original creators of Kafka. In this episode, he talks about the evolution of Kafka from in-house infrastructure to a managed cloud service and discusses what’s next for infrastructure engineers who used to self-manage the workload. 

    Kafka started out at LinkedIn as a distributed stream processing framework and was core to their central data pipeline. At the time, the challenge was to address scalability for real-time data feeds. The social media platform’s initial data system was built on Apache™Hadoop®, but the team later realized that operationalizing and scaling the system required a considerable amount of work. 

    When they started re-engineering the infrastructure, Jay observed a big gap in data streaming—on one end, data was being looked at constantly for analytics, while on the other end, data was being looked at once a day—missing real-time data interconnection. This ushered in efforts to build a distributed system that connects applications, data systems, and organizations for real-time data. That goal led to the birth of Kafka and eventually a company around it—Confluent.

    Over time, Confluent progressed from focussing solely on Kafka as a software product to a more holistic view—Kafka as a complete central nervous system for data, integrating connectors and stream processing with a fully-managed cloud service.

    Now as organizations make a similar shift from in-house infrastructure to fully-managed services, Jay outlines five guiding points to keep in mind: 

    1. Cloud-native systems abstract away operational efforts for you without infrastructure concerns
    2. It’s important to have a complete ecosystem for Kafka, including connectors, a SQL layer, and data governance
    3. A distributed system should allow data to be accessible everywhere and across organizations
    4. Identifying a reliable storage infrastructure layer that is dependable, such as Amazon S3 is critical
    5. Cost-effective models mean sustainability and systems that are easy to build around


    EPISODE LINKS

    • Building Real-Time Data Systems the Hard Way
    • Kris Jenkins Twitter
    • The Hitchhiker’s Guide to the Galaxy
    • Hedonic treadmill
    • Watch the video version of this podcast
    • Join the Confluent Community
    • Learn more with Kafka tutorials, resources, and guides at Confluent Developer
    • Live demo: Intro to Event-Driven Microservices with Confluent
    • Use PODCAST100 to get an additional $100 of free Confluent Cloud usage (details)

    SEASON 2
    Hosted by Tim Berglund, Adi Polak and Viktor Gamov
    Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
    Music by Coastal Kites 
    Artwork by Phil Vo 

    •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
    • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
    • 👍 If you enjoyed this, please leave us a rating. 
    • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
    47 min

About Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

From the publisher's feed

Hi, we’re Tim Berglund, Adi Polak, and Viktor Gamov and we’re excited to bring you the Confluent Developer podcast (formerly “Streaming Audio.”) Our hand-crafted weekly episodes feature in-depth…

More shows like Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

Software Engineering Radio - the podcast for professional software developers by team@se-radio.net (SE-Radio Team)

Software Engineering Radio - the podcast for professional software developers

273 Listeners

The Changelog: Software Development, Open Source by Changelog Media

The Changelog: Software Development, Open Source

286 Listeners

Software Engineering Daily by Software Engineering Daily

Software Engineering Daily

623 Listeners

Data Engineering Podcast by Tobias Macey

Data Engineering Podcast

144 Listeners

The Daily by The New York Times

The Daily

111,766 Listeners