This Week in Open Lakehouse

This Week in Open Lakehouse

By Lisa Cao and Scott HainesNewsTechnologyTech News
Download on the App Store

This Week in Open Lakehouse episodes

  • Iceberg 1.12, OIDC Credentials, Spark Connect Gateway, Arrow JSON Schemas, and Release Management

    In this episode, Lisa and Scott explore the latest developments across the open data ecosystem, including Delta Rust, Apache Iceberg, Apache Spark, Rust-native infrastructure, and the challenges of keeping rapidly evolving technologies interoperable. They discuss release velocity, governance, security, schema evolution, standardization, and the maintenance work required to make open-source innovation reliable in production.

    • Delta Rust 1.0, DataFusion integration, column mapping, schema evolution, and reduced JVM dependencies
    • Apache Iceberg 1.12, variant types, variant shredding, V4 foundations, and performance improvements
    • Spark security proposals, row-level filters, column masking, trusted execution engines, and policy enforcement
    • OIDC credential propagation, short-lived credentials, identity management, and modern authentication
    • Spark Connect Gateway, proxy-driven architecture, multi-tenancy, routing, rate limits, and observability
    • Rust adoption, modular infrastructure, smaller binaries, faster startup times, and composable systems
    • Interoperability across Spark, Flink, DuckDB, Polars, Arrow, Iceberg, and Delta
    • Schema evolution, column mapping, positional assumptions, nested data, and logical versus physical representations
    • Standardized schemas, shared expression models, JSON representations, and reducing translation layers
    • Release candidates, backports, compatibility matrices, security patches, and the role of open-source maintainers
    1 hr 11 min
  • XML Functions, Polars 2.0, delta-rs, and Agent Thrashing Policies

    In this episode, Lisa and Scott explore recent updates in ecosystem technologies including Apache Iceberg, Apache Flink, Polars, delta-rs, highlighting their frustrations in agent workflows and a potential oversight on the importance of XML.

    • Iceberg 1.12: REST catalogs, C++ support, equality deletes, and interoperability
    • Polaris catalog governance, lineage, pagination, and deletion policies
    • Flink updates, including State Fun deprecation and native XML functions
    • Polars 2.0, streaming execution, lazy frames, and memory efficiency
    • Delta Rust releases, partition handling, correctness fixes, and migration
    • Parquet performance, FastLanes benchmarks, and Arrow’s Intel macOS support
    • AI agent orchestration with Omnigent, Kubernetes sandboxes, Spark pipelines, and thrash detection
    1 hr 7 min
  • FileTypes, Language-Agnostic UDFs, Parquet Versioning, and Iceberg Updates

    In this episode, Lisa and Scott explore recent updates in data ecosystem technologies including Apache Iceberg, Apache DataFusion, Apache Parquet, and the Lakekeeper Iceberg Rest Catalog, highlighting their implications for data management and interoperability.

    • Iceberg Data Fusion integration
    • Parquet 2.14 release and versioning
    • Nanosecond timestamp support in Parquet
    • Iceberg 1.12 release and V4 specification
    • Iceberg rest catalog support in Unity Catalog
    • Lakekeeper updates and multimodal support

    24 min
  • Nanosecond Timestamp Precision, Iceberg-Datafusion,Spark Connect Rust Client, and Parquet Types

    The biggest throughline this week is format-layer consolidation. Apache Parquet 2.14.0 shipped as a final release, carrying chronological ordering for INT96 timestamps and Variant documentation fixes, while a parallel dev-list debate over versioning semantics and reader behavior for unsupported versions shows the community is actively thinking about long-term compatibility contracts, not just feature additions.

    One layer up, Lance v12 is in a sustained beta sprint, shipping six beta releases in a single week. The work spans file format stabilization (resolving the stable format to 2.2), distributed vector search in Java, in-memory WAL improvements, and IVF index optimization. This is not routine maintenance; it reads as a deliberate push to lock down a production-ready v12 surface before a GA cut.

    At the query engine and integration layer, two parallel votes to move iceberg-rust's DataFusion integration into the Apache DataFusion project are the most structurally interesting governance event of the week. If both votes pass, Iceberg catalog access becomes a first-class DataFusion concern rather than a satellite project, which changes how engine builders think about Iceberg adoption. Separately, the Spark Connect Rust client reached RC2 for its 4.2.0 release, and Spark 4.3.0 RC1 is in vote, keeping the Spark release cadence unusually active.

    1 hr 17 min

About This Week in Open Lakehouse

From the publisher's feed

This Week in Open Lakehouse is a weekly data engineering podcast discussing the latest in open source news, including query engines, data catalogs, open table formats, orchestrators, and streaming…