MBA Training Data

MBA Training Data

Download on the App Store

MBA Training Data episodes

  • AI agents querying your warehouse: measure before rebuilding

    Fivetran and dbt Labs used dbt Summit to argue that AI agents, not human analysts, are now the main consumer of enterprise data. The position here is that the architectural shift is real but the urgency is a sales pitch, and that re-platforming off a keynote is the expensive mistake.

    You come away able to test dbt Labs' claim that 30% of warehouse queries are machine-generated against your own logs, spot the failure mode where agents pass nonsense results downstream with no scar tissue, make the case for a written semantic layer over new tooling, and set a query budget before an agent touches production. References include O'Reilly Radar and MIT Sloan Management Review.

    Key takeaways

    • Pull last month's query logs and calculate what share came from service accounts rather than named humans to find your real agent load.
    • Treat dbt Labs' 30% machine-generated query figure as directional, since the vendor sells the tooling that raises it.
    • If agent traffic is small, spend the quarter writing business definitions down instead of changing architecture.
    • Build the semantic layer so terms like active customer are defined in writing, because an agent has no memory to fill the gaps.
    • Set a spending cap on agent query volume before a pilot touches production data to avoid a five-figure compute surprise.

    Chapters
    0:00 Agents as the new data consumer
    0:44 What changes when a bot queries
    1:12 dbt's 30% machine-generated query claim
    2:03 Metadata gaps and the semantic layer
    3:04 Compute costs and query budgets
    3:49 Check service account query share

    Go deeper, free lessons

    • dbt (data build tool): industrialized SQL transformation
    • Active metadata and the data catalog
    • The metrics & semantic layer
    • Data lineage & metadata management: knowing where your data was born
    • Data products: definition, design & lifecycle management

    Sources

    • AI coding agents need a secrets-safe context boundary
    • Fivetran + dbt Labs Announces New Capabilities to Make Enterprise Data Agent-Ready at dbt Summit 2026
    • Everything we announced at dbt Summit and why it matters
    • We built dbt State to stop rebuilding what hadn't changed
    • Celebrating the 2026 dbt partner of the year winners
    • Enterprise Analytics Beyond Dashboards: Intelligent Data Orchestration with LLMs
    • Spot New Tech Skills Emerging From the Workforce
    • Building on AI’s Unfinished Foundation

    Full article and transcript: https://www.mba-training.com/blog/agent-ready-data-infrastructure-dbt

    MBA Training, mba-training.com

    5 min
  • AML pipelines: three design calls examiners test

    Why do fintech AML pipelines fail regulatory examination when the models themselves are sound? The position here is that examiners do not score area under the curve, they ask reproducibility questions, and most data architectures cannot rebuild a March alert with March data. Three design decisions decide the outcome: immutable timestamped feature storage, documentation that records who approved a feature and which regulation it maps to, and explanation captured at inference time rather than manufactured later.

    You walk away knowing why dbt column level lineage maps transformations but not business intent, why SHAP values may be plausible rather than faithful, and how to run a one hour reproduction test on your last high value alert.

    Key takeaways

    • Freeze every model input as it existed on the decision date so an alert can be rerun with that day's data.
    • Keep design records for each feature covering who built it, when, why, and which regulation it maps to.
    • Log contributing features live at inference time instead of reconstructing a reason with post hoc explainers.
    • Treat vendor lineage statistics, such as dbt Labs' 80% column level lineage claim, as marketing until checked against your own graph.
    • Take your last high value AML alert and try to reproduce inputs, output and reason in under an hour.

    Chapters
    0:00 The March alert nobody could explain
    1:08 Why model accuracy does not pass exams
    2:03 Lineage tooling cannot show business intent
    3:18 Post hoc explainers versus inference-time logging
    4:16 Retrofit costs and a Monday reproduction test

    Go deeper, free lessons

    • AML, KYC and sanctions screening in practice
    • Reading the transaction ledger: what payment and behavioral data reveal
    • Governing fintech data: consent, lineage, and regulatory defensibility
    • the regulatory map every fintech data leader must carry
    • The regulatory bodies map: who enforces what and how examinations work

    Full article and transcript: https://www.mba-training.com/blog/fraud-kyc-aml-detection-pipeline-fintech

    MBA Training, mba-training.com

    6 min
  • DuckDB and text-to-SQL: redesign the analytics team

    If a manager can type "show me last quarter's churn by region" and get an answer in ten seconds, what is the analytics team for? The position here is that DuckDB plus language models that write SQL is a workforce redesign, not a tooling upgrade. The ad hoc request queue goes first, dashboard building is half gone, and the analytics engineer who owns the semantic layer becomes the centre of the team.

    You get a concrete redesign of a team of eight, the reason a model will compute churn three contradictory ways, the caveat on dbt Labs and O'Reilly survey numbers, and a Monday test: count how many live definitions exist for your five most quoted metrics.

    Key takeaways

    • Treat natural-language BI as a job redefinition rather than a headcount cut.
    • Check that generated SQL is visible so analysts can verify the number and the query behind it.
    • Fund owners of the semantic layer, the single blessed definition of churn, revenue and active user.
    • Rebalance a team of eight from five juniors, two seniors and one engineer toward two hard-analysis roles and four definition owners.
    • Count how many distinct definitions are live for your five most quoted metrics; more than one means you have a dictionary problem, not a tooling problem.

    Chapters
    0:00 Why analysts fear natural-language BI
    0:41 What changed: DuckDB plus text-to-SQL
    1:25 The standard analytics team org chart
    2:21 Why dashboards and metric definitions survive
    3:42 Redesigning a team of eight
    4:18 Monday morning: audit your five metrics

    Go deeper, free lessons

    • Modern BI: tools, maturity & the semantic layer
    • The metrics & semantic layer
    • Self-serve analytics: architecture, data catalog & data literacy
    • The lakehouse: unifying analytics & ML
    • Data warehouse, lake & lakehouse: choosing the right architecture

    Sources

    • Building a Data Lakehouse with DuckDB and DuckLake
    • What’s So Good About ChatGPT Work? Here’s What I Found
    • 5 Free Zoomcamps From Data Pipelines to AI Agents
    • “Everyone’s in a race to replace GitHub”: Zed launches Delta because agents made pull requests obsolete
    • Fivetran + dbt Labs Announces New Capabilities to Make Enterprise Data Agent-Ready at dbt Summit 2026
    • Everything we announced at dbt Summit and why it matters
    • We built dbt State to stop rebuilding what hadn't changed
    • Celebrating the 2026 dbt partner of the year winners

    Full article and transcript: https://www.mba-training.com/blog/natural-language-bi-analyst-role-duckdb

    MBA Training, mba-training.com

    5 min
  • Grey market leaks: Richemont's serial number fix

    Genuine Cartier and IWC pieces were surfacing in unauthorised Asian markets at 20 to 35 percent discounts, and the villain was Richemont's own distributors. The position taken here is blunt: anti-counterfeiting gets the attention, but the grey market is the larger financial wound, and distribution integrity is a data governance problem wearing a legal costume.

    You get the sequence Richemont followed to build a product-level data layer where each serial number is a live record updated at every handoff, the 40-odd maison fight over one shared definition of a serial event, and why enforcement worked as deterrence rather than analytics. Referenced sources include dbt Labs on 18-month data modeling rebuilds, O'Reilly Radar and MIT Sloan Management Review on ownership failures.

    Key takeaways

    • Agree one shared definition of a product record and serial event across every brand before spending on traceability tooling.
    • Give the owner of the serial schema enough authority to overrule a maison CEO, or the system will go unenforced.
    • Treat contract clauses as decoration unless you can name and prove the specific authorized account behind a leak.
    • Use traceability as a deterrent: once distributors know over-ordering can be traced and contracts terminated, behaviour changes.
    • Pull ten units of your top product and write down which authorized account sold each one; failure means you have a data problem today.

    Chapters
    0:00 Genuine IWC watches discounted in Bangkok
    1:00 Why wholesale arbitrage is hard to admit
    1:57 The governance cost of one serial schema
    2:57 Deterrence beat detection in the grey market
    3:44 Define the product record before buying tooling
    4:25 Trace ten units of your top product

    Go deeper, free lessons

    • Why counterfeiting is a legal battle, not just a copycat problem
    • Controlling distribution to protect desirability
    • Sell-in vs sell-out: reconciling wholesale and retail truth
    • The anti-money-laundering checks behind a six-figure watch sale
    • Modeling scarcity and allocating waitlisted hero products

    Full article and transcript: https://www.mba-training.com/blog/richemont-authentication-grey-market-data

    MBA Training, mba-training.com

    6 min
  • Data flywheel: why loop latency beats data volume

    Board decks show a circular arrow labelled data flywheel, but few teams can say what spins it. This episode separates compounding from speed: a faster pipeline processes the same low value data quicker, while a real flywheel makes each new record more valuable because of the data already held, as Google search ranking does with click data. It examines the dbt Labs claim that pipeline automation cuts time to insight by around 60 percent, plus findings from MIT Sloan Management Review on flattening accuracy gains and O'Reilly Radar on warehouses that never feed decisions.

    You leave with three failure modes to test for, data saturation, data decay and a loop that never reaches the product, and a single metric to measure this week: the days between data arriving and a decision changing.

    Key takeaways

    • Judge a flywheel by whether each new data point is worth more because of the data you already hold, not by pipeline speed.
    • Check for data saturation: if marginal accuracy gains have flattened, stop paying to store more of the same pattern.
    • Treat ageing data as a liability, since behaviour from 2023 predicts little about 2026 spending after the inflation shock.
    • Pick your most important data driven decision, such as pricing, recommendations or churn, and measure the days between data arriving and the decision changing.
    • If that latency is over a week, cut it in half before spending anything on collecting new data.

    Chapters
    0:00 What the data flywheel actually claims
    0:51 dbt Labs pipeline speed claim, examined
    1:55 Three places the flywheel breaks
    3:14 Why Amazon and Netflix loops compound
    4:00 Measure your decision latency Monday

    Go deeper, free lessons

    • Data monetization: three modes & the data flywheel
    • CDO in retail & e-commerce: the data flywheel
    • Calculating the ROI of data initiatives
    • Data as a strategic asset: how to put a number on it
    • Communicating with the board and c-suite

    Full article and transcript: https://www.mba-training.com/blog/data-flywheel-compounding-advantage

    MBA Training, mba-training.com

    5 min
  • MES and ERP integration: the real Siemens Amberg win

    Siemens Amberg is usually told as a robotics story, with 99.99885% quality and a heavily automated line. The position here is different: the recent gains came from making the Manufacturing Execution System and ERP agree on reality, not from more sensors. The route was treating manufacturing data as a product with owners, contracts and quality checks, rather than chasing one monolithic platform.

    You get the specific failure mode to look for, where MES counts a unit complete when it leaves the line and ERP counts it when booked to inventory, and what that gap does to supply forecasts. References include MIT Sloan Management Review, O'Reilly Radar and dbt Labs, plus a first project any CDO can start with a meeting.

    Key takeaways

    • Agree on shared definitions of a unit, a batch and a defect across MES and ERP before funding any dashboard.
    • Treat manufacturing data as a product with named owners, data contracts and quality checks instead of building one monolithic platform.
    • Measure success as speed of production response, hours instead of weeks, rather than reporting cycle time.
    • Pull MES and ERP side by side and compare yesterday's unit counts; the discrepancy is your first project.
    • Copy the reconciliation logic, not the robot fleet, since the portable part costs a fraction of the automation spend.

    Chapters
    0:00 What the Siemens Amberg case studies leave out
    1:09 Why more automation is the wrong lesson
    2:13 Hours instead of weeks: the real Amberg effect
    2:47 Dashboards fail without data contracts
    3:59 Does this transfer to mid-sized plants?
    4:33 Monday morning MES and ERP unit count

    Go deeper, free lessons

    • Master data done right: materials, BOMs, and equipment hierarchies
    • Mapping the manufacturing data landscape: sources, systems, and silos
    • Building end-to-end traceability across the supply chain
    • Measuring what matters with OEE and quality metrics
    • Data quality metrics that matter on the shop floor

    Full article and transcript: https://www.mba-training.com/blog/mes-erp-unified-manufacturing-data-backbone

    MBA Training, mba-training.com

    6 min
  • ML in production: why models die after the notebook

    Why do most machine learning projects never leave the lab? The position here is blunt: the model is about ten percent of the work and the other ninety percent is pipelines, monitoring and ownership. The episode walks through the three failure points at the notebook to production handoff, data drift, unowned pipelines after launch, and systems never designed for retraining, and why a drifting model costs money quietly instead of paging you at midnight.

    You come away able to name the roles that fix this, the ML engineer and the data engineer, read vendor claims from DBT Labs against O'Reilly Radar and MIT Sloan Management Review, apply data as a product, and ask the accountability question before funding an AI project.

    Key takeaways

    • Before approving any model, ask who by name is on the hook to keep it running six months after launch.
    • Budget for an ML engineer who sits between the data scientist and the software team, not just modellers.
    • Treat each dataset as a product with an owner, a quality standard, documentation and a promise it will not silently change.
    • Monitor for data drift, because a degrading model keeps returning confident answers instead of crashing.
    • Discount vendor productivity claims such as DBT Labs surveys and cross check them against O'Reilly Radar or MIT Sloan research.

    Chapters
    0:00 Why ML projects die after the notebook
    1:10 Three places the handoff snaps
    2:02 The ML engineer nobody budgets for
    2:25 Data as a product, not byproduct
    3:25 Why buying a platform does not help
    4:07 The one question to ask before greenlighting

    Go deeper, free lessons

    • Closing the POC-to-production gap
    • Models in production: drift, monitoring & MLOps
    • MLOps: monitoring, retraining & drift
    • The lakehouse: unifying analytics & ML
    • Data observability: detect problems before your users

    Sources

    • Building a Data Lakehouse with DuckDB and DuckLake
    • What’s So Good About ChatGPT Work? Here’s What I Found
    • 5 Free Zoomcamps From Data Pipelines to AI Agents
    • “Everyone’s in a race to replace GitHub”: Zed launches Delta because agents made pull requests obsolete
    • Fivetran + dbt Labs Announces New Capabilities to Make Enterprise Data Agent-Ready at dbt Summit 2026
    • Everything we announced at dbt Summit and why it matters
    • We built dbt State to stop rebuilding what hadn't changed
    • Celebrating the 2026 dbt partner of the year winners

    Full article and transcript: https://www.mba-training.com/blog/ml-pilots-production-field-guide

    MBA Training, mba-training.com

    5 min
  • Data literacy: measure decision quality, not training

    Why do data literacy programs teach medians, filters and chart types and still leave the business running on gut feeling? The position here is blunt: knowledge and behavior are different things, and almost nobody measures the second one. The fix is to build training around real recurring decisions like pricing, hiring and inventory, and to require anyone proposing a decision to state in advance what evidence would change their mind.

    You come away with a way to score decision quality across ten recurring decisions before and after a program, a read on Booking.com's experimentation default, and evidence from MIT Sloan Management Review, O'Reilly Radar and DBT Labs on why culture and trust, not skills, are the bottleneck.

    Key takeaways

    • Pick ten recurring decisions, record whether each was made on gut, politics or evidence, then measure again six months later.
    • Require the person proposing a decision to write down what evidence would change their mind before the meeting.
    • Start the literacy program with the twelve executives who actually make the calls, not with five thousand employees.
    • Tie learning to the actual workflow, like a plant manager changing one reorder decision on a forecast, rather than a training calendar.
    • Treat vendor data such as DBT Labs' claims with care and cross-check against independent work like O'Reilly's annual surveys.

    Chapters
    0:00 Why most data literacy programs fail
    1:01 Designing training around real decisions
    1:50 Booking.com experimentation and workflow learning
    2:22 Measuring decision quality over six months
    3:15 Executive committee is the real bottleneck
    4:06 Monday morning test for any decision

    Go deeper, free lessons

    • Building a data literacy program
    • Data literacy: building analytical capability across the organization
    • Calculating the ROI of data initiatives
    • Communicating with the board and c-suite
    • Driving cultural change to data-driven

    Full article and transcript: https://www.mba-training.com/blog/data-literacy-behavior-change-roi

    MBA Training, mba-training.com

    5 min
  • GDPR retention and minimization: the real compliance gap

    Is a consent banner and a privacy notice enough to call a GDPR program finished? The position here is no. Consent management got budget because customers see it, while retention and data minimization stayed invisible and unfunded. Modern data tooling in 2026, including lineage work pushed by vendors like dbt Labs, now makes every copy of a customer email visible, and research from O'Reilly and MIT Sloan Management Review points to personal data being kept with no defined deletion schedule.

    You walk away able to treat unused personal data as a liability, defend against data subject access requests, hash or aggregate identifying fields at ingestion, and set a retention schedule on your largest personal data table.

    Key takeaways

    • Check whether every row in your largest personal data table has a defined deletion date, and set a retention schedule on that one table this week.
    • Treat unused personal data as a growing liability rather than an appreciating asset, because retained records can be stolen, subpoenaed or fined.
    • Strip or hash identifying fields at ingestion so analytics keeps its value while the liability drops, instead of cleaning up data later.
    • Discount vendor framing on lineage and compliance, since data transformation companies sell the tooling that exposes the problem.
    • Assume data subject access requests will be unmanageable if personal data sits in dozens of systems nobody has mapped.

    Chapters
    0:00 Why GDPR compliance claims are self-deceived
    0:54 How 2026 data lineage tools expose hoarding
    1:48 Retention with no deletion schedule
    2:19 Storage is cheap but breaches are not
    3:15 Minimization at ingestion beats deleting later
    4:06 Monday morning: set one retention schedule

    Go deeper, free lessons

    • GDPR in practice: the 10 mistakes CDOs make most often
    • CCPA, LGPD, AI Act: navigating the global regulatory patchwork
    • Data classification & access control: the zero-trust data approach
    • Data lineage & metadata management: knowing where your data was born
    • Data ethics: beyond compliance, toward institutional trust

    Full article and transcript: https://www.mba-training.com/blog/gdpr-retention-minimization-compliance

    MBA Training, mba-training.com

    5 min
  • Alternative credit scoring: how Tala lends without bureaus

    Can you underwrite a loan for someone who has no credit bureau file at all? Tala says yes, and the episode explains how: permissioned smartphone data such as contact counts, airtime top-up timing and money movement frequency, used as behavioral proxies for the habits a bureau score normally captures. The position taken here is that Tala's edge was never a cleverer algorithm. It was owning a data source nobody else collected, a point MIT Sloan Management Review has argued for years.

    You come away knowing how model drift erodes a live lending model, why retraining and freshness tests matter more than model choice, and the limit of behavioral data once borrowers start performing for the algorithm. It closes with a one-week audit of your three most important data sources.

    Key takeaways

    • Treat "no credit history" as a format problem, not an absence of data, and go looking for behavioral signals elsewhere.
    • Build the advantage in the data pipeline, not the algorithm, because proprietary data beats a marginally smarter model.
    • Retrain constantly and assume a signal that worked in Nairobi in 2022 may be noise by 2026.
    • Run automated freshness and quality tests on input data, or accept that your scoring is guesswork with a green dashboard.
    • Pick your three most important data sources and answer how you would know within a week if they went stale.

    Chapters
    0:00 Lending to 2 billion people without credit files
    0:42 Smartphone data as a behavioral proxy
    1:24 Why the data pipeline is the moat
    2:14 Model drift and testing your inputs
    3:27 Can a copycat buy this data advantage
    4:21 Audit your three most important sources

    Go deeper, free lessons

    • Underwriting the thin-file customer with alternative data
    • Consumer protection law: fair lending and disclosure rules
    • Reading the transaction ledger: what payment and behavioral data reveal
    • Data privacy and open finance: GDPR, CCPA and data-sharing consent
    • Bias, explainability & model cards

    Full article and transcript: https://www.mba-training.com/blog/tala-alternative-data-thin-file-underwriting

    MBA Training, mba-training.com

    6 min

About MBA Training Data

From the publisher's feed

MBA Training Data. Daily strategy in data governance, architecture, analytics and AI, for data leaders and aspiring CDOs. New episode every day at mba-training.com.