MBA Training Data

MBA Training Data

Download on the App Store

MBA Training Data episodes

  • Luxury waitlist allocation: building a defensible scarcity model

    Who decides that a Hermes Birkin buyer waits six years? The position here is blunt: scarcity is an asset a data team manages, and every "nothing available" is an allocation decision someone modeled. The episode separates manufactured scarcity from the allocation logic that ranks clients by lifetime spend, purchase breadth and loyalty, and looks at the US class action arguing that tying Birkin access to other purchases is illegal bundling.

    You come away with a three-layer design, scarcity forecast, eligibility stripped of protected proxies, and allocation ranking, plus why transformation tools like dbt give you lineage as a legal shield. Also covered: the Patek Philippe Nautilus 5711 discontinuation and reading waitlists as intent data for product planning.

    Key takeaways

    • Keep the scarcity forecast, the eligibility layer and the allocation ranking as three separate layers so any single decision can be audited.
    • Strip proximity variables such as postcode from eligibility rules, since they can smuggle protected characteristics into the model.
    • Document transformations so every allocation traces back to the rule that produced it, and treat vendor claims like dbt Labs' incident reduction figure as marketing to cross-check.
    • Feed waitlist data into product planning as a demand sensor by model, region and client tier instead of treating it as an apology backlog.
    • On Monday, test whether you can explain in writing why the last person got an offer and the person behind them did not.

    Chapters
    0:00 Why a Birkin waitlist is a data decision
    1:13 Client scoring and the Hermes bundling lawsuit
    2:04 Three-layer allocation architecture and dbt lineage
    3:06 Patek Nautilus 5711 and waitlists as demand sensors
    4:05 Keeping the algorithm out of the boutique

    Go deeper, free lessons

    • Why selling less makes luxury worth more
    • Building the single client view for high-net-worth luxury buyers
    • Controlling distribution to protect desirability
    • Advanced analytics: CLV, churn prediction & demand forecasting
    • Data-driven authentication and grey-market leakage tracking

    Full article and transcript: https://www.mba-training.com/blog/scarcity-modeling-waitlist-allocation-luxury

    MBA Training, mba-training.com

    6 min
  • Mistral's $3.5B raise: why open weights favor portability

    Mistral's September 2026 raise of $3.5 billion made open-weight models a credible enterprise option, and most data leaders read that as a signal to stop buying and start building. The position here is the opposite: capability parity raises the ceiling, not the floor, and the hard part was never the model. It is the data plumbing and the retrieval layer, a point O'Reilly's 2026 adoption research supports.

    You get the three conditions that justify self-hosting, the token volume crossover point DBT Labs published (with the caveat that they sell data tooling), and the MIT Sloan Management Review argument that open weights are worth having for optionality. You leave with a lock-in audit to run on your top three AI use cases.

    Key takeaways

    • Build with open weights only if you have proprietary data, an existing platform team that already operates services, and high predictable usage; missing one means don't build.
    • Check your own inference bills against the DBT Labs crossover numbers before assuming self-hosting is cheaper than paying per call.
    • Treat open-weight models as insurance and switching optionality, using them as leverage in vendor negotiations rather than running them yourself.
    • Audit your top three AI use cases and ask whether you could swap the underlying model in under a month without a rewrite.
    • Fix the data plumbing and retrieval layer first, since the model is rarely the bottleneck.

    Chapters
    0:00 Mistral's $3.5 billion raise in context
    0:29 Open weights versus open source explained
    0:56 The hidden cost of self-hosting models
    2:07 Three conditions for building in-house
    2:41 Token volume crossover: API versus owning hardware
    3:19 Optionality as insurance and Monday's audit

    Go deeper, free lessons

    • Build vs buy: RAG vs fine-tuning
    • CDO AI strategy: prioritization, build/buy & value chain
    • Generative AI in the enterprise: RAG, risks & governance
    • LLMOps & evaluation
    • Measuring AI ROI

    Full article and transcript: https://www.mba-training.com/blog/enterprise-generative-ai-build-vs-buy

    MBA Training, mba-training.com

    5 min
  • Decision intelligence: engineer the decision, not the dashboard

    Why do organisations with hundreds of dashboards still decide badly? The answer here is that the numbers arrive nowhere near the moment of choice. A support agent deciding on a refund needs lifetime value, refund history and churn risk inside the ticket, not in a tool they have no login for. Embedded analytics is the plumbing; decision intelligence is treating the recurring decision as the thing you design around, then working backwards.

    You walk away with three tests for picking a decision worth engineering (frequent, consequential, currently made blind), a rule for what to leave alone, and a warning on reading DBT Labs vendor figures, with supporting points from O'Reilly Radar and MIT Sloan Management Review.

    Key takeaways

    • Design backwards from a recurring decision instead of building a dashboard and hoping someone consults it.
    • Apply three tests before investing: is the decision frequent, is it consequential, and is it currently made with no data at the point of choice.
    • Encode the logic for repeatable approve or reject calls and route the odd 5 percent to a human, rather than automating messy judgement calls.
    • Leave rare one-off strategic bets alone; those are argued in a room, not instrumented.
    • Sit beside the person making your most expensive recurring decision and check whether the data they need is on screen at that second.

    Chapters
    0:00 Why dashboards get admired and ignored
    1:04 Embedded analytics versus decision intelligence
    2:09 Why full automation is the trap
    2:43 DBT Labs numbers and vendor caution
    3:30 Three tests for decisions worth engineering
    4:13 Sit with the decision maker Monday

    Go deeper, free lessons

    • Decision intelligence: decision architecture & embedded analytics
    • Embedding analytics into workflows
    • From dashboards to decisions
    • Decision rituals: getting data into the room
    • The metrics & semantic layer

    Full article and transcript: https://www.mba-training.com/blog/decision-intelligence-embedded-analytics-cdo

    MBA Training, mba-training.com

    5 min
  • GDPR versus GxP audit trails: the deletion trap

    Pharma data sits under three rulebooks at once: GxP quality regulation, the FDA's 21 CFR Part 11 electronic records rule, and GDPR. The first two say keep everything and prove everything. The third says erase on request. This episode takes the position that the damage happens where they conflict, when a right to be forgotten workflow reaches into a Part 11 audit trail and makes it look modified to an inspector.

    You come away able to audit your own deletion path, apply pseudonymization with a separately held key, use the GDPR research exemption by design, and argue why the chief data officer should own all three regimes. References include MIT Sloan Management Review and dbt Labs.

    Key takeaways

    • Trace your GDPR deletion workflow end to end and block it from touching any GxP or Part 11 audit trail.
    • Delete the link between person and record, never the record itself.
    • Use pseudonymization with the identity key stored separately and locked down so trial data stays intact.
    • Design for the GDPR research exemption upfront rather than discovering it during an FDA inspection.
    • Give one person, usually the chief data officer, authority across quality, privacy and records so no department owns the conflict alone.

    Chapters
    0:00 Three regimes pointing at one record
    1:03 What GxP demands from the audit trail
    1:54 GDPR erasure against permanent retention
    2:39 How deletion scripts break Part 11 records
    3:13 Pseudonymization and the GDPR research exemption
    4:18 Who owns the conflict, and Monday's fix

    Go deeper, free lessons

    • GxP explained: the quality rulebook behind every batch of pills
    • Global privacy regimes and what they mean for pharma data flows
    • Running a data audit: from access logs to inspection readiness
    • Consent, de-identification and the limits of anonymous data
    • Building a pharma data governance operating model

    Full article and transcript: https://www.mba-training.com/blog/gxp-21cfr-part11-gdpr-pharma-cdo

    MBA Training, mba-training.com

    6 min
  • Open data portals: publish fewer datasets as products

    Government portals hold thousands of datasets that get downloaded once and never opened again. The position here is that dataset count is a vanity metric and repeat use is the real measure of a portal's health, citing Data.gov's 300,000 plus datasets and their dead long tail. The answer is to treat a dataset like a product with an owner, a release schedule and documentation, and to retire the rest.

    You will learn how to separate the needs of journalists, advocates and oversight bodies, why Transport for London's live feed produced CityMapper and tens of millions of pounds in value, how to read public records requests as a demand signal, and why a plain language data dictionary decides whether a dataset gets cited.

    Key takeaways

    • Measure repeat downloads and living audiences instead of the number of datasets published.
    • Pick one audience per dataset, since journalists, advocates and auditors need opposite things.
    • Retire 90 percent of datasets and fund the ten people fight over, such as spending, permits, inspections and police stops.
    • Use recurring public records requests as a free backlog of what to publish well.
    • Open your download logs, find the five datasets people return to, and give each a named owner and publishing schedule.

    Chapters
    0:00 Why open data portals become graveyards
    0:44 Repeat downloads beat dataset counts
    1:12 Journalists, advocates and auditors want different things
    2:08 Transport for London and CityMapper
    3:20 Use records requests to pick datasets
    3:46 Data dictionaries and named dataset owners

    Go deeper, free lessons

    • Freedom of Information and open records: what becomes public and when
    • Data quality scorecards for government datasets
    • Mapping the public data landscape: registries, admin records, and survey data
    • Governing privacy and equity on legacy systems
    • Managing stakeholders and public accountability

    Full article and transcript: https://www.mba-training.com/blog/open-data-publishing-public-sector-playbook

    MBA Training, mba-training.com

    6 min
  • Fanatics data platform ROI: making the board care

    How does a chief data officer get a board to fund a data platform instead of tolerating it? The answer here is attribution. Fanatics stopped presenting infrastructure cost and presented measurable business output, treating the platform as a profit and loss line. The episode walks through the use cases they attached numbers to, including pricing, inventory markdowns and fraud detection during big jersey drops, and flags that the widely cited dbt Labs figures come from a vendor and need cross-checking against sources like MIT Sloan Management Review.

    You finish able to pick one money decision, price the delta your data creates, and present it as a conservative, likely and optimistic range rather than a single hero number a CFO can dismantle.

    Key takeaways

    • Present the data platform as a profit and loss line tied to output, not an IT cost line tied to spend.
    • Attribute value to specific decisions such as pricing frequency, inventory markdowns and fraud detection during high volume drops.
    • Treat vendor case studies, including dbt Labs figures, as marketing until independently verified against sources like MIT Sloan Management Review.
    • Give the CFO conservative, likely and optimistic ranges instead of one precise number you cannot defend under questioning.
    • Pick one business decision that already involves money, calculate what your data changed about it, and take that number to the CFO first.

    Chapters
    0:00 How Fanatics won board attention for data
    0:37 Reframing the data platform as a P&L line
    1:34 Why vendor case studies need cross-checking
    2:38 Building a defensible dollar figure from one decision
    3:34 Accountability to the P&L changes data culture
    4:18 The one number to bring your CFO

    Go deeper, free lessons

    • Calculating the ROI of data initiatives
    • Securing executive buy-in: the boardroom pitch framework
    • Internal data platforms as products
    • Data as a strategic asset: how to put a number on it
    • CDO in retail & e-commerce: the data flywheel

    Sources

    • After nine years as HashiCorp CEO, Dave McJannet now wants to “unblock” enterprise AI agents
    • This founder is making cheaper, cleaner steel
    • How to find failures without drowning in tracing data
    • Spot New Tech Skills Emerging From the Workforce
    • Building on AI’s Unfinished Foundation
    • Databricks processes your data. dbt defines what it means
    • dbt Core v1.12 is GA
    • Model for the token, not the table

    Full article and transcript: https://www.mba-training.com/blog/fanatics-data-platform-roi-board

    MBA Training, mba-training.com

    5 min
  • Load forecasting: why DER signals break utility models

    Utility teams still treat load forecasting as a regression against temperature and industrial schedules. This episode argues that world is gone. Rooftop solar can flip the sign of the weather relationship, a passing cloud can return 40 megawatts in 90 seconds, and a midnight EV rate discount can synchronize 10,000 chargers into a spike that never existed before. Bolting DER signals onto the old model produces a more complicated wrong answer.

    You come away knowing where the money leaks, through NERC reliability reports and rate cases where a 2% error becomes a nine-figure conversation, why endogeneity breaks price-sensitive models, what probabilistic forecasting with confidence bands gives operators, and why dbt Labs style version control keeps models auditable without telling you they are wrong.

    Key takeaways

    • List every rate change and demand response program launched in the last 18 months, then check whether your model has a variable that reacts to them.
    • Stop treating weather as a one-directional driver, since rooftop solar can cut afternoon demand instead of raising it.
    • Model EV charging as behavior shaped by your own price signals, not as a fixed 6 p.m. plug-in assumption.
    • Publish a range with confidence bands rather than a single megawatt number, so operators know when to keep a reserve plant warm.
    • Use pipeline tooling like dbt for auditability and speed of fixes, but cross-check vendor claims and treat model correctness as a human judgment.

    Chapters
    0:00 Why load forecasting stopped being solved
    0:46 Rooftop solar flips the weather relationship
    1:35 EV charging timers and rate-driven spikes
    2:25 Forecast error in NERC reports and rate cases
    3:30 Probabilistic forecasting and confidence bands
    4:26 Audit your rate changes Monday morning

    Go deeper, free lessons

    • From smart meters to grid telemetry: the energy data stack
    • Grid reliability rules: NERC standards and the cost of a blackout violation
    • Rate cases decoded: how utilities justify prices to their regulator
    • MLOps: monitoring, retraining & drift
    • How wholesale power markets set the price of electricity

    Full article and transcript: https://www.mba-training.com/blog/load-forecasting-utility-der-signals

    MBA Training, mba-training.com

    6 min
  • ELT stack: where dbt, ingestion and orchestration break

    How do the three layers of a modern ELT stack actually connect, and why do pipelines still produce wrong numbers when every tool in them is good? The position here is that ingestion, dbt and orchestration are each mature, but the joints between them are untested, and that is where bad data comes from.

    You get a plain definition of extract, load, transform and why cheap storage in Snowflake, BigQuery and Databricks flipped ETL around. You also get the cost picture on Fivetran monthly active rows versus open source Airbyte, what dbt adds through version control, tests and documentation, what Airflow and Dagster decide about run order, and a one hour test to catch upstream schema changes.

    Key takeaways

    • Load raw data untouched first so a logic error means retransforming, never refetching from the source.
    • Watch Fivetran monthly active row billing, which can reach six figures a year for a midsize company, and price Airbyte against it.
    • Use dbt to put transformations under version control, review and automated tests instead of leaving loose SQL queries around.
    • Let an orchestrator like Airflow or Dagster hold dbt back until all sources have landed, rather than letting ingestion trigger it directly.
    • Write a test at the ingestion to dbt boundary that fails loudly when a source column disappears or changes type.

    Chapters
    0:00 What ELT means and why it replaced ETL
    1:09 Ingestion tools: Fivetran pricing versus Airbyte
    1:50 dbt: SQL transformations treated as software
    2:44 Orchestration with Airflow and Dagster
    3:20 Schema drift at the ingestion to dbt seam
    4:15 The boundary test to build Monday

    Go deeper, free lessons

    • dbt (data build tool): industrialized SQL transformation
    • Data pipelines: ETL/ELT, batch, streaming and the Medallion architecture
    • Data observability: detect problems before your users
    • Data contracts: the new standard for quality agreements between teams
    • Data FinOps: controlling cloud data cost

    Full article and transcript: https://www.mba-training.com/blog/modern-elt-stack-dbt-ingestion-orchestration

    MBA Training, mba-training.com

    5 min
  • Data ROI to the board: why attribution models backfire

    Why do chief data officers present cost savings and revenue attribution to the board and still lose budget? The position here is that the problem is the mental model, not the numbers. Attribution percentages like "our model drove 12% of pipeline" invite the board to argue about credit, and manufactured precision reads as hiding rather than rigour. Cost savings is the language of a department fighting for survival.

    You walk away able to reframe the conversation around decisions improved and bets de-risked, using a retailer example where a demand model caught a forecast inflated by 30% before a nine-figure inventory commitment, and able to read Snowflake and Databricks ROI calculators with proper scepticism.

    Key takeaways

    • Stop framing data spend as justification, because justifying already concedes the cost centre frame the CFO never accepts.
    • Drop attribution percentages such as "12% of pipeline", since everyone in the room knows the number is manufactured.
    • Present a decision the company almost got wrong and the capability that changed the odds, rather than credit for a single outcome.
    • Treat vendor ROI calculators from platform companies as sales material, because their business case assumes perfect usage.
    • Before the next board meeting, pick the single biggest upcoming decision and show one clear before and after.

    Chapters
    0:00 Why cost savings dashboards lose budget
    0:44 Justifying spend loses the frame
    1:43 Retailer catches an inflated demand forecast
    2:29 Selling capability instead of claiming credit
    3:18 Why vendor ROI calculators mislead boards
    4:10 One decision to bring to the board

    Go deeper, free lessons

    • Calculating the ROI of data initiatives
    • Communicating with the board and c-suite
    • Measuring the business value of analytics: ROI and the business case
    • Data as a strategic asset: how to put a number on it
    • The data P&L

    Full article and transcript: https://www.mba-training.com/blog/proving-data-roi-board

    MBA Training, mba-training.com

    5 min
  • Data contracts: how JPMorgan Chase named data owners

    Who owns a dataset, and what happens when it breaks? JPMorgan Chase could not answer that for years across hundreds of business lines, so it replaced governance policy documents with data contracts covering more than 50 domains, each with a named person who signs off on format, quality, freshness and escalation. The position here: accountability without a name attached is theater.

    You get the design that avoids organizational paralysis, contracts strict at the core and light at the edges, where consumers subscribe to a published interface instead of renegotiating each request. You also get the failure mode being avoided, silent schema changes, the lineage case regulators care about, a caution on the vendor-quoted 60% cleanup figure from Monte Carlo and Anomalo, and a one-page contract to write this week.

    Key takeaways

    • Name a specific person who signs off that a dataset meets its contract, since making the data team responsible means nobody is.
    • Define a contract once at the publishing domain and let consumers subscribe, so you negotiate the interface rather than every table request.
    • Write the contract to cover format, field meanings, update frequency and who gets paged when it breaks.
    • Treat an undeclared schema change, like renaming customer_id, as a breach that is flagged before it ships.
    • Start with the one dataset that already causes fights between finance and marketing, one page and two signatures, instead of launching a 50-domain program.

    Chapters
    0:00 What a data contract actually is
    0:40 Why nobody owned the data
    1:17 Named owners beat team accountability
    2:19 Silent schema changes break dashboards
    2:50 Lineage, regulators and analyst cleanup time
    3:41 One-page data contract to start Monday

    Go deeper, free lessons

    • Data contracts: the new standard for quality agreements between teams
    • Data ownership, stewardship and accountability across the org
    • CDO in financial services: when regulation is your architecture
    • Data lineage & metadata management: knowing where your data was born
    • Data mesh: principles, success conditions & criticisms

    Full article and transcript: https://www.mba-training.com/blog/data-contracts-ownership-domains-jpmorgan

    MBA Training, mba-training.com

    5 min

About MBA Training Data

From the publisher's feed

MBA Training Data. Daily strategy in data governance, architecture, analytics and AI, for data leaders and aspiring CDOs. New episode every day at mba-training.com.