Leaders Insights — Data

Leaders Insights — Data

By Leaders InsightsBusinessEducationManagement
Download on the App Store

Leaders Insights — Data episodes

  • SB 947: proving who reviewed an AI firing call

    California's No Robo Bosses Act bars employers from disciplining or firing a worker on an automated system's output alone, and from 1 July 2027 a human must independently corroborate that output in writing. The position here is that this reads as an HR matter but lands on data teams, because the statute covers statistical modeling and data analytics, which means attrition models and attendance scores, not only frontier AI.

    You come away knowing what SB 947 and SB 951 demand, how Cal-WARN notices must now name AI-driven displacement, what the Chamber of Commerce and Senator McNerney argued, why the SB 7 inventory clause was dropped, and how to audit every system scoring a named employee.

    Key takeaways

    • Inventory every system that outputs a score, ranking or flag about a named employee, including attendance, quality and attrition models.
    • For each system, record in writing whether you retain the input features and output for that individual on that date, and who the human reviewer was.
    • Stop scoring pipelines overwriting prior outputs, since a decision may be questioned years later.
    • Treat internal automation ROI decks and headcount-avoided trackers as future legal exhibits under SB 951's Cal-WARN disclosures.
    • Check models for proxy inference of protected characteristics such as postcode or shift pattern, which SB 947 prohibits.

    Chapters
    0:00 Why SB 947 is a data team problem
    0:53 What the No Robo Bosses Act requires
    1:52 Proving model lineage for firing decisions
    4:17 SB 951 and Cal-WARN AI layoff notices
    5:46 Chamber opposition and the vetoed SB 7
    7:40 Enforcement odds and your logging checklist

    Full article and transcript: https://www.mba-training.com/blog/sb-947-ai-firing-human-review

    Leaders Insights, mba-training.com

    11 min
  • Churn models in telecom: why uplift beats accuracy

    Why do churn models produce accurate predictions that never move retention numbers? The position here is that most telcos predict churn the week it happens, when the customer has already price-shopped and gone. Telefonica's work came from years of accumulating unglamorous signals such as degraded data speed at a single cell tower, payment timing drift and the gap between plan and actual usage, combined with contract stage and competitor pricing by postcode.

    You walk away able to separate accuracy from uplift, size the movable middle, defend a holdout control group to a revenue VP, and wire a score into a pre-approved retention offer the agent can act on. References include MIT Sloan Management Review on unactionable analytics and dbt Labs on data cleaning time.

    Key takeaways

    • Stop scoring churn in the week it happens and build features from months of early signals like payment timing drift and network quality.
    • Treat network quality as a retention input, not only an engineering metric, and get those two teams talking.
    • Measure uplift against a holdout control group instead of accuracy, so you spend only on the movable middle.
    • Wire the churn score into a pre-approved offer the frontline agent can give with no approval chain.
    • Audit your last retention campaign for a control group; without one you have a discount program, not a churn program.

    Chapters
    0:00 Catching churn twelve hours too late
    1:00 Boring signals that predict churn early
    1:54 Why years went into data plumbing
    2:31 Accurate scores nobody can act on
    3:10 Uplift and the movable middle
    4:04 Monday check: did you run a control group

    Go deeper, free lessons

    • Decoding the telecom data goldmine: CDRs, network telemetry, and usage signals
    • Advanced analytics: CLV, churn prediction & demand forecasting
    • Models in production: drift, monitoring & MLOps
    • Governing rich telecom data under GDPR, ePrivacy, and lawful intercept
    • Benchmarking network and service KPIs that leadership tracks

    Full article and transcript: https://www.mba-training.com/blog/telefonica-churn-prediction-subscriber-behavior

    Leaders Insights, mba-training.com

    5 min
  • AI slop detectors: why filtering training data backfires

    Should you strip suspected AI-generated text out of a training set? The position here is no, not blind. Slop detectors are classifiers that flag one to five percent of clean human writing as fake, and they misfire on dense expert prose and standardized legal boilerplate, so teams in financial services and legal have deleted the backbone of their corpus and watched accuracy fall. A recent experiment found filtering degraded performance more than the contamination removed.

    You leave with a Monday plan: quarantine and hand-sample flagged items rather than deleting, tag provenance at ingest instead of at training time, and read the unknown-provenance estimates from dbt Labs and MIT Sloan Management Review with the vendor incentive in mind.

    Key takeaways

    • Run any slop detector on a batch of known-good human data first, and if it flags more than a few percent as fake, the detector is the problem.
    • Quarantine flagged records instead of deleting them, then hand-review a sample of about a hundred to see whether the tool is catching slop or catching your experts.
    • Stamp every source at the point of ingest with where it came from and whether a human or a machine produced it, rather than reconstructing provenance at training time.
    • Treat filtering as a governance decision about unknown-provenance data, not a question of which detector to buy.
    • Discount the dbt Labs figure of over sixty percent unknown-provenance data as a vendor number and cross-check it against independent work such as MIT Sloan Management Review.

    Chapters
    0:00 Why a filtered dataset made accuracy drop
    0:50 How often detectors flag human writing
    1:36 Contract summarization corpus gutted by filtering
    2:26 Data provenance as a governance problem
    3:28 Quarantine flagged data instead of deleting
    4:00 Test the detector on known-good data

    Go deeper, free lessons

    • Data quality dimensions: why 'good enough' destroys trust
    • Shift-left data quality: embedding governance in the engineering pipeline
    • Data lineage & metadata management: knowing where your data was born
    • Data observability: detect problems before your users
    • LLMOps & evaluation

    Full article and transcript: https://www.mba-training.com/blog/ai-slop-training-data-quality-cdo

    Leaders Insights, mba-training.com

    5 min
  • Cross-agency data: why 70% of failures are semantic

    When two government agencies define "household" differently, one by mailing address and one by shared cooking facilities, no integration platform reconciles the numbers. This episode argues that roughly 70 percent of cross-agency integration failures trace to the semantic layer rather than the plumbing, and that the median time to reach a trusted cross-agency data agreement is 18 months.

    You get a worked example from benefits coordination, where food assistance uses an economic household and Medicaid often uses the tax household, plus the trap of mapping household_id to household_id. You leave with a sequencing rule for governance spend, a named-owner test, and a three-entity definition exercise. References include MIT Sloan Management Review and dbt Labs.

    Key takeaways

    • Spend the first dollar on written definitions, not an integration platform, because mapping fields that mean different things creates agreement without substance.
    • Treat any vendor promise of interoperability in a quarter as selling the connection, not the consensus; plan against an 18-month median for a real cross-agency agreement.
    • Ship one agreed entity, such as household, case or resident, with both agencies signing the exact wording including edge cases, then publish it.
    • Name a single person who can say no on each definition instead of a working group, since shared ownership means the definition drifts.
    • Ask each partner agency to define your three most-shared entities independently in one sentence, and treat every disagreement as the actual project.

    Chapters
    0:00 Two agencies, two definitions of household
    0:58 Why 70% of integration failures are semantic
    2:03 Food assistance and Medicaid household mismatch
    3:02 The 18-month median for data agreements
    3:40 Shipping a signed definition, not a dashboard
    4:18 Monday action: write definitions independently

    Go deeper, free lessons

    • Interoperability and shared standards across agencies
    • Mapping the public data landscape: registries, admin records, and survey data
    • Governing privacy and equity on legacy systems
    • Writing consent and data-sharing agreements that survive an audit
    • Building open data that citizens and journalists actually use

    Full article and transcript: https://www.mba-training.com/blog/public-sector-interoperability-shared-data-infrastructure

    Leaders Insights, mba-training.com

    6 min
  • Amazon's data flywheel: twelve years of engineering

    Does more data automatically compound into advantage? This episode argues it does not. Drawing on an MIT Sloan Management Review breakdown tracing Amazon's flywheel to its origins, the discussion shows the compounding came from roughly twelve years of feedback loop design, architecture and organizational choices, including teams owning data as a product with an internal consumer they answered to.

    You come away with a test for your own setup: time to feedback, meaning how long from a customer action until that signal changes what the next customer sees. You also get a check on modern data stack claims from tooling vendors such as dbt Labs, and a working definition of data monetization measured on the P&L.

    Key takeaways

    • Measure time to feedback: how long from a customer action until that signal changes the next customer's experience.
    • Before buying another dataset, trace one repeated decision such as pricing or restocking and confirm existing data actually changes it.
    • Treat transformation tooling as necessary but not sufficient, and cross-check vendor speed claims against how the business acts on output.
    • Define data monetization as internal decisions that lift the P&L, such as inventory placement and fewer returns, rather than selling data externally.
    • Assign ownership: make teams accountable for their data as a product with a named internal consumer.

    Chapters
    0:00 Why the data flywheel needs building
    1:14 More data does not improve decisions
    1:45 Modern data stack is only half
    2:53 Data monetization is internal, not external
    3:26 Time to feedback as the metric
    4:05 Wire one closed loop first

    Go deeper, free lessons

    • Data monetization: three modes & the data flywheel
    • CDO in retail & e-commerce: the data flywheel
    • Recommendation systems: architectures & ethical personalization
    • The data-as-a-product mindset
    • Data productization: pricing, distribution & business case

    Full article and transcript: https://www.mba-training.com/blog/amazon-data-flywheel-compounding-advantage

    Leaders Insights, mba-training.com

    5 min
  • Boeing component traceability: pay for the birth record

    When a 737 MAX fastener is installed without a birth record, who owns the defect? The answer is the assembler, not the tier-3 shop that made the part. This episode takes apart Boeing's multi-year traceability program, the digital genealogy record that follows one component from raw metal to installed part, and why three-quarters of the cost sits in supplier process rather than software. DBT Labs' claimed 30% reconciliation tax is weighed against MIT Sloan Management Review's manufacturing data work.

    You leave knowing how to write supplier clauses that pay for clean records and penalize broken ones, how to scope traceability to your five most recall-critical parts, and why orphaned part numbers break chains of custody.

    Key takeaways

    • Create the birth record as a single object the moment the part exists, not as a report generated later.
    • Rewrite supplier contracts so the tier-3 shop is paid for a clean record and penalized for a bad one, before buying any platform.
    • Scope traceability to the five parts that would sink you in a recall and trace those end to end.
    • Treat vendor figures like DBT Labs' 30% reconciliation tax as a sales claim until independent work such as MIT Sloan confirms it.
    • Expect most of the cost in supplier process and incentives, not in the database or tooling.

    Chapters
    0:00 Why a 737 fastener lacks a birth record
    0:48 Digital genealogy across 12,000 suppliers
    1:23 Orphaned records and the reconciliation tax
    2:22 Why buying a traceability platform fails
    3:10 What a mid-sized parts maker should copy
    3:57 Cost of building traceability versus ignoring it

    Go deeper, free lessons

    • Building end-to-end traceability across the supply chain
    • Master data done right: materials, BOMs, and equipment hierarchies
    • Product liability and safety recalls: who pays when a product hurts someone
    • Governing data shared with suppliers, customers, and machine OEMs
    • Integrating MES and ERP for a unified data backbone

    Full article and transcript: https://www.mba-training.com/blog/discrete-manufacturing-supply-chain-traceability

    Leaders Insights, mba-training.com

    5 min
  • Semantic layers: one revenue number for the board

    What happens when finance, sales and product each compute revenue their own way? The episode opens on a board meeting where three presenters defend three different quarterly numbers, and argues that the cost is not bad data but frozen decisions while teams re-litigate whose spreadsheet counts. Finance recognizes revenue on fulfillment, sales on signature, product on ship date, and each team's bonus depends on its version.

    You come away with a build sequence for a semantic layer: inventory the 12 to 15 metrics that appear in board decks and budget fights, lock each to one written formula with finance and sales signing off together, then choose tooling. The discussion weighs a dbt Labs vendor claim against MIT Sloan Management Review's work on trust in shared numbers, and covers churn, customer acquisition cost, and the CDO's refereeing role.

    Key takeaways

    • Inventory only the 12 to 15 metrics that show up in board decks and budget fights, not all 400.
    • Lock each metric to a single written formula and make finance and sales leads sign off in the same room.
    • Define the five metrics in your next board meeting, ship those, and use the momentum for the rest.
    • Treat the dbt Labs figure on conflicting metric definitions as vendor marketing, not evidence.
    • Email each department head before the next board meeting asking for their exact formula in one sentence, then compare the replies.

    Chapters
    0:00 Three revenue numbers in one board deck
    0:47 Why finance, sales and product disagree
    1:24 What a semantic layer actually is
    2:03 Inventory metrics and lock the formulas
    2:40 dbt Labs claims and the trust problem
    3:20 Shipping five metrics before committee kills it

    Go deeper, free lessons

    • The metrics & semantic layer
    • Modern BI: tools, maturity & the semantic layer
    • dbt (data build tool): industrialized SQL transformation
    • Designing a KPI tree
    • Data contracts: the new standard for quality agreements between teams

    Full article and transcript: https://www.mba-training.com/blog/semantic-layer-metric-consistency-playbook

    Leaders Insights, mba-training.com

    5 min
  • Identity resolution in pharma: one HCP, six records

    Why does a single cardiologist show up as six separate entities across a pharma company's CRM, purchased claims data and analytics warehouse, and what does that cost? The position here is blunt: identity resolution is not a tool purchase, it is plumbing. Bad matching splits one high prescriber into six low ones, which skews sample allocation, wastes temperature controlled product, and can hide adverse event patterns in pharmacovigilance reporting that regulators will not excuse.

    You get a working approach: make the federal NPI registry the source of truth, run nightly matches, write down survivorship rules, and name one accountable owner with budget for the prescriber master record. References include DBT Labs and MIT Sloan Management Review, plus a master data example at Novartis.

    Key takeaways

    • Make the federal NPI registry the source of truth for prescriber identity and force every other system to point at it.
    • Run a match against the federal file nightly instead of treating deduplication as a one-off project that decays within 18 months.
    • Write down explicit survivorship rules so you know which record wins when two versions disagree on address or name.
    • Name one person accountable for the prescriber master record, with budget, before buying any matching tool.
    • Pull every record for your top 10 prescribers by revenue by hand and count the duplicates to get the dollar figure your board needs.

    Chapters
    0:00 Six records for one cardiologist
    1:09 Sampling waste from split prescriber records
    2:02 Pharmacovigilance risk and regulator warning letters
    2:53 Identity as plumbing: NPI golden record
    3:51 Naming an owner for prescriber master data
    4:12 Monday morning audit of top 10 prescribers

    Go deeper, free lessons

    • Master Data Management in practice: styles, tools, and the Golden Record
    • Data quality metrics that matter: completeness, latency and lineage
    • Mapping the pharma data landscape: sources, vendors and standards
    • Consent, de-identification and the limits of anonymous data
    • Marketing on a leash: promotional compliance and the anti-kickback minefield

    Full article and transcript: https://www.mba-training.com/blog/hcp-patient-identity-resolution-pharma

    Leaders Insights, mba-training.com

    5 min
  • AI agents querying your warehouse: measure before rebuilding

    Fivetran and dbt Labs used dbt Summit to argue that AI agents, not human analysts, are now the main consumer of enterprise data. The position here is that the architectural shift is real but the urgency is a sales pitch, and that re-platforming off a keynote is the expensive mistake.

    You come away able to test dbt Labs' claim that 30% of warehouse queries are machine-generated against your own logs, spot the failure mode where agents pass nonsense results downstream with no scar tissue, make the case for a written semantic layer over new tooling, and set a query budget before an agent touches production. References include O'Reilly Radar and MIT Sloan Management Review.

    Key takeaways

    • Pull last month's query logs and calculate what share came from service accounts rather than named humans to find your real agent load.
    • Treat dbt Labs' 30% machine-generated query figure as directional, since the vendor sells the tooling that raises it.
    • If agent traffic is small, spend the quarter writing business definitions down instead of changing architecture.
    • Build the semantic layer so terms like active customer are defined in writing, because an agent has no memory to fill the gaps.
    • Set a spending cap on agent query volume before a pilot touches production data to avoid a five-figure compute surprise.

    Chapters
    0:00 Agents as the new data consumer
    0:44 What changes when a bot queries
    1:12 dbt's 30% machine-generated query claim
    2:03 Metadata gaps and the semantic layer
    3:04 Compute costs and query budgets
    3:49 Check service account query share

    Go deeper, free lessons

    • dbt (data build tool): industrialized SQL transformation
    • Active metadata and the data catalog
    • The metrics & semantic layer
    • Data lineage & metadata management: knowing where your data was born
    • Data products: definition, design & lifecycle management

    Full article and transcript: https://www.mba-training.com/blog/agent-ready-data-infrastructure-dbt

    Leaders Insights, mba-training.com

    5 min
  • AML pipelines: three design calls examiners test

    Why do fintech AML pipelines fail regulatory examination when the models themselves are sound? The position here is that examiners do not score area under the curve, they ask reproducibility questions, and most data architectures cannot rebuild a March alert with March data. Three design decisions decide the outcome: immutable timestamped feature storage, documentation that records who approved a feature and which regulation it maps to, and explanation captured at inference time rather than manufactured later.

    You walk away knowing why dbt column level lineage maps transformations but not business intent, why SHAP values may be plausible rather than faithful, and how to run a one hour reproduction test on your last high value alert.

    Key takeaways

    • Freeze every model input as it existed on the decision date so an alert can be rerun with that day's data.
    • Keep design records for each feature covering who built it, when, why, and which regulation it maps to.
    • Log contributing features live at inference time instead of reconstructing a reason with post hoc explainers.
    • Treat vendor lineage statistics, such as dbt Labs' 80% column level lineage claim, as marketing until checked against your own graph.
    • Take your last high value AML alert and try to reproduce inputs, output and reason in under an hour.

    Chapters
    0:00 The March alert nobody could explain
    1:08 Why model accuracy does not pass exams
    2:03 Lineage tooling cannot show business intent
    3:18 Post hoc explainers versus inference-time logging
    4:16 Retrofit costs and a Monday reproduction test

    Go deeper, free lessons

    • AML, KYC and sanctions screening in practice
    • Reading the transaction ledger: what payment and behavioral data reveal
    • Governing fintech data: consent, lineage, and regulatory defensibility
    • the regulatory map every fintech data leader must carry
    • The regulatory bodies map: who enforces what and how examinations work

    Full article and transcript: https://www.mba-training.com/blog/fraud-kyc-aml-detection-pipeline-fintech

    Leaders Insights, mba-training.com

    6 min

About Leaders Insights — Data

From the publisher's feed

Leaders Insights — Data. Daily strategy in data governance, architecture, analytics and AI, for data leaders and aspiring CDOs. New episode every day at mba-training.com.