
Sign up to save your podcasts
Or


In this episode of The Deep Edge Podcast, host Ray Mota sits down with Ari Weil, Vice President of Cloud Computing and Delivery Product Marketing at Akamai, to examine the real economics of AI inference.
Token pricing may be the most visible line item, but it does not tell the whole story. Ray and Ari discuss how network egress, cross-region data transfers, latency, model placement, and increasingly complex agentic and RAG workflows can reshape the total cost of running AI in production.
They also explore when centralized inference makes sense, when distributed and edge architectures offer an advantage, and the essential questions enterprises should ask before committing to an AI infrastructure provider.
Read Ari Weil’s article, “Your AI Cost Model Stops at the Token Price. The Bill Doesn’t”:
https://www.akamai.com/blog/ai/ai-cost-model-stops-token-price-bill-doesnt
Topics covered:
• Why token pricing provides an incomplete view of AI costs
• The hidden impact of egress and data movement
• How agentic AI and RAG increase infrastructure interactions
• Latency, locality, privacy, and data sovereignty
• Centralized versus distributed AI inference
• Choosing the right model for each workload
• Questions enterprises should ask AI infrastructure vendors
Chapters:
00:00 Introduction
01:25 Welcome — Ray Mota and Ari Weil
02:04 Why token pricing misses the real cost
03:53 Agentic AI, RAG, and infrastructure interactions
06:18 Why egress, data movement, and latency matter
08:36 Centralized versus distributed inference
10:26 Questions enterprises should ask vendors
12:01 Closing thoughts
Subscribe to The Deep Edge Podcast for conversations about AI infrastructure, cloud computing, networking, edge architecture, and the technologies shaping our connected world.
#AI #AIInfrastructure #CloudComputing #EdgeComputing #Akamai #AgenticAI #TheDeepEdgePodcast
What if a message could be protected by the laws of physics—not just by encryption?
In this video, we explore the real promise of quantum networking and the science behind the no-cloning theorem, quantum entanglement, trusted nodes, quantum repeaters, and the race to build a true end-to-end quantum internet.
We also look at what is working today, including China’s large-scale quantum communication network, and what is still holding the industry back. The biggest challenges are not the theory, they are the engineering bottlenecks involving distance, quantum memory, hardware interoperability, switching, and network emulation.
You will learn:
• Why quantum data cannot be copied
• How eavesdropping leaves a detectable fingerprint
• Why trusted relay stations remain a security risk
• How entanglement swapping and quantum repeaters work
• Why quantum memory is one of the biggest technical obstacles
• How digital twins and virtual emulators are helping engineers test quantum networks
• The three-stage roadmap toward a global quantum internet
• Why “harvest now, decrypt later” is accelerating investment in quantum security
Quantum networking is not simply a faster version of today’s internet. It represents an entirely new approach to communication, security, and distributed computing.
For the last century, we built the internet around the ability to perfectly copy information. The engineering race of the next century may be defined by the ability not to.
#QuantumNetworking #QuantumInternet #QuantumComputing #QuantumSecurity #Cybersecurity #QuantumCommunication #PostQuantumCryptography #Technology #FutureOfNetworking #ACGResearch
For the last two years, the AI infrastructure conversation has been dominated by GPUs. But as AI clusters scale to tens of thousands — and eventually hundreds of thousands — of GPUs, the real bottleneck is shifting.
The next AI infrastructure challenge is not just compute.
It is interconnect.
In this episode of The Deep Edge Podcast, Ray Mota breaks down why optical networking is becoming the foundation of AI data centers. From 800G and 1.6T optics to Linear Pluggable Optics, Co-Packaged Optics, Optical Circuit Switching, and silicon photonics, the future of AI depends on how fast, efficiently, and reliably GPUs can communicate.
A GPU waiting on data is stranded capital.
And optical is what keeps the AI factory moving.
Topics covered:
Key takeaway:
Compute gets the headlines, but optical holds the AI factory together.
Everyone is talking about the GPU shortage. GPUs are expensive, scarce, power-hungry, and central to AI infrastructure. But the real issue is much deeper.
In this episode of The Deep Edge Podcast, Ray Mota breaks down why AI infrastructure is not fundamentally a GPU problem — it is a network economics problem. As AI workloads scale, the hidden bottleneck is not just compute capacity, but the ability of the network to move data efficiently, keep GPUs utilized, reduce latency, control power consumption, and avoid massive infrastructure waste
Ray explains why enterprises, service providers, and hyperscalers must rethink AI infrastructure through the lens of network design, utilization economics, east-west traffic, data movement, automation, and total cost of ownership.
The key message: buying more GPUs will not solve the problem if the network cannot support the economics of AI at scale.
Starbucks has recently pulled the plug on their AI inventory tool, after just 9 months in 11,000+ shops.
The bigger lesson for me is simple:
AI cannot solve an operations problem it does not understand.
The technology was designed to utilize computer vision and spatial intelligence to count inventories faster and more precisely. On paper it sounds powerful. But in the real world, it apparently confused products, failed to find items on shelves, and made stockout concerns worse, not better.
This is not a Starbucks problem. It’s a larger lesson for enterprise AI.
Many firms are trying to expand AI without even understanding the operational challenge, the workflow, the human behavior and execution environment. AI doesn’t instantly produce process discipline. It magnifies the quality, or the fragility, of the underlying operational model.
Some major take-aways:
Speed is not preparation. Scaling AI to thousands of locations fast might seem daring, but success in the pilot does not always translate to success in production.
AI requires operational context. If the fundamental cause is supply chain consistency, replenishment timeliness, store level execution, or process variation, then counting faster is not the issue.
Frontline feedback is important. The failure sites are commonly seen by employees closest to the work. That signal is ignored, delaying adjustment and increasing risk.
The design of the integration is as crucial as the AI model. Enterprise AI success isn’t about claims of accuracy. It’s about workflow fit, process reform, data quality, governance, training and adoption.
Enterprise AI’s biggest danger isn’t the adoption. Scaling prematurely without properly diagnosing the business and operational problem.
AI is powerful, but we need to relate it to the realities of how work really gets done.”
What’s your take: are corporations rushing too fast to implement AI before they properly grasp the operational difficulties they are seeking to solve?
#AI #EnterpriseAI #DigitalTransformation #RetailTech #SupplyChain #Operations #TechLeadership
In this episode of The Deep Edge Podcast, Ray Mota breaks down one of the most important questions facing network architects, CTOs, and business leaders building next-generation AI data centers: how to make the right infrastructure investment decisions with real confidence.
Ray explains why the economics of AI data centers are fundamentally different from traditional enterprise networks. In the AI era, the biggest investment is often the GPU cluster, and the network’s role is no longer just to provide connectivity. Its job is to keep high-value GPU resources fully utilized. When the network underperforms, the result is not just lower performance. It can mean lost revenue, delayed service delivery, and direct destruction of capital efficiency.
The episode explores why financial modeling tools such as economic digital twins are becoming essential for planning AI infrastructure. Ray also outlines the critical questions every organization should quantify before committing capital, including the cost of GPU idle time, the CapEx-to-OpEx tradeoffs of different architectures, and the break-even point for transport decisions under varying growth scenarios. This is a practical and executive-level discussion on why AI infrastructure decisions must now be guided by modeled economics, not vendor slides.
In this episode of The Deep Edge Podcast, host Ray Mota, Ph.D. speaks with Will Eatherton, Cisco’s engineering leader driving the future of AI-optimized networking.
They dive deep into:
• Cisco’s AI-ready data center strategy
• Silicon One evolution & optical innovation
• Adaptive routing, congestion control & GPU fabric design
• Cisco + NVIDIA partnership and what it means for AI clusters
• SONiC openness, multi-cloud integration & operational simplification
• HyperFabric, LPO/CPO optics, and the shift from training to inference at scale
Video link: https://youtu.be/39X2-BIL9sE
WIll's Bio:
As the SVP of the Data Center, Internet & Cloud Infrastructure Engineering team, Will leads a wide-ranging group responsible for data center, web, and service provider solutions. This group encompasses cutting-edge technologies, longstanding Cisco product lines, and the recently launched Cisco 8000 series.
Will initially joined Cisco in 2000 with the acquisition of Growth Networks. During his early years at Cisco, Will expanded his expertise in routing systems architecture and co-authored several patents on packet processing. In 2013, Will co-founded Skyport Systems, a cloud technology pioneer addressing industry gaps in cloud-managed secure virtualized systems. In 2018, Will returned to Cisco with the acquisition of Skyport Systems. With his return, Will has led the Cisco Networking Engineering team, bringing with him a wealth of industry knowledge and experience.
Will is an active technology researcher, author, and speaker. His most recent speaking engagements include sessions on high-performance ethernet fabrics for AI infrastructure and enabling enterprise generative AI with ethernet AI networking.
Will has a Master’s degree in Electrical Engineering from Washington University in St. Louis, MO, and a Bachelor’s in Electric Engineering from the Missouri University of Science and Technology.
I use AI to create this podcast, let me know what you think?
The ACG Research study highlights how application-driven traffic is reshaping the economics of fixed and mobile networks. While fixed networks benefit from lower unit costs ($0.06/GB vs. $0.33/GB for mobile), far higher household data consumption drives a 393% higher TCO per subscriber ($21.22 vs. $5.40). A small set of applications accounts for the vast majority of costs: 9 apps drive 92% of fixed expenses and seven drive 96% of mobile led by YouTube, Netflix, TikTok, Facebook, and Steam. Cost burdens are concentrated in the access layer for fixed networks (75% of TCO) and in the RAN for mobile (85%), underscoring structural vulnerabilities as video, gaming, and latency-sensitive applications proliferate.
For CSPs, these dynamics create rising margin pressure, which are exacerbated by OTT players that capture revenue while operators absorb infrastructure costs. To respond, CSPs must adopt application-aware traffic management to optimize QoE and defer CapEx, pursue commercial or regulatory frameworks to rebalance costs with OTT providers, and deploy CSP specific TCO models to guide investment and policy strategies. Sustaining profitability in an application-centric era will depend on combining smarter traffic engineering, fairer cost sharing, and tailored economic modeling.
Download report:
https://www.acgcc.com/reports/the-economics-of-application-traffic-tco-benchmark/
For more information, contact Ray Mota or Peter Fetterolf.
Join Ray Mota, CEO of ACG Research and host of the Deep Edge Podcast, for an in-depth conversation with Kevin Vachon, of Mplify (formerly the MEF). In this episode, Kevin walks us through:
• The Story Behind the Rebrand: Why “Mplify” was chosen, how market shifts toward multi-cloud, AI, and automation drove the decision, and the legacy of LSO and network standardization.
• Building a Global NaaS Ecosystem: Strategies for fostering collaboration among service providers, vendors, hyperscalers, data-center operators, and enterprise customers—plus the engagement models that make it all work.
• Certifications & Interoperability: How trusted transparency and standardized APIs reduce risk, boost confidence, and drive real interoperability across layer-2, APIs, and beyond.
• Networks for AI vs. AI for Networks: Clarifying where Mplify plays in enabling AI-centric use cases, from on-demand GPU services to elastic bandwidth and consumption-based models.
• Inside the Global NaaS Event (GNE) 2025: What to expect at this year’s Dallas gathering—real demos, rapid-fire user stories, business-value case studies, and deep dives into next-generation NaaS capabilities.
LinkedIn:
https://www.linkedin.com/in/kevin-vachon-5675ab1/
Kevin's bio
https://www.mplify.net/people/kevin-vachon/
Mplify Board Link
https://www.mplify.net/about-mplify/officers-board-of-directors/
#mplify #acgresearch #NaaS #LSO
Join host Ray Mota as he sits down with Matt Biringer, Co-Founder of North Cloud, to explore the evolving world of public cloud finance. In this episode, we dive into:
• Why cloud cost management matters – The shift from upfront data-center purchasing to utility-based billing and its impact on budgets.
• The rise of FinOps – How finance, engineering, and product teams can collaborate around cloud spend—and the cultural and political barriers they face.
• Visibility & automation – From 50-page invoices to real-time dashboards, and the AI-driven tools that right-size resources, alert on anomalies, and allocate costs automatically.
• Agent North: Cloud FinOps AI – A sneak peek at North Cloud’s upcoming “Agent North,” which ingests your billing telemetry and technical docs to answer finance and engineering questions in natural language.
• Security & governance – Balancing automation with safeguards, air-gaps, and scenario planning to keep your financial automation platform secure.
About Matt Biringer
Matt Biringer is the co-founder and CEO of North.Cloud, the AI-powered platform that helps companies cut cloud waste and control costs, without the spreadsheets. Before launching North in 2023 (literally from his garage), Matt spent 12 years selling datacenter technology, leading growth at Pure Storage, CDI, and SHI.
When he’s not scaling North, you’ll find him chasing his kids, training jiu-jitsu, lifting weights, or perfecting a new recipe. He also angel invests in early-stage startups, backing companies like NodeECO, Light, Data Herald, and Pipedream Labs.
Find Matt online:
LinkedIn
https://www.linkedin.com/in/mattbiringer/
Twitter
https://x.com/biringer_
About North.Cloud
North.Cloud is an AI-powered cloud optimization platform tackling the rising costs and inefficiencies of modern cloud infrastructure. By eliminating manual FinOps processes and rigid cost models, North delivers real-time, automated savings across AWS and GCP. Its dynamic platform helps companies reduce waste, improve efficiency, and gain transparency over their cloud spend. To learn more,
visit North.Cloud
https://www.north.cloud/
From the publisher's feed
The Deep Edge Podcast discusses technologies powering a cloud-native, fully connected world: Virtual and physical compute; virtual and physical networking; orchestration; mobile/wireless, wired;…