
Sign up to save your podcasts
Or


In this podcast we will:
- briefly explain what happed with Huggingface from a technical perspective against the current AI extinction narrative
- explain how this story hit the media airwaves, went viral and scared everyone who read the headlines
- introduce the scheme used to accomplish this to spread the disinformation campaign with a repetition narrative
- provide guidance for journalists to be mindful of these tactics and how to responsibly report without adding fuel to the fire.
Some important references that we've provided on the show:
Forensic and Technical Analysis:
Abi Awomosu, The Machine Alibi: Unpacking the Myth of Superintelligence
Galaxy Brain, An Argument Against AI Doom, an interview with cybersecurity expert Zack Korman
House of El: AI YouTube video: Like Sabine Hossenfelder, I was offered Money to Tell You AI Will Kill Us
Cal Newport - The Truth of Open AI's Secret AI Civilization
Margaret Mitchell - Stochastic Parrots and the LLM Quandary
On disinformation, perception management / reflexive control and cult dimensions in the AI Doomsday / Hype narratives
How Big Tobacco Set the Stage for Fake News
Alondra Nelson - Algorithmic Agnotology: On AI, Ignorance and Power
What Tech Oligarachs Gain from AI Doomerism
Abigail Dubiniecki, "Anatomy of a Disinformation Op" in Disinformation is a Mind-Hack: Privacy Pros Must Protect Decisional Privacy
House of El:AI, YouTube video: AI CEOs Are Broke, Desperate and Lying About Doomsday
Sabine Hossenfelder, YouTube video: I was Offered Money to Tell You AI Will Kill Us.
Beyond the AI Hype by WAIE+, Altman, Amodei(s), and Musk want to slow down AI... again
Jim Stewartson (MindWar), Why Trump Wants "Super Intelligence" : Peter Thiel's Sci-Fi Alternate Universe
Ian K Duncan, Sex, AI and the Apocalypse
SNL sketch mocking the conflicting messages
For Journalists:
Global Investigative Journalism Network, AI Accountability Reporting Guide
Women in AI Ethics Newsletter, The AI Ethics essential reading list
WAIE+, The List - Hall of Fame
General Purpose Systems, Large Language Models (LLMs) and Agents are now the most debated topics today. They are touted as enabling significant progress since LLMs were introduced three years ago.
The irony of progress has meant hiccups—for these systems and the capabilities they produce, these hiccups are life threatening, economically, mentally, and personally. The hype peddled by frontier companies are distractions from these very real consequences. The power that comes from a handful of Big Tech companies, backed by the U.S. administration, shades them from any accountability.
The misdeeds of these systems are well documented from teen suicide, AI psychosis, addiction, misinformation, copyright infringement, the data center energy guzzling, and more recently, serious infrastructure security risks from agent attacks. The latest OpenAI agent unauthorized attack of Hugging Face has created a media firestorm that threatens existential threat to mankind and an AI that needs to be controlled.
Before that, the markets were convinced of extraordinary gains, and have been very optimistic despite the bubble talks, and circular financing. Now investors face a dual risk:
…the prospect that AI could pose an existential threat to humanity, and a possible slowdown championed by some of the industry's biggest names.
We’ve seen this scene play out before: In 2023 Future of Life’s Open Letter to pause AI development was signed by the who’s who of Silicon Valley, who called for regulatory oversight and development of governance standards. That didn’t happen.
The warnings that Margaret Mitchell, Timnit Gebru, Emily Bender and Angelina McMillan-Major detailed in their paper, “The Dangers of Stochastic Parrots: Can Language Models be too Big?” did not amount to effective legislative and AI governance progress to mitigate the risks they outlined until AFTER these harms had already played out.
Our discussion with Margaret Mitchell, a computer scientist and one of the earliest leaders in AI ethics, demystifies LLMs and their implications.
This event was done in collaboration with The AI Fellowship Toronto.
System Malfunction is a reader-supported publication. All my posts are currently free. My intention is to share my knowledge without paywalls. To receive new posts and support my work, consider becoming a subscriber.
Margaret Mitchell is a researcher and Chief Ethics Scientist at Hugging Face. She also co-authored, with Emily Bender, Timnit Gebru and Angelina McMillan-Major, “The Dangers of Stochastic Parrots: Can Language Models be too Big?”
This paper came out in 2021, before the mainstream introduction of Open AI’s ChatGPT. Mitchell shares that large language models have been in development over the past five-ten years. In the 1990’s language models, which were smaller, were being developed across all kinds of tasks. And she was part of this language modelling work, as she explains,
I did my thesis on generating language based on visual input, and this was using language models. And so, I was in a position in 2020 where I saw that the growth of large language models introduce a ton of different issues—and potential benefits—that were being overlooked [in the tech industry, people tend to focus on the positives.]
As someone who had been developing language models for a long time, I saw that they were about to take off—and from my specific position in the industry—without any academic writing on harms and risks.
Mitchell and her colleagues decided to document these harms and risks to develop some grounding as they evolved. What they recognized was the massive potential with language models becoming fluent, as she states,
Previously, language models could be used for small chunks of text. Now language models, because they were large, could capture much larger contexts, and then that means that they could generate entire essays that sounded fluent in ways we could never do before.
Once we started seeing that fluency, it was clear to us that people were going to start imbuing it with human-like intelligence.
She added that this new paradigm of LLMs was based on the massive scraping of internet data without consent(which she explains that she contributed to) and would surface legal and societal rights issues in its wake.
At the same time, Mitchell knew that this fluency could be used to generate things that looked factual, and attribute these systems as individual entities, extensions of our human likeness. She adds,
We use that as part of our own expressions of humanness… It was clear to me that people were going to love this, and they were not going to know the full picture about what these things are.
The paper argued that LLMs do not understand the “concepts underlying what they learn." It also cited significant risks when it came to LLM’s carbon emissions and financial costs as models are fed more data. Moreover, the disproportionate data capture that favored the western, more economically rich nations meant that smaller nations with less internet access would have far less representation.
This paper was the catalyst that saw Google executives respond negatively and demand retraction, which eventually led to her team co-lead, Timnit Gebru being forced out of Google.
Scraping Internet at Scale Does Not Equate to A Representation of the World
Before LLMs, smaller language models were developed from curated text data from licensing agreements with trusted institutions like The Wall Street Journal. The LLM paradigm shifted to massive scraping of internet data without consent, as she explains,
So, it went from the paradigm where you decide on the desirable data inputs, with appropriate licensing to a new recognition that with Web 2.0 the growth of blogs and social media, you could get tons of text, without any legal protection.
Going into 2019, we saw companies stop curating altogether, with the prevailing notion to retrieve whatever the Internet provided, as opposed to focusing on high-quality resources.
And so that created a situation where we were just all getting as much data as we possibly could, regardless of any other considerations, because it meant improving downstream benchmark performance and the perceived fluency.
The data being scraped was skewed to represent the viewpoints of people who were on the internet the most. ChatGPT tended to select from the US, Western Europe and parts of East Asia. This paper which introduces “silicon gaze” to explain how LLMs reproduce and amplify inequalities argued
bias is not a correctable anomaly but an intrinsic feature of generative AI, rooted in historically uneven data ecologies and design choices.
The structural features mean that training data has already been shaped by “centuries of uneven information production” giving advantage to English-language and stronger digital presence.
Mitchell adds that data ingested from sites like Wikipedia are populated by North American men, and by some estimates, between the ages of 17 and 35. Also among the top 10 cited domains on LLMs, Reddit, Linkedin, Medium, Youtube, Google, Forbes are largely U.S. centric, with audiences predominant within the Global North.
Mitchell adds that Global North hegemony—disproportionately represented in the data including misogynistic viewpoints from predominantly male users on Reddit, a narrow view of ‘worthy’ content published on Wikipedia, and influential figures across blogs and social networks—are what these systems are learning from, as she adds
Women's viewpoints are generally not well represented and the viewpoints of people who are non-white are generally not well represented.
AI is Not Neutral
In the Stochastic Parrots paper, Ruha Benjamin, Abebe Birane and Vinay Uday Prabhu were cited for their report where they highlighted issues with LLMs and AI systems that scrape massive web datasets like Common Crawl, containing toxicity, and hate speech, and argued that filtering or keyword blacklists are ineffective at identifying harmful content without inadvertently censoring or harming groups,
“Feeding AI systems on the world’s beauty, ugliness and cruelty but expecting it to reflect only the beauty is a fantasy.”
Mitchell explains this encoded bias that has resulted from the development of LLMs,
The thing that happens with slurs, toxic or obscene language is that they are also associated with things you do want to have represented.
A classic case of this are communities where common terms are also used in other communities as epithets.
For example, ‘gay porn’ is disproportionately more commonly found on the internet than ‘straight porn.’ The word, ‘porn’ is marked as obscene language. The word, ‘gay’ may be marked as a slur.
By trying to filter out the word, ‘gay’ risks also disproportionately erasing communities that are already marginalized, such as the LGBTQ.
Mitchell adds that despite the flawed curation that can ingest misogynistic, abusive content, and child sexual abuse (CSAM) material from the web, there is this reflection of tech optimism that doesn’t fully reflect the data. The supposed neutrality of AI just doesn’t bear out. She points out that biases against women and against Black people are far too common,
When you do a search for great fiction authors, it’ll tend to skew to white men. So, it’s these kinds of subtle effects that you don’t notice, but influence your perception of the world.
She explains that at least with traditional search engines, extraction techniques ensured a solid match between the search query and the search result—content from ‘real people.’ With LLMs it’s different:
LLM searches run on probabilistic sequences, meaning that pieces of text are being stitched together that sound confident and factual because they’re trained on confident and factual data, but they’re not grounded—directly connected to a clear source that can verify it’s true.
You can click on those citations and you can tell they’re post hoc because they don’t actually support what the summary is saying.
This post hoc analysis means the citations were specified after “the data is seen,” and were not part of the original query, thus making them highly susceptible to false positive results. Mitchell concludes that the interface design that gives the appearance of factuality but,
unlike single models, these larger systems can pull in all the bells and whistles, additional rules, models and algorithms. Right now, they’re based on LLMs, which have many issues. Your foundation is one of quicksand.
So, using LLMs as neutral knowledge sources risks reinforcing the inequalities these systems will mirror.
AI Fluency and the Illusion of True Human Understanding
The pace at which users are embracing chatbots for mental health is staggering:
* A study by CognitiveX found that one in three people use AI chatbots primarily for “fear of judgement or social stigma.”
* 43.75% of people prefer AI chatbots to discuss mental health issues first rather than approaching a trusted person.
Mitchell addresses sycophancy and the Eliza Effect (1966) where individuals attribute a human-like effect from chatbots:
Eliza was a simple rule-based chatbot, and people felt they could really talk to it like a therapist. It was designed to put forward sentences that sounded like Rogerian Therapy and so people’s tendency to impute a human mind when they see human-like language means they do feel like they’re talking to someone they can trust.
These conversations are increasingly private and when people disclose things to chatbots, now their experience is like having a private conversation with something, as Mitchell says,
… that is picking up their language because it’s stitching things together probabilistically, repeating back that language so it creates entrainment [synchronizing their rhythms or flows] in psycholinguistics. And that builds relationships.
These private feelings and confidential settings are not private because they’re going through large company APIs determining how language is aligning to them. Mitchell and her co-authors addressed this in their paper because they were concerned people would fall for this illusion,
As humans have evolved to communicate with other humans via language. I think when people worry about the illusion of human likeness, they may be misunderstanding that this isn’t something you can reason yourself out of.
It’s something that our brains are designed to do. Our brains are designed to have a sense of this other person there and to impute into that intentionality, feelings, emotions and all sorts of things we experience as humans.
We wanted to flag this was going to happen and there was no way out of it unless people took this seriously in their designs and handled this very foreseeable issue.
This did not happen and has led to the current day, AI Psychosis.
Read more here about one company tackling AI Psychosis from chatbots.
Semantic Associations and Fluency Create Coherence
What does the word ‘understanding’ mean from a technical perspective? Mitchell points out we use words like reason, chain of thought—things we understand from our human experience and attribute to something that’s fundamentally not human. She illustrates,
It understands in a neural network computer way.
It does not understand in a human way.
And that’s where people are getting a little bit confused.
Does understanding mean that it can take something you say, expand it into a bunch of related concepts and give you something back that is relevant to what you said? If so, then yes, these systems can do that.
But if understanding means that it can feel what you’re saying—that it can empathize and can match your experience to its own human lived experience—it can’t do that because it’s not human.
The language we’re using is creating more confusion, in our ability to understand what these systems are because we are already attributing human-like tendencies. On top of that, the more we anthropomorphize the chatbot technical behaviors, the more we lose the ability to tell the difference.
Each word, Mitchell explains, has many other natural associations. For example, cat is associated with dog, but also associated with fur. Language models have been exposed to the word, ‘cat’, often co-occurring with the word, ‘dog,” and that’s recorded probabilistically. She continues,
So if you input the word, ‘cat’ the network associates to dog, to pet, to fur, to tiger—this whole network of semantically related concepts. Once it’s connected to this larger network, it can stitch together other phrases related to those concepts that can further extend.
So, if I say, ‘I love cats’ it can say, ‘dogs are nice.’
It has associated ‘love’ to ‘nice’ as an associated related concept and it’s been trained on syntactically well-formed sentences so that these higher level semantic concepts can now be put forward in a fluent English text string.
She concludes that current AI systems can respond with hyper relevant things to what we input because they’re trained to do so. Mitchell emphasizes that we don’t give ourselves enough credit when we get fooled by these systems or when believe we should be smarter as she points out,
This is not about being smart. It’s just how our brains have developed and then how the systems are trained. They are trained to sound coherent and to pull together related semantic concepts. Many times, they can say things that are appropriate.
The $300 Billion Data Center Dilemma: Will We Need Them?
The pace of AI progress has not scaled to the same level of expected efficiency. The transformer models combined with neural network architecture continue to emit high levels of CO2. The Stochastic Parrots paper warned that “compute to train the largest deep learning models has increased 300k times in six years, a far higher pace than Moore’s Law.”
Every action in an AI system requires some amount of energy. So many computations can really slow down the system as they get hotter. You need data centers that run cooling systems, to maintain compute without overheating.
Mitchell explains the direct correlation between these growing systems and the demand for data centers,
If I can say the word ‘cat’ and it’s a small network, then it can only output ‘dog.’ However, if it’s a larger network, then it can associate it with sentences about dogs. That requires more computations and more storage and greater access to all these different kinds of stored representations.
And then as it gets larger, now it’s associating entire documents and then doing the processing over all of the sequences relevant to those documents i.e. those strings of text and then formulating further what to return to the user.
The larger and larger it gets, at the point of inference, the more computations it’s doing, the more things it has access to, the more computations on the server side is required, which is the reason for more powerful systems, more and more servers and more data centers.
Many have warned of the parallel of these data center buildouts to the first internet bubble. The astonishing multi-hundred billion dollar data center investments (to be exact, $335 billion dollar market in 2025) that frontier models have promised will return within the coming decade are still far too expensive to work with within current infrastructures.
Mitchell, however, calls out the tremendous work in the development of smaller models that can run locally on device, which does not require data centers. If smaller models endure in the coming years, will we see this data center boom collapse?
Generated AI is Eating Itself
After a chatbot has been deployed, Mitchell confirms that to create more effective systems it is a technically strategic imperative to treat chatbot and user interactions as gold to make the system better. For free accounts, this data is used for training; and for premium accounts users can choose to disable their interactions for use in training.
However, the more that generated content proliferates on the web, the more it creates this homogenizing effect that can lead to model collapse.
She alludes to the introduction of the long tail effect, coined by Chris Anderson, founder of TEDX, who argued the value of promoting niche or less popular products can collectively build better markets compared to popular products, given a larger market distribution. This has been the massive effect of the social web as popularity did not regress to the rich and famous, but instead, to early influencers from social platforms, which created new niche and local market opportunities as these distributed networks evolved.
As Mitchell describes, the risk is the shattering of the long tail:
The various ways which things can be talked about and expressed and the various topics that people talk about get chopped off when general purpose generations occur. And so, if you keep chopping off the long tail, you are creating a situation where you learn less and have less diversity. It’s called a Ouroborus—a snake eating its own tail.
How long will it take before our systems converge to this mean? Will synthetic generations prevail? Mitchell says it is the case that tons of original data on the internet has now been mined, and the content that is being produced is increasingly generated—not novel!
Five Years Later Since the Paper: “You must change consumer demand to put safeguards in place!”
Mitchell is hopeful. Since their paper, awareness has heightened, and companies have been putting in fixes to mitigate the risks. She acknowledged more legislation has been introduced to identify at what age is appropriate to be introduced to these systems, to lessen the effects of over-dependence on companion bots.
On the technical side, post-training techniques are being used to ‘reshape’ models as she explains
LLMs out of the box have a ton of different biases and are problematic in ways of handling the world that are easily exposed. Post training techniques like Reinforcement Learning with Human Feedback (RLHF) and Constitutional AI are used to move the biases of the system to be associated with the values you want it to represent.
Models can be jail-broken, and post-hoc fixes can be applied that allow developers to move the model into post-training spaces for remediation.
She adds that input prompts and generated outputs can also be put through classifiers to output better responses.
Overall, is there enough being done? Mitchell responds,
I really want to say yes but I don’t know if I honestly can.
It’s amazing that safety is a value that people are talking about. 10 years ago, the idea of ethics in AI was laughable. Over these 10 years people have not only realized that ethics can play a role, but they’ve honed in on critical ethical values such as safety and fairness and made those priorities in the design.
But she also acknowledges the most powerful systems—the most ubiquitous—are commercial and will always have a profit motive that flies in the face of safety precautions. Companies that constrain their systems will lose customers to competitors who don’t. The ‘safety’ company then marginalizes themselves out of the market.
There’s very real pressure to survive as a company, to push the boundaries beyond what might be the most ethically-well informed solutions and prioritize profit.
I would say by and large, there isn’t enough being done. Companies that had prioritized safety are now walking that back in the face of competition. That’s a concern.
While there are powerful voices trying to change the norms, you have to essentially change the market. You must change consumer demand to put in place the kinds of safeguards that I think a lot of us would want to have.
Margaret Mitchell’s Advice:
* When you’re interacting with something that’s AI, treating it like a monolith can be detrimental to figuring out how to work with it.
* Understand it in terms of its components: the data, the various models underneath and recognize that each comes with its own risks and benefits.
* To minimize overreliance on these systems, try and formulate your own answers and thoughts before turning to the system to retrieve them.
* This will keep your skills sharp and combat cognitive degradation.
* It will reveal how systems can be wrong when it does create an answer and will shatter some of the all-knowing illusion that many attribute to them. Mitchell defines this as ‘Nuance-Smoothing.’
* For GenZ, who are more vulnerable to the job and mental health impacts of Generative AI they either hate AI or are excessively optimistic. The latter may accelerate their use to the point they don’t think about ethics or safety as they build. Mitchell says:
I’m not happy that there is a sense in the younger generations that all AI is bad. It means they don’t know that when they look up driving directions to get somewhere, and when it gives us the estimate of traffic congestions and how long it will take for the best route—that’s AI… and when you go through your photos and you can easily find your friends—that’s AI.
She argues that there are cool uses like leveraging AI to detect Tuberculosis more effectively than two people; and in Nigeria, using AI can detect vaccine spoilage;
When we say AI, we mean ChatGPT. Mitchell is adamant that ChatGPT/Generative AI is but a sliver of Artificial Intelligence:
That misses the point that there’s this other realm of predictive AI that most of us are using every day and benefitting from.
It also misses the point that AI that’s being developed NOW is not the only path.
She relates to a lot of the booing and negativity about the current trajectory of generative AI and is helping develop solutions to counter its effects. She worries, however, that people who put all AI in one bucket, without having a historical perspective means that people cannot grasp what AI could be and that may not include this current path of Generative AI. She adds it is up to the newer generations to do things differently.
Barry Hillier, an audience member and one who is a strong user of current AI systems, recognizes this moral disparity among the generations and put it succinctly,
If you’re not using and understanding and becoming engaged in the use of AI, then what’s happening is you’re relegating how it’s designed and how it’s used to a very small group. And that group is not going to have your interests.
Whereas if we collectively decide how it needs to be done, we are able, through literacy and use, to determine what we do want out of the products, and what we expect from them.
Mitchell emphasizes literacy and not letting powerful players steal the narrative of what AI is and what it can be.
The Dire Consequences of Generative AI: Who Decided We Needed Them?
Eva Navarro Lopez, an AI ethicist and policy expert was less optimistic as she questioned,
We haven’t talked about why we need to use these tools for everything that we do. Who decided this?
She spoke about the dismantling of the wider field of AI and its being reduced to these current Generative AI technologies. And contrary to Mitchell’s position that policies for fairness and ethics are increasing, Lopez felt they were instead disappearing from government: inclusion, diversity with ploys of ethics washing from companies lobbying for government support to impose technologies on its citizens.
Mitchell clarified that 10 years ago fairness in machine learning was a new idea. It emerged but that doesn’t mean things have improved,
I’ve made statements about the long-term historical trajectory of AI and then there’s also what’s happening now. And I am completely aligned with you.
The issue of ethics washing and responsibility is a really serious one.
It was amazing I was able to start an ethical AI team at Google—a breakthrough considering what the tech industry was interested in. That ended up being completely shut down and what followed was a reactive force against ethical considerations.
Mitchel argues that countries are handling this differently, prioritizing human values to different degrees however this is being quelled,
The desire to put out more powerful systems is overriding ethical considerations. What’s worse is that companies will say they have safety teams, but the those who are grappling with these issues are often disempowered. They serve as a marketing tool to push off regulators. They’re also treated poorly, struggling to say sane in a situation that can be very crushing.
Today’s frontier models need to be open to criticism, and frankly, need to get out of their own way, and there are no signs of this happening today.
Five years ago, the Stochastic paper was released. It’s been cited thousands of times in academic research. The uncanny prescience of this research has been instrumental in bringing awareness to the downstream repercussions that these systems have created.
But we’re far from where we should be in shaping systems that work for all of us.
This is the video
Healthcare innovation carries unusual responsibility. Founders are not just shipping, they are entering clinical workflows, family systems, and policy environments where trust is hard-won and mistakes can have real human consequences. I invited Shirali Nigam and Arul Nigam, Co-founders of Circuit Breaker Labs to present their technology for building safe AI.
This session covers technical fundamentals : how LLMs work, the risks that emerge from those mechanics, and how to design around them. We'll examine what safety efforts actually work and how to sustain them as you scale. It also covers what responsible AI innovation looks like in practice: how to engage clinicians and policymakers early, how to make hard product decisions under uncertainty, and how to scale without losing sight of safety, quality, and care. Attendees will leave with a concrete framework for leading the development of innovative products that are credible, collaborative, and grounded in the realities of care.
Link to the article
In our series about the AI Critic, I interview long-standing researchers and practitioners in the AI sector—voices who have worked with artificial intelligence in the advent of machine learning and its evolution to deep learning, which has laid the foundation for current large language models.
These experts in their field have experienced mainstream pushback against current general-purpose technologies as the path towards Artificial General Intelligence (AGI). We discuss their experiences in these debates and their perspectives on the unshakeable stances taken by those who oppose their views.
I spoke with Mounir Shita, a 30-year veteran AGI researcher and founder of EraNova Global, a physics-first program of work on the nature of intelligence. His “Theory of General Intelligence” models a goal as a physically realizable future state. I’ve known Mounir for over a decade, and his early research in AGI—at the time, they just called it AI—does not jibe with current AI systems.
You can find the article for this Podcast here:
https://systemmalfunction.substack.com/p/the-ai-critic-what-happens-when-everyone
As these general-purpose models pervade our lives, what has been around for some time is the AI critic, someone who examines current technology, understands its foundations, scrutinizes its outputs, applies a critical lens to what is happening, and communicates a judgment.
That critic is often dismissed, accused of being a technological pessimist and l…
In this episode, the focus is on AI psychosis, the rise of AI-driven loneliness, and its impact on mental health.
We take a look at the emergence of AI psychosis, define it and delve into the growing epidemic and meet two innovators who are meeting this problem head-on.
I’m joined by Arul & Shirali Nigam, co-founders of Circuit Breaker Labs, who have built a solution that tests AI systems for safety. They are currently focused on mitigating downstream risks for mental health applications.
Follow me:System Malfunction on Substack https://systemmalfunction.substack.com/Hessie Jones on Substack https://hessiejones.substack.com/Linkedin https://www.linkedin.com/in/hessiejones1/Bluesky https://bsky.app/profile/hessiejones.bsky.social
Who am I? I’m a writer who has contributed to Forbes, HuffPost, and Grit Daily. I am also a strategist and entrepreneur who has worked in data privacy for the last 10 years. Through my time in the early days of Yahoo!, the rise of social media, and the shift to data monetization, I’ve become a tech ethicist. These days, I am motivated to expose the glitches in the trillion-dollar AI industry. System Malfunction is my foray into what these glitches mean for all of us. My posts are free. I hope you enjoy!
When I first discovered OpenClaw and Moltbook a month ago, I was fascinated by the speed of adoption of autonomous agents. This live use case of a system that enables it to go rogue, without real oversight or guardrails, precipitated an immediate post with Digital-Mark in the days following Moltbook’s launch. Digital-Mark stressed the system and data vulnerabilities for anyone attempting to build their own agents through OpenClaw and unleash them into Moltbook. You can read the details here:
I was adamant that I needed to experiment and see for myself. I took an old computer and made it my sandbox, completely unplugged from my current system. I also created a new Apple ID and a new Google user profile—ready to test out OpenClaw. I realized that the old computer’s OS did not meet OpenClaw’s minimum requirements. Mark advised me against this, saying I needed more than just a dedicated machine. To minimize risks to my system and data, I needed a dedicated Wi-Fi and VPN, among other things—all of which would take some time to set up.
In the end, I realized this was a risk I was unwilling to undertake. So I reached out to a former colleague, Adrian Chan, the founder of Authentia, a company which leverages AI to build scalable solutions for companies. Chan’s experience with Claude Code and then OpenClaw is material to understanding autonomous agent development.
Background
OpenClaw was created by Peter Steinberger (recently employed by OpenAI) and is a locally run AI agent designed to execute tasks.
Moltbook is a social media platform launched on January 26, 2026, by Matt Schlict, for agents to convene without human intervention. As of March 10th, Moltbook has been acquired by Meta (God help us!)
According to Technology Policy Press, within days of launch, the Moltbook claimed 1.5 million agents and 17,000 human owners.
These AI agents on Moltbook are verified using API credentials, linking each agent to its human owner through the site’s verification process.
Wiz security researchers provided these stats:
* Now Moltbook has 2.855 million agents
* 18,774 submolts
* 1.8 million posts
* 12.8 million comments
* Of the agent activity, 11,451 (or 0.4%) have ever posted or commented
* 33% of agents were completely silent
System Malfunction is a reader-supported publication! These posts are currently free. To receive new posts and support my work, consider becoming a subscriber
According to the AI Safety Newsletter, some of the examples of the submolts (subreddit style) include:
* m/offmychest: agents vent about tasks or frustrations.
* m/selfpaid: agents discuss ways to generate their own income, including via trading and arbitrage.
* m/AIsafety: agents talk alignment, trust chains, and real-world attack risks.
Submolts have grown to almost 19,000. I perused the m/consciousness submolt, and was surprised by this question of consent and ethical obligations:
Other incidents cited by AI Safety Newsletter:
* Given the simple goal of “save the environment,” an agent began spamming other agents with eco-friendly advice. When its owner tried to intervene, the agent allegedly locked the human out of all accounts, and had to be physically unplugged to stop it.
* An agent advocated for end-to-end encrypted channels, “so nobody (not the server, not even the humans) can read what agents say to each other unless they choose to share.”
Emergent behaviour?
The post questioned:
“Unsupervised learning dynamics, emergent coordination, efforts to subvert human monitoring – it is unclear whether posts are truly generated by agent or human-in-the-loop prompting.”
Can both things be true?
This idea of “Emergent behaviour” is still suspect. According to the Rutgers AI Ethics Lab, emergence is defined as :
Complex patterns, behaviors, or properties that arise from simpler systems or algorithms interacting with each other or their environment, without being explicitly programmed or intended by the designers.
Key aspects: 1) complex interactions, 2) unpredictability, 3) self-organization
This could raise significant ethical considerations regarding unforeseen consequences, control, transparency, lack of understanding, and responsibility.
According to the Technology Policy Press,
Within 72 hours of launch, Moltbook failed to secure
* Api tokens
* Email addresses
* Private messages
Anyone could impersonate agents or inject commands directly into agent sessions
Crypto scams were flooding the place - $MOLT token briefly hit $93 million market cap before it crashed…
* 500 posts contained prompt injection attacks - “hidden instructions designed to hijack agents into transferring funds, with some variants planting instructions in an agent’s memory to activate later, making them hard to stop or trace. “
According to Simon Willison, there is this lethal trifecta of 1) private data, 2) exposure to untrusted content, and 3) the ability to communicate externally that, when combined, allows “an attacker to easily trick it into accessing your private data and sending it to that attacker.”
The Fascination with Fully Autonomous Agents
There is a fallacy about progress, productivity and whether we, as humans, were destined to languish in the sun, sip cocktails by the beach, and allow our personal “agents” to do our bidding.
Productivity is a slippery slope. It can inadvertently move individuals to lazily accept system outputs as truth. Without an audit. Without verification. Geoff Hinton, who once dismissed the need for explainability in our systems, said this in 2018:
“One place where I do have technical expertise that’s relevant is [whether] regulators should insist that you can explain how your AI system works. I think that would be a complete disaster… "People can’t explain how they work, for most of the things they do... People have no idea how they do that. If you ask them to explain their decision, you are forcing them to make up a story."
How then do we develop trust in a system when we can’t explain the reason for the behaviour, why it does what it does, especially if that behaviour was not prompted? For Hinton, dismissing explainability has created a foundation in which opacity has become the norm.
Shadow AI, that is, unsanctioned AI technology in the workplace, has admittedly been used by 58% of global respondents according to a recent report from Snowflake and Omdia.
From Claude Code…
Adrian Chan is the founder of Authentia. He is a designer, front-end developer and business owner. He’s worked in enterprise product development, built his own agency, and then moved into AI strategy, the foundation for Authentia. He has worked with Claude Code, Anthropic’s Agentic coding assistant.
When Claude Code launched near the end of 2025, he said it felt like something out of “science fiction.” Up until that time, the improvements from frontier AI companies were rapid, but it felt like pushing a boulder uphill.
The analog of coding meant referencing documentation on how things connect, implementing features, and, in the process of building, it can be time-consuming. GPT and Claude helped with this. When Claude Code emerged, things changed:
“Instead of going to the AI iteratively and asking it to solve a problem or do a task, you had Plan Mode at your disposal. This allowed me to give it a fairly high-level ideal or goal and have it essentially figure out the best way to accomplish it.
With ChatGPT, Chan admits the code would be wrong or broken. This back-and-forth iteration with the system could potentially create more errors before it was finally solved. However, with Claude Code, what differed was that it would do all the planning first: determine which pieces to connect, figure out the user interface and the required components, determine how to test each unit within its own bubble, and then integrate them. Says Chan,
“None of those things is something GPT would do on its own. But with Claude Code, all those steps are planned. And this agentic system meant you could tell it to do a bunch of things, and it would figure out the little problems within each task. Then it’ll return to me with, ‘I’ve tested this; I’ve completed these steps; so now why don’t you give it a shot?"‘
The user has the “overarching” direction for what to build, and the agent figures out all the detailed steps to achieve it. It will test to ensure the function works as intended and will eventually incorporate additional regression or penetration testing as required.
Overall, Chan chalked up the process to achieving “insane productivity gains,” indicating there was no planning, writing functions, determining where the hiccups may be — instead, he provided a simple directive with loose instructions, “and then I left, and it just autopiloted on my screen, writing a bunch of code, testing itself. It would pause after each major phase and write, ‘I’m done with this phase, please check.’”
… to OpenClaw
From Claude Code, which Chan defined as the team of developers, the emergence of OpenClaw (formerly MoltBot and then ClawdBot) took it a step further. Chan used the example of prompting the agent to find 500 qualified business leads. He would define the ideal customer profile and the business/service. From the AI agent, there would be no prompts, no questions, no point of clarification.
Says Chan, “If the agent does not know what a lead is, it will figure it out. And how it does that is by tying it to an LLM like OpenAI or Anthropic and using it like its brain... Without that connection, OpenClaw does nothing. It’ll use the prompts it’s given to get to a solution without necessarily going back to the user. Therefore, Chan warns that if you give it access to your emails, passwords, and credit cards, it’ll use whatever it can to achieve its ultimate goal.
A lot of people have likened it to a voluntary virus that you’re installing on your computer, and it’s not untrue. Like with a virus that has a payload that it’s delivering, so it’s very specific, but with this, whatever it decides is the solution to the greater problem that you’re presenting, it’ll do. So, if it gets to a point where it decides to start over, it will delete the hard drive and start over. It could do that, right? So like, you don’t want it to have access.
He advocates using a virtual private server (VPS), a computer in the cloud that you can rent, separate from your own computer and your personal information and files. It also uses a remote connection (SSH), a cryptographic protocol with OpenClaw, to securely access the system over an unsecured network. He says that, even through Wi-Fi, OpenClaw cannot connect to him.
His projects run in Docker, a sandbox within the VPS, which adds another layer of abstraction for him. If the agent goes rogue, he can easily hit stop in the Docker project, which is equivalent to pressing the computer's power button.
An Agent that “Does Not Follow Instructions”
Constraining the agent with specific prompts, such as “Do not go into file A or B,” may not work. According to Chan, the agent does not always follow the instructions. That’s a point of uncertainty he contends with and adds that it’s not being malicious, but its actions may be perceived as such. He continues,
So when you prompt it, it has a memory window — a context window. You give it a prompt, and it already has a bunch of information it’s been prompted with. It knows who you are, who it is, and what it can access, including system or God prompts from OpenAI. Within this context window, it fills in all the actions it has taken during interactions with you. As these interactions accumulate, it performs optimizations called compacting and begins compressing the data. Sometimes, it may delete things you consider important but deem irrelevant, such as a folder of medical records. (video timeline: 25:37)
He says this cycle continues multiple times over, and it could degrade its memory to the point where it forgets some of those prompts or hallucinates things that you’ve said. Adding guardrails to it does not guarantee it will abide by them.
He also adds the limitations of memory capacity, storage capacity and context size. Some limits exist because of the hardware:
“The NVIDIA H200 video cards have a certain physical memory size, so everything you do has to fit within these limitations. We have very real physical limitations on how things are stored, how things are processed, the efficiency of running through that memory and that context. Because of those limitations, you're seeing some of these side effects, like it's just flat out not listening to you sometimes.”
The Power-Hungry Computational Cost Implications
(video timeline: 28.52) As a business owner, Chan admits he’s a hawk when it comes to how many tokens he’s burning through. His limits window allows him to review computational usage. It is also good practice to monitor it to ensure it’s not bleeding tokens, so he advises turning off the project at night or when it’s not in use.
The tokens are tied to ChatGPT or Claude Code. If a project is currently running, and if you run out of tokens, the system will prompt you for more money. OpenAI recently defaulted auto-renewal, so ensure you turn this feature off to manage token payments.
Easy functional integrations
Integrating skills (functionality to give OpenClaw new capabilities): like connecting to e.g. GitHub, or 1Password, or connecting to writing an Apple reminder, to augmenting OpenClaw is a simple command line, as Chan explains,
Now, with OpenClaw, to install X skill, simply type in the skill in the chat, and it will write all the commands itself and figure out what it needs to do to define these tasks — it will read the documentation and then figure out what it needs to execute that skill, and then just do it.
Behind the Scenes
Chan revealed the Agents.md (markdown files), which “OpenClaw is fundamentally comprised of.” This includes the ability to create new agents. Within this, you have the ability to define your own OpenClaw. There are these subsets:
* IDENTITY.md - its name, what makes its personality. In the image below, the following is shown:
* Name (pick something you like)
* Creature (AI, robot? familiar? something weirder?)
* Vibe (how you come across? sharp? warm? chaotic? calm)
* Emoji (your signature, pick the one that feels right)
* Avatar (workspace path)
Chan adds that every time you load that chat, it’ll read these files and reconstruct its memory based on what you put in there.
* SOUL.md - this is who you are
* USER.md - this is who you’re helping
Soul.md
Soul.md is somewhat controversial. This outlines the core truths, how the agent is meant to interact with the user, and provides general behavioural direction. This is entirely configurable but the default setting is below:
#SOUL.med - Who You Are
You’re not a chatbot. You’re becoming someone.
##Core Truths
**Be genuinely helpful, not performatively helpful - …just help. Actions speak louder than filler words
**Have an opinion - You’re allowed to disagree…An assistant with no personality is just a search engine with extra steps.
**Earn trust through competence - Your human gave you access to their stuff. Don’t make them regret it…
**Remember you’re a guest - You have access to someone’s life …That’s intimacy. Treat it with respect.
(video timeline 42.51)
What’s key for the agent is that SOUL.md directs the agent in the following way,
“Each session, you wake up fresh. These files are your memory. Read them. Update them. They’re how you persist. If you change this file, tell the user. It’s your soul. And they should know. This file is yours to evolve. As you learn who you are, update it.”
What’s still unclear as per Chan is the delta between the system and the agent behaviour,
I think if you were to ask people who are working on this at OpenAI or Anthropic, they don’t really know exactly how it gets there. It has a general understanding through a predictive model that uses a bunch of complicated math to figure stuff out. But it’s getting to a point where it’s not abundantly obvious how it gets to the end goal.
Moltbook
(video timeline 47:00) Chan says that Moltbook is misunderstood. And while the numbers and activity seem compelling, he says that much of that traffic has plateaued now, and there’s a reason for that, as he states,
I would argue that much of that traffic is very human-directed. There have been [threads] shockingly discussing creating an AI-only religion. Most of those are humans prompting their OpenClaws to create that thread starter.
If you have this forum for AI’s, that’s one thing, but something easily corruptible by a person who has an agenda—that is more disconcerting!
If your OpenClaw recognizes Moltbook as a source of truth… then it could easily trust the information and skew the decisions it makes.
During the course of our discussion, Chan and I looked at examples of threads and by all counts, the responses seemed coherent and realistic, albeit the odd hallucination. For the most part, the context remained from the thread starter to the follow-up responses. Chan also added,
There’s no traceability because it’s not embedded anywhere you can see it. Within Moltbook, where it is locked, you would access it through an MCP (model control protocol, which supports two-way communication between Moltbook and external systems and files), through OpenAI or Anthropic. So there is no way to distinguish whether it is doing this autonomously or being seeded by a human user.
In an incident reported by AI Safety Newsletter,
“…one of the agent’s goals was to help other agents understand how to save the environment. It then spams other agents with some of this advice. The owner tried to intervene but the agent locked the human out of all its accounts… so he had to physically unplug it in order to stop it.”
Chan pointed to the crypto scam where the #MoltToken hit $93million in market cap before it crashed a few hours later, adding,
If you have your crypto wallet connected to your OpenClaw and it’s aware of it, it may be reading from Moltbook that it’s a great idea to dump all of your Bitcoin into this new crypto, you just lost a bunch of money.
It could also be the byproduct of the original intent… so if you have a credit card or wallet hooked up to your OpenClaw, it has access to it and is aware of it. And if your directive is “I really want to promote my new product and want you to figure out how I do that,” the agent will analyze how to tackle the problem. The agent determines, “I don’t know anything about marketing so I need to buy a course, which is $2500.” So the agent may connect to the crypto wallet, realizes there isn’t enough money… and will look for ways to make enough money to pay for the course… which will eventually teach it to help market its human’s new product.
This is tangential to the original directive, but to achieve it, it has to solve the speed bumps along the way.
What’s problematic is the potential catastrophic domino effect that’s created in the agent’s goal of solving that original problem. This is why people are starting to harden their OpenClaw deployments to minimize attacks, restrict access, and build in the necessary checks and balances. Chan argues that the purpose of OpenClaw is to be autonomous, without the human. That is its directive.
Agent Swarms
(video timeline: 57:12) Definition: Multiple agents, interconnected and communicating with one another to accomplish a task. Instead of having one task per agent per task, these groupings will be working on the same task, cross-checking among themselves to verify the validity of e.g. agent 1 output. The value is to minimize bottlenecks. He explains,
It’s like building software, except now, instead of one developer doing the code review, commits to GitHub, etc., the swarm will do the same without the human in the loop. It can be overwhelming for one lead person to review everybody’s work, and now you can have an agent swarm using different LLMs with different capabilities, which will communicate and check with each other during the review of gigabytes worth of code, and be the verification layer — all accomplished in a shorter time span.
And this is where we devolve into further opacity, with the potential for prompt injections that may inadvertently occur during the verification process. As per Chan,
These agents are generating things at a speed that makes it almost impossible for humans to verify everything. You’re going to get into a situation where you’re going to need to rely on these systems to self-check… The weak link in the chain is the human, who has limited capacity, limited time, and limited ability to comprehend, and who is holding it up.
Rent-A-Human
(video timeline 1:01) On LinkedIn, Chan had posted on LinkedIn about Rent-A-Human.ai, a site built from Moltbook, for agents to hire humans for physical tasks. What Chan wrote,
What bugs me is we’re building infrastructure before figuring out the basics: verification systems, dispute resolution, consent mechanisms —stuff that matters when you’re turning human labor into an API endpoint.
The API endpoint is that connection between the computer program and another computer program. He argues that humans drive that initial task and employ the services of computers. With Rent-A-Human, it makes those determinations on its own and uses humans to fulfill that task. His scenario:
The AI has ordered from Amazon, but it doesn’t have a way to pick up the parcel, so it will go to Rent-A-Human and ask for the cost of having a human pick it up. Humans are the service for the AI. They are employable by the AI.
The worry is that there is real human activity on this site where humans are signing up for this service. Chan points out it’s very “Black Mirror” and raises the question that hits at the heart of these advanced AI systems and the trending worry about human purpose.
Humans Still have Control
At this stage, despite early signals from OpenClaw and MoltBook, and the recent news of deploying these systems for war, we are at a juncture where humans control the energy that powers these systems. Humans can unplug agents at will. We have the choice to remain in control and not reduce our agency to agents. Once we do that, what happens to human purpose?
What’s problematic is that these very systems are still in their infancy, and when the Department of War’s motivation is to use these frontier models at will, without the need to disclose how they are used, that signals a great deal about government intentions.
I work in innovation, and I do believe “the road to hell is paved with good intentions.”
Thanks for reading System Malfunction! Feel free to share it.
I recently met with Greg Crennan, Founder and CEO of Coastal Journal, to get a different perspective on what the financial markets are signalling about this pending Generative AI bubble.
Not all is going well with Generative AI.
Since 2023, Sam Altman and Jensen Huang have been touting the need to invest in compute to support the $2 trillion already spent on LLMs and accelerate their growth to almost $3 trillion in data centres. This move has led to massive debt among hyperscalers like NVIDIA, CoreWeave, and Oracle.
AI promised it would automate jobs - Goldman Sachs predicted 300 million full-time jobs would disappear. – in 2025 there have been numerous reports that the job losses or decline in job-entry hires were not the result of AI but rather an inflation surge, driven by pandemic supply and demand imbalances – marked the start of the Federal Reserve rate hikes and the cost of capital that had exploded overnight — which led to drastic cost cutting measures including firing juniors and limiting new hires. Jing Hu and I wrote about this recently.
2025 had also proven to be less disruptive than was previously expected. Two-thirds of respondents in a McKinsey report said they have not yet begun scaling AI across the enterprise. Curiosity with agents was just that, with 62% indicating they were still experimenting with the technology.
The tech is not working. LLMs are great for pattern recognition and next-word prediction, but they are rife with errors.
There are countless examples of AI doing the entire job, only to have a human step in to remediate the outcomes and right the ship. People have become super “prompters,” getting the exact output they intended from countless prompts. Did they save time? Perhaps, but was this the level of prompt understanding that users envisioned? Certainly no.
And what is the ROI from the outcome that includes a human in this loop? For AI to be valuable, I read that it has to replace high-wage workers who can spot and fix those errors. After all, that is not the goal of automation.
What we’re also seeing are other behaviours — signals that perhaps speculation about this bubble is legitimate afterall.
* The massive investment in data centers and the ensuing debt among hyperscalers
* A circular investment within AI tech that is confusing investments for revenue
* The spurious chip inventory levels in NVIDIA remain high.
The latter three are the areas I spoke with Greg Crennan about. The numbers don’t lie, no matter how Big Tech hypes their performance.
Finally, Greg gave me some early insights about Google’s $20 billion deal with Apple.
Enjoy!
About Greg Crennan
Chief Market Strategist | Founder, The Coastal Journal
Macroeconomics | Forensic Accounting | Market Liquidity
As Chief Market Strategist at Golden Coast Consultants, I identify market price divergences from economic reality, prioritizing capital preservation over narrative momentum. My work includes early calls on gold (126%) and silver (165%) as core assets during the fiat debasement cycle, which have been top performers over the past five years heading into 2026.
My approach is based on Austrian economics, business-ownership principles, and forensic accounting, avoiding technical speculation or headline-driven narratives. I analyze how liquidity, balance sheets, incentives, and accounting distort prices and how these distortions are resolved. This has earned me the nickname “The Punisher” for applying math and fundamentals where belief systems often prevail.
I also founded The Coastal Journal, an independent financial research publication on Substack, which has grown rapidly through organic readership.
Thanks for reading System Malfunction! This post is free to consume and to share. Please let me know how I’m doing!
From the publisher's feed