Ep. 61: Everyone Gets an Agent: Meta Muse and the Ongoing Race to AGI
In this episode I talk about the eleventy billion AI models that dropped in the past few weeks, including Anthropic’s Fable 5.1 and Mythos 5.1, OpenAI’s GPT-6 Astra, and Meta’s new Muse agent. I get into former Anthropic researcher Jacob Coxon’s resignation and his warning about AI risk, plus what it actually means that OpenAI just rated Astra “Critical” for cybersecurity capability. I also share why I think Meta Muse is a terrible idea and what all of these model drops say about the direction AI is heading.
Many Truths All at Once
Before we hop into today’s main topic and discuss the the eleventy billion models that have been released in the past 5 minutes, I want to address something that you’ve likely seen in the media as of late: the resignation of former Anthropic researcher, Jacob Coxon, and his statement on X that, “The people building AI earnestly believe that it could kill us all by the end of the decade.”
Following his resignation, Coxon shared a 7-part thread on X, and shortly after, the bigwigs of AI released their own statements, calling for a slowing of the pace of development of AI as a means of ensuring public safety.
I could probably dedicate a full episode to all of this,, and likely will, but what I want to drive home right now is simply the reality that many things can be true at once:
- AI regulation is, and has always been, needed. I’ve spent the past few episodes highlighting the fact that I’m more concerned about AI’s hacking capabilities than anything else, and I continue to hold that position. This technology is extremely capable and will only get better, and we absolutely need regulation.
- The AI leaders calling for the slowing of AI advances are telling the truth AND also likely acting in their best interest, spurred by financial motivations.
- Jacob’s warning should be heeded, but he doesn’t deserve extra applause or absolution for alerting us to a fire that he helped spread.
Like I said, I’ll likely do an episode sharing more of my thoughts, but I continue to believe that aside from completely pulling the plug on AI (which ain’t gonna happen), the best thing we can do is continue to educate ourselves and stay informed so that we can advocate accordingly.
Part of what Jacob wrote in his X post was: “A common response is ‘if they truly believe this, why are they still building it?’ At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first – they believe no one else will act responsibly, so they must do it themselves, despite the risk.”
Honestly, that’s kinda where my head is at with things. I absolutely believe that individuals can bring about change, but this (AI) feels like way too big of a ship to sink (AI is literally propping up our economy), so the best move feels like learning as much as we can.
So Let’s Talk About the Recent Model Drops
So, in the spirit of learning more, I want to chat about the recent model drops and what I think it demonstrates that these companies are up to.
Here’s what was released during the first week of September:
- Anthropic: On Step 1st they released Fable 5.1 and Mythos 5.1; same underlying model but targeted updates
- OpenAI: On Sept 3rd they released GPT-6 (Astra), labeled its most capable model yet, including critical cyber capabilities
- Meta: On Sept 8th they released Meta Muse; everyone gets an agent!
Diving a Little Deeper: Anthropic
The Anthropic releases weren’t super revolutionary, and it wasn’t an entirely new model that was released, but rather a “targeted update, not a uniform intelligence jump.” With the release of Fable 5.1 and Mythos 5.1 we saw:
- Scientific research benchmark: more than doubled
- Business workflow automation: up 84% (relative)
- Agentic coding (Terminal-Bench): up almost 14 points
- Short, simple tasks: barely moved, a few points at most
The standout story to me is that Mythos 5.1 (still limited to vetted organizations) has better “covert capabilities” than its predecessor, and its cyber skills are described as “getting close to” the next danger tier without crossing it yet. :::Insert Ralph Wiggum “I’m in danger” GIF:::
For both Fable and Mythos, the longer and more autonomous the task, the bigger the jump in capability with this release. Will you actually notice the difference? If you’re just using it for simple Q&A, likely not. If you’re performing long-running agentic work, research, coding? Very likely.
The TL;DR – agentic workflow improvements.
Diving a Little Deeper: OpenAI
On September 3rd Open AI released what some are calling their most capable model to date: GPT-6, also known as Astra. Astra isn’t necessarily “smarter” as assessed via benchmarks, but it’s a big jump in capability, particularly on a computer-use benchmark.
“Computer use” means the model can take a goal, then autonomously navigate software, take actions, see what happens, adjust, and keep going, all without a human directing each individual step.
I know that was a lot of words, so here’s what I need you to understand: the same autonomous multi-step operating ability that lets it fill out a spreadsheet unsupervised is what lets it hunt down and weaponize a security hole unsupervised. (Did someone say hacker?!)
So again, who notices this difference in capability? Well, for starters, Astra is only available on paid tiers via Work and Codex, so that’s the floor as it relates to who can even access it in order to notice any changes. But, as it relates to anyone with a paid account, casual chat users likely won’t notice any difference. Developers, researchers, and anyone running long multi-step or multi-app work, they gonna notice.
As for the critical cyber capabilities that I mentioned in the bulleted list two sections ago (and honestly this is probably the most important thing to flag with this model release): on Sept 1, OpenAI reported, “We now believe Astra meets the Critical cybersecurity capability threshold under our Preparedness Framework. It is the first model we are designating at this level.”
OpenAI has an internal risk ladder for how dangerous a model’s cybersecurity skills could be. The top rung, “Critical,” means the model can find and build a working exploit for an undiscovered flaw, or plan and execute a whole cyberattack, on its own, no human walking it through each step.
Every prior OpenAI model topped out one rung below that Critical level. Astra is the first model that OpenAI flagged, before they even launched it, and then after Astra scored a perfect 100% on ExploitBench, OpenAI gave it the Critical rating. Again :::insert Ralph Wiggum “I’m in danger” GIF:::
At the Astra launch press briefing following the release, OpenAI’s president, Greg Brockman, was quoted as saying, “Welcome to the AGI era”, and told reporters he personally believes the company is “there.”
Quick Refresher: What Is AGI?
Throwing it back to episode 7, artificial general intelligence, or AGI, is AI that can understand, learn, and perform any intellectual task a human can, across different domains, not just the narrow things it was trained for. (Think, Skynet. Where my Terminator fans at?)
From minute one, AGI has been the holy grail of what these AI companies have been chasing, seemingly largely so they can make a ton of money by owning a technology that can replace workers.
So with this most recent release we have OpenAI saying they’re there, safety-concerned researchers moving the goalposts and are warning about superintelligence (beyond human-level, smarter than the best humans at basically everything), and an Anthropic employee resigning and issuing a formal warning on X because he says these companies won’t stop until it’s too late.
Meta Impersonates Oprah
So, amidst all the warnings and concerns, what does Meta do? They go and perform their best Oprah impersonation and everyone gets an agent!
On September 8, Meta launched Meta Muse, Meta’s standalone personal AI agent. It’s not a chatbot, but rather an agent that acts on your behalf. “It can take real actions across the web and connected apps, work through multi-step tasks on its own, and keep running in the background even after you close the app.”
In an August essay, Zuckerberg wrote, “Everyone will have an exceptionally capable personal agent that understands you, your goals, and everything you care about.” What could possibly go wrong?!
Per the release info page, Muse can answer questions, complete tasks, browse the web, make purchases, generate images, create documents, and connect with your favorite apps and services via Connectors. It can also set reminders, track goals, and monitor things you care about, all working in the background on your behalf.
The Details on Meta Muse
How to actually access it:
- Download the Muse app; it also works inside WhatsApp; runs on Meta’s Muse Spark model family
- US only at launch, 18+ only
- Free tier reportedly caps around 100 million tokens/week
- Paid tiers reportedly $20/month (“Power”) and $100/month (“Maximum”)
- Reportedly requires a payment card on file, even for the free tier
The privacy/security pitch:
- Isolated virtual machine per user, your agent and data aren’t pooled with others
- You choose which apps/services it can access, disconnect them at anytime
- Can opt out of having your Muse conversations used to train future Meta models
- Encrypted version promised later in 2026
What’s still unproven:
- No independent benchmarks yet against Astra, Claude, or Grok on agentic tasks, capability claims are Meta’s own
- Early internal testing reportedly flagged dropped sessions and sensitive-data exposure concerns (this agent touches inbox, calendar, payment info)
- Mandatory card on file for a “free” tier raises red flags
- Regulatory attention will likely come down the pipeline as cross-app data reach across Facebook/Instagram/WhatsApp plus email and payments has drawn scrutiny in the past
I think that Meta Muse is a terrible idea, and honestly it feels like a push to get agents into the hands of a group of people who haven’t bought in yet, but absolutely have money: Boomers. Not that Boomer and tech are inherently a bad thing, but the same way our parents are the biggest offenders when it comes to AI generated images (both creating them and falling for them), I think that this will end poorly.
How I Used AI This Week
Each episode I share a quick example of how I used AI that week. And yes, I’m aware of the absolutely diabolical dichotomy of writing that these AI companies basically don’t care if they kill us, and then in the next sentence sharing how I used AI. Gimme some time to collect all of my thoughts and organize them into an episode.
This time I just want to highlight that I’ve been using the Gemini AI Overviews in Google way more. I do find it VERY annoying that when you hit “show more” it takes you into a different window if you’re on mobile, but overall I’ve definitely been utilizing the Overviews way more.
This uptick is possibly because my questions have been more related to looking for information as opposed to something like a store, but either way, I’m definitely noticing that I’m opting for the AI overview way more. What about you? Using them? Still hate them? Wishing they were different? Hit me back and let me know. For real though. I’m genuinely interested.
Da Wrap-up
Overall we see that these companies are continuing to push full steam ahead, despite any and all concerns, and I do believe that money is the ultimate driving force.
I’ve always been a proponent of looking at actions more than just listening to what is being said, and the speed of these model releases, along with the type of capabilities that are being focused on (agentic), speaks to the priorities of these companies. Hint: it’s not the betterment of society. Cheers to staying informed.
As always, endlessly grateful for you and your curiosity.
Catch you next Thursday.
Maestro out.
