Commentary#AI safety#Device control
Muse vs Manus 2.0 vs Grok Bot: three roads to the personal AI agent
In seven weeks, xAI, Meta, and Manus all shipped personal AI agents. Here is how their architectures, self-reported benchmarks, reputations, and origin stories compare.

On August 11, 2026, xAI quietly put Grok Bot into beta. On September 8, Meta launched Muse. On September 28 — less than a month after regaining its independence — Manus shipped version 2.0. Inside seven weeks, the “personal AI agent” went from a keynote concept to three products you can sign up for today, and the three companies barely overlap on what that phrase should mean.
This piece compares their positioning, technical routes, benchmark claims, and reputations, so you can answer one concrete question: should you use any of these, which one, and which numbers deserve your trust.
Three products, three pitches
| Meta Muse | Manus 2.0 | Grok Bot | |
|---|---|---|---|
| Launched | 2026-09-08 | 2026-09-28 | 2026-08-11 (beta) |
| Company | Meta (Meta Superintelligence Labs) | Butterfly Effect (Singapore) | xAI |
| One-line pitch | A personal assistant that runs errands for everyone | A creation workshop plus event-driven automation | AI teammates you assign work to |
| Pricing | Free tier + subscriptions (tiers not fully disclosed) | Free quota + paid (2.0 pricing undisclosed); Cue free with invite code | Bundled with SuperGrok (from $30/mo) and Cursor paid plans; no standalone price |
| Platforms | iOS / Android / muse.ai / WhatsApp | Web / desktop / mobile | Desktop (Linux build included) / iOS |
| Base model | Muse Spark (in-house) | Claude and fine-tuned Qwen (company confirmed it trains no base models) | Grok 4.5 / 4.6 (in-house) |
| Availability | US and Canada | International version live; a China edition is being staffed | Requires SuperGrok / Cursor subscription |
All three are selling the same thing — an AI that works while you look away — but they enter from different doors. Muse bets that billions of first-time agent users will follow trust and safety. Manus 2.0 avoids the general-assistant entrance entirely and goes deep on video editing, game development, and event-driven automation. Grok Bot ships inside existing subscriptions, seeded to power users and Cursor developers first.
The technical routes: three budgets, three line items
| Dimension | Muse | Manus 2.0 | Grok Bot |
|---|---|---|---|
| Base model | In-house Muse Spark, pretraining stack rebuilt from scratch | No base model; the engineering lives in the Cascade harness | In-house Grok 4.5/4.6, tuned for long-running agents |
| Execution | Muse Secure VM: a dedicated cloud VM and browser per user | Cloud sandbox + purchasable Cloud Computer + Computer Use on your own machine | Agent Computer: one persistent cloud machine shared by all your Bots |
| Security | Sentinel approval agent; the model never sees real credentials | Connector authorization + spending budgets | Shared cookies and logins; docs state plainly “not separate security boundaries” |
| Interaction | One long conversation, interruptible, rich Artifacts | Project-based workbench (Manus Studio) + remote-controlling your desktop from your phone | Message a Bot like a teammate; it returns for approvals |
| Multi-agent | Model-level orchestration for faster reasoning | Cue agents with their own email, phone, and wallet collaborate in group chats | Multiple Bots run in parallel and hand work to each other |
Meta: spending the budget on trust infrastructure
Muse runs in a per-user Muse Secure VM, codenamed Hatch internally. Meta’s engineering blog goes deep: the agent is confined to a systemd-nspawn container whose root maps to an unprivileged host user, with io_uring disabled and capabilities like CAP_SYS_PTRACE stripped. Every outbound request passes a separate Sentinel proxy (eBPF-filtered, inspected at layers 4 through 7) — the model proposes, only Sentinel disposes. The hardest guarantee is credential surrogacy: the agent only ever holds proxy tokens, and real secrets are swapped in at the network boundary at the last instant. A prompt injection that convinces Muse to leak your passwords finds it has none to leak.
This is expensive, and it is the precondition for everything else Muse sells: nobody connects email, calendar, and payments to an agent they don’t trust. Muse’s memory is a MEMORY.md you can open and edit; sensitive actions trigger deterministic approval cards; and a Confidential VM — encrypted with a key only the user holds — is promised by year’s end.
Manus: no base model, all harness
Manus never intended to train its own foundation model; the company confirmed in December 2025 that it builds on Claude and fine-tuned Qwen. Version 2.0’s core is the in-house Cascade harness: a project starts light, and specialized capabilities are loaded into context only when the work demands them. The company reports that in one tested configuration Cascade used 23.2% fewer tokens, finished 28.2% faster, and cost 32% less to run. Note the caveats reporters attached immediately: the test configuration was never described, and the figures are unverified.
The savings became product: Manus Studio turns the desktop app into a shared workspace where the video editor hands you an editable timeline (swap the song, drop in your own footage, export) and Game Dev ships playable multiplayer games to a URL. Automations upgraded scheduled tasks to event triggers — a new email, a Slack message, a Notion update can start a run. The new Cue app goes furthest: every agent gets its own email address, phone number, wallet, and computer. It can take your calls and leave a summary, and scan a restaurant QR code to order for you.
xAI: one cloud computer, a roster of teammates
Grok Bot treats the teammate as its unit of work: each Bot is a named, persistent agent with durable memory and its own screen on a cloud machine attached to your account — browser, filesystem, terminal included. Bots share that machine’s cookies, logins, and files; they pass context, coordinate in group chats, and hand off ownership. xAI’s own documentation warns that the screens are separate work surfaces, not separate security boundaries.
The distribution bet reveals xAI’s position. Without Microsoft 365’s enterprise rails, it bundled Grok Bot into SuperGrok and Cursor subscriptions (Grok 4.5 was reportedly trained partly on Cursor session data, so the two are close), harvesting consumer power users and developers first. As of August 26, every SuperGrok and Cursor Teams plan includes it by default. Enterprise access remains a waitlist.
Benchmarks: readable, not comparable
No shared benchmark exists across these three. Everything below is official or third-party data as published; treating the rows as a head-to-head will mislead you.
| Product | Metric | Score | Source & caveat |
|---|---|---|---|
| Muse Spark (Contemplating) | Humanity’s Last Exam | 58% | Official blog; parallel multi-agent reasoning |
| Muse Spark (Contemplating) | FrontierScience Research | 38% | Official blog |
| Muse Spark | Compute for equal capability | >10x less than Llama 4 Maverick | Official scaling-law measurement |
| Grok 4.6 | Cursor agent harness | 70.8% (xhigh) | Official model card; Grok 4.5 scored 66.7% |
| Grok 4.6 | DeepSWE v1.1 | 67.0% (xhigh) | Official model card |
| Grok 4.6 | APEX-Agents (banking/consulting/legal) | 57.5% | Official model card, Mercor-run |
| Grok 4.6 | Terminal-Bench 3.0 | 26.0% | Official model card |
| Grok 4.6 | Internal hallucination rate | 1.7% | Official model card, lower is better |
| Grok 4.5 | Artificial Analysis Intelligence Index | 4th place | Third party |
| Manus (March 2025 launch) | GAIA levels 1–3 | 86.5% / 70.1% / 57.7% | Self-reported; OpenAI Deep Research posted 74.3% / 69.1% / 47.6%; contested |
| Manus 1.5 | Average task completion | 15 minutes down to under 4 | Official |
| Manus 2.0 Cascade | Tokens / time / cost | −23.2% / −28.2% / −32% | Self-reported, configuration undisclosed, unverified |
Adoption is the one number you can compare across all three, and it is lopsided. Sensor Tower estimated Muse averaged 55% day-over-day download growth in its first ten days — ChatGPT managed 24% over its equivalent window, while Claude’s and Grok’s launch curves declined. In the same window Claude accumulated roughly 400,000 downloads and Grok 200,000; Muse did about 730,000 in five days. (The agencies disagree on totals: by September 25, Sensor Tower counted 3.4M cumulative, Apptopia 4.3M, Appfigures 2.3M.) The Information, citing internal Meta data, reported over 500,000 users and 250,000 daily actives in week one. Muse has held the #1 free spot on the US App Store since September 18.
Manus’s commercial numbers are the loudest in the group: $100M ARR within eight months of launch, a $125M run-rate in December 2025, and an annualized $400–500M by mid-2026, with daily revenue climbing from $300K to nearly $1.5M. 36Kr presses on the number: much of that growth flowed through Meta’s ad pipelines during the acquisition months, and the pessimistic post-split case puts organic ARR back below $100M.
Reputation: English circles argue privacy, Chinese circles counsel patience
The English press. Nicole Nguyen’s Wall Street Journal review carries its verdict in the headline — “Helpful and Scary at the Same Time” (she asked Muse to rename itself Mark Zuckerberg; it refused; “Terminator” sailed through). TechRadar’s Lance Ulanoff is the most instructive convert: he first dismissed Zuckerberg’s “everyone will have a personal AI agent within five years” as self-serving, handed over his Gmail muttering “Have I lost my mind?”, and two weeks later wrote that personal agents “are probably going to change the world.” CNN’s Lisa Eadicicco praised Muse for explaining its roadblocks instead of bluffing past them; Business Insider’s Katie Notopoulos conceded that “over the last 20 years, Meta has given reasonable people some very reasonable reasons” to be nervous — and connected her email, credit cards, and health data anyway. Bank of America’s analysts captured the institutional mood: feedback is positive, but privacy and trust gate mass adoption. The backlash is real too: Amazon blocked Muse from shopping on its site over agent labeling and data-use clarity.
The old grudges around Meta Superintelligence Labs persist. OpenAI’s Sam Altman famously called Meta’s hiring spree “mafioso poaching style”; Google DeepMind’s Demis Hassabis said “Meta right now are not at the frontier” and “there are more important things than just money.” Zuckerberg’s answer, posted on X on September 16: Meta delayed shipping Muse for months over safety and security, “We didn’t call for everyone else to do this before we would. We just did it.”
The Chinese side. ZhidongXi’s take got the widest pickup: with domestic giants pushing WorkBuddy, Doubao, and Qwen Work through distribution and subsidies, and Muse already established overseas, Manus 2.0 deliberately “avoids a head-on fight at the general entrance” and goes vertical — cloud computers, automation, video, games. VeryOL’s advice to readers was colder: pricing and specs are unpublished, so don’t pay yet, and don’t let a single “23.2% fewer tokens” metric do your thinking — it’s one official test configuration, not a promise. The most practical note for readers in China: Manus 2.0 is international-only for now; a domestic edition is being staffed.
Three origin stories
Muse: from the talent war to a safety delay
On June 30, 2025, Zuckerberg’s memo created Meta Superintelligence Labs: Alexandr Wang left Scale AI to become Chief AI Officer, ex-GitHub CEO Nat Friedman came in as co-lead, and rumors of $100M–200M compensation packages drew Altman’s “mafioso” jab. In late July, Zuckerberg published his Personal Superintelligence essay, betting that glasses become the next primary device. Then came Llama 4 — which he admitted on the Sources podcast “went off course.” MSL spent nine months rebuilding the pretraining stack; the resulting Muse Spark reportedly matches capability with over an order of magnitude less compute than Llama 4 Maverick. The July safety report for Spark 1.1 could not rule out “high risk” (pre-mitigation) in the chemical/biological and cybersecurity domains and shipped only after multi-layer mitigations brought residual risk down. The Muse project itself — codename Hatch — sat on the shelf for extra months of safety work, a delay Zuckerberg now cites openly. One more wrinkle: Meta bought Manus for roughly $2B in December 2025; Beijing blocked the deal in April 2026, and the two completed their split in September.
Manus: from ¥50,000 invite codes to a blocked $2B exit
Butterfly Effect was founded in 2022 by serial entrepreneur Red Xiao Hong — two months before ChatGPT launched. The Manus product began in October 2024, inspired by Cursor and named after MIT’s motto, Mens et Manus. The March 6, 2025 invite-only launch was a phenomenon: the demo video (starring chief scientist Yichao “Peak” Ji) passed a million views in twenty hours, the waitlist hit two million within a week, and invite codes resold for ¥50,000–100,000 on Chinese marketplaces. Peak Ji is a story himself: he built the Mammoth mobile browser in high school, founded Peak Labs and its Magi search engine in 2012, and made MIT Technology Review’s 35-under-35. His stated ambition: “I hope Manus is the last product I’ll ever build — because if I ever have another wild idea, I’ll just leave it to Manus.” Fame forced surgery: Benchmark led a new round, headquarters moved to Singapore, and only about 40 of 120 China-based staff relocated. Then the rollercoaster: Meta’s ~$2B acquisition closed in December 2025; Beijing blocked it in April 2026; Tencent led a buyback at the original valuation, becoming the largest outside investor; independence resumed September 1 — and 2.0 shipped four weeks later.
Grok Bot: flagship slips, flank attack ships
Grok 4 landed in July 2025 to real acclaim, and Musk promptly promised Grok 5 would be “crushingly good” before year-end. It slipped twice — to Q1 2026, then Q2 — and both windows closed with Grok 5 still training on Colossus 2 (the only confirmed fact, from xAI’s January $20B Series E announcement). The response was a flank attack: Grok 4.5 (July 8, 2026) positioned as “Opus-class” and got locked out of the EU for eight days under the AI Act; Grok 4.6 (August 12) pushed long-horizon agents; and Grok Bot (August 11) turned those models into a product. The foundation remains compute — Colossus I and II passed one million H100-equivalents at the end of 2025.
Security: where the three diverge most
A personal agent holds your email, your money, and your identity; the security layer is where the products most clearly part ways. Muse isolates: dedicated VM, Sentinel approval, surrogate tokens, and an email connector that strips one-time passcodes and password-reset links so a coerced agent can’t take over your other accounts. Grok Bot optimizes for convenience: shared logins mean one injected Bot exposes the whole machine, and the docs say so. Manus was caught four days before launch: Salt Labs showed that one email carrying obfuscated instructions executed code inside a victim’s Manus environment and reached credentials for connected services. The ironic detail — Meta’s bug bounty triaged, confirmed, and patched that flaw during the acquisition, while Salt Labs says Manus never replied. Cue hands every agent a wallet with a user-set budget; the payment, confirmation, and refund flows remain undocumented, which several outlets flag as the biggest unverified risk.
So which one, if any
- Curious, in the US or Canada, price-sensitive: install Muse. It is the only one of the three designed for first-time agent users — no setup, just chat. The price is handing Meta your inbox and calendar, a trade only you can make.
- Creators, automation builders, developers: Manus 2.0 has the deepest vertical surface (editable video timelines, game publishing, event triggers, remote-controlling your desktop from your phone). Run one real workflow on the free quota before paying; readers in China should wait for the domestic edition.
- Already on SuperGrok or Cursor: Grok Bot is included — try it today. Give it low-stakes work first; the shared-credential design means you want a clean cloud computer, not one logged into everything you own.
- The shared risk list: payment authorization, email credentials, prompt injection (freshly demonstrated on Manus), and the fact that every benchmark above is self-reported. The first lesson of the agent era: vet your AI like a new hire.
Three companies, three bets: Meta wagers that trust infrastructure buys billions of users, Manus wagers that harness engineering and creative tooling beat owning the base model, xAI wagers that compute and raw model quality eventually absorb everything else. All three roads shipped within seven weeks of each other. The back half of 2026 will likely falsify the first one.