All Articles
6 minUpdated

Can You Build Your Own Meta Muse With Claude Sonnet 5.5 and MCP?

Most RecentTrendingAI Agents

Short answer: Claude Sonnet 5.5 plus MCP servers covers the model and the tool layer of Meta Muse today. The runtime around them (policy gate, token vault, memory, job queue, audit log, injection defenses and a browser sandbox) is still yours to write. This post lists each part and shows the loop that ties them together.

Why should builders care about this launch week?

Two launches landed three weeks apart. On September 8, 2026 Meta announced Muse, a personal agent that browses, sends email, shops and keeps working after you close the app. On September 28 Anthropic shipped Claude Sonnet 5.5 at $2 per million input tokens and $10 per million output tokens, with an 80.1% score on the OSWorld 2.1 computer-use benchmark.

If you build products, the useful question is how much of Muse you could ship inside your own app next month. I split Muse into the parts Meta actually shipped, marked the ones Claude and MCP cover today, and sketched the rest as code.

What did Meta actually ship?

  • Muse Spark, the model.
  • Muse Secure VM: one isolated virtual machine per user, with its own browser and storage.
  • Sentinel: a separate policy agent that allows, denies or escalates each connector request and outbound call.
  • Approval prompts before email or purchases, and a full audit trail.
  • Memory of routines and one-off details, used for proactive suggestions.
  • Background execution that continues after the app is closed.
  • Payments through Stripe's Link with single-use card numbers, with Shop Pay and 1Password announced.
  • Surfaces: iOS, Android, the web and WhatsApp, rolling out in the US.

Only the first item is a model. The other seven are infrastructure.

Animated diagram of a Meta Muse browser task passing through the Secure VM, Sentinel and user approval
How Muse runs a browser task: plan, read the accessibility tree, fill the form, Sentinel check, user approval, audit log.

What does Claude Sonnet 5.5 bring to an agent?

Sonnet 5.5 is the planner in this design. The numbers that matter for an agent are these: output about 30% faster than Sonnet 5, up to 30% lower cost per task because it uses fewer tokens and tool calls, 80.1% on OSWorld 2.1 against 81.8% for Opus 5.5, adjustable effort levels, and availability on the Claude API, AWS, Google Cloud and Microsoft Azure. Lovable reported roughly one third fewer tool calls after switching.

For an agent that makes dozens of tool calls per task, the drop in calls matters more than a benchmark point. Every call you avoid is one less network round trip to an MCP server and one less place for the run to fail.

What does MCP cover out of the box?

The Model Context Protocol is the open standard Anthropic created in 2024 and handed to the Linux Foundation's Agentic AI Foundation in December 2025. It gives the agent one way to reach every tool. A starter set for a Muse-like assistant looks like this:

NeedMCP serverStatus
EmailGmail connector or remote serverReady
CalendarGoogle CalendarReady
DocumentsGoogle DriveReady
Browser@playwright/mcpReady, you host it
Codegithub-mcp-serverReady
Data and memory rowsPostgres serverReady
Internal APIsSmall custom server with the TypeScript or Python SDKYou write it
Animated diagram of Claude Sonnet 5.5 routing tool calls through an MCP client to six MCP servers
One tool call at a time: Claude picks the tool, the MCP client routes it, the server calls the real API, the result returns to context.

Which seven parts do you still write?

  1. Policy gate. Tag every tool as read, write or money. Allow reads, send writes to an approval queue, block anything unlisted. A tool_policies(tool, risk, requires_approval) table is enough to start.
  2. Token vault. Encrypt refresh tokens per user with a KMS or libsodium, refresh them in a worker, and never put them in the prompt.
  3. Memory. Postgres with pgvector, a memories table that stores the source and the timestamp of each fact, and a settings page where users can read and delete entries.
  4. Jobs. A queue such as BullMQ on Redis or a managed queue such as SQS, plus a scheduler for requests like "in five days".
  5. Audit log. An append-only agent_events table with task, tool, input hash, result summary, approver and time.
  6. Injection defenses. Mark tool results that come from the web as untrusted. After the agent reads untrusted content, allow only read tools until the user approves the next write.
  7. Browser sandbox. One container per user session running Playwright, with a live view and a kill switch.
Architecture diagram of a Muse-style agent with Claude Sonnet 5.5 at the center, MCP servers, APIs, isolated browser, memory, workers, policy gate and audit log
The full stack: the model in the middle, and everything around it is infrastructure you run.

What does a minimal agent loop look like?

This is the core of the runtime in TypeScript. The model and the MCP servers do most of the work. The loop adds the policy check, the log and the pause.

import Anthropic from "@anthropic-ai/sdk";
const claude = new Anthropic();

export async function runTask(task: Task) {
  const messages = await memory.load(task.userId, task.goal);
  const tools = await mcp.listTools(task.userId); // tools from every connected MCP server

  for (let step = 0; step < 40; step++) {
    const res = await claude.messages.create({
      model: "claude-sonnet-5-5",
      max_tokens: 4096,
      tools,
      messages,
    });
    messages.push({ role: "assistant", content: res.content });
    if (res.stop_reason !== "tool_use") return finish(task, res);

    const results = [];
    for (const call of res.content.filter((b) => b.type === "tool_use")) {
      const verdict = policy.check(task.userId, call); // "allow" | "ask" | "deny"
      if (verdict === "ask") return queue.pauseForApproval(task, messages, call);
      const output = verdict === "allow"
        ? await mcp.callTool(task.userId, call.name, call.input)
        : "Blocked by policy.";
      await audit.log({ task: task.id, tool: call.name, verdict, output });
      results.push({ type: "tool_result", tool_use_id: call.id, content: output });
    }
    messages.push({ role: "user", content: results });
    await memory.checkpoint(task.id, messages);
  }
}

pauseForApproval saves the conversation and the pending call. When the user taps approve, a worker runs that call, appends the result and starts runTask again from the checkpoint. That small function is how "keeps working after you close the app" happens.

How does the job-hunt task look in the audit log?

Here is the example request "find remote React jobs, check my calendar, prepare applications, follow up in five days" as the log would record it:

09:02  plan        search remote React roles, posted in last 7 days
09:02  browser     navigate jobs board        allow   42 listings (untrusted)
09:03  postgres    read prefs                 allow   remote, React, salary floor
09:03  calendar    find_free_slots(14d)       allow   6 slots
09:05  drive       create_doc x5              allow   5 cover letters
09:05  gmail       create_draft x5            ask     waiting for user
09:41  approval    user approved 4, rejected 1
09:42  browser     submit application x4      allow   4 submitted
09:42  scheduler   follow_up in 5 days        allow   job #8812
day 5  gmail       search replies             allow   1 reply, 3 silent
day 5  gmail       create_draft x3            ask     waiting for user

Each line marked ask is a point where Muse's Sentinel would also pause. Each line with a job number is something a chat window cannot do on its own.

Should you use Muse or your own stack?

Use Muse if you are one person who wants results and you are fine with Meta's data terms, its US-only rollout and its current integration list.

Build on Claude and MCP if the agent has to live inside your product, reach internal systems, follow your retention rules, or switch models later without rewriting every integration.

Verdict

Sonnet 5.5 gives you a planner that sits close to the top model at a mid-tier price. MCP turns the tool layer into configuration. Muse's real lead is its runtime: the VM, the policy agent, the vault, the scheduler and the log. That layer is where agent products will compete next, and it is also the layer you can build yourself.

Sources