Can You Build Your Own Meta Muse With Claude Sonnet 5.5 and MCP?
Short answer: Claude Sonnet 5.5 plus MCP servers covers the model and the tool layer of Meta Muse today. The runtime around them (policy gate, token vault, memory, job queue, audit log, injection defenses and a browser sandbox) is still yours to write. This post lists each part and shows the loop that ties them together.
Why should builders care about this launch week?
Two launches landed three weeks apart. On September 8, 2026 Meta announced Muse, a personal agent that browses, sends email, shops and keeps working after you close the app. On September 28 Anthropic shipped Claude Sonnet 5.5 at $2 per million input tokens and $10 per million output tokens, with an 80.1% score on the OSWorld 2.1 computer-use benchmark.
If you build products, the useful question is how much of Muse you could ship inside your own app next month. I split Muse into the parts Meta actually shipped, marked the ones Claude and MCP cover today, and sketched the rest as code.
What did Meta actually ship?
- Muse Spark, the model.
- Muse Secure VM: one isolated virtual machine per user, with its own browser and storage.
- Sentinel: a separate policy agent that allows, denies or escalates each connector request and outbound call.
- Approval prompts before email or purchases, and a full audit trail.
- Memory of routines and one-off details, used for proactive suggestions.
- Background execution that continues after the app is closed.
- Payments through Stripe's Link with single-use card numbers, with Shop Pay and 1Password announced.
- Surfaces: iOS, Android, the web and WhatsApp, rolling out in the US.
Only the first item is a model. The other seven are infrastructure.
What does Claude Sonnet 5.5 bring to an agent?
Sonnet 5.5 is the planner in this design. The numbers that matter for an agent are these: output about 30% faster than Sonnet 5, up to 30% lower cost per task because it uses fewer tokens and tool calls, 80.1% on OSWorld 2.1 against 81.8% for Opus 5.5, adjustable effort levels, and availability on the Claude API, AWS, Google Cloud and Microsoft Azure. Lovable reported roughly one third fewer tool calls after switching.
For an agent that makes dozens of tool calls per task, the drop in calls matters more than a benchmark point. Every call you avoid is one less network round trip to an MCP server and one less place for the run to fail.
What does MCP cover out of the box?
The Model Context Protocol is the open standard Anthropic created in 2024 and handed to the Linux Foundation's Agentic AI Foundation in December 2025. It gives the agent one way to reach every tool. A starter set for a Muse-like assistant looks like this:
| Need | MCP server | Status |
|---|---|---|
| Gmail connector or remote server | Ready | |
| Calendar | Google Calendar | Ready |
| Documents | Google Drive | Ready |
| Browser | @playwright/mcp | Ready, you host it |
| Code | github-mcp-server | Ready |
| Data and memory rows | Postgres server | Ready |
| Internal APIs | Small custom server with the TypeScript or Python SDK | You write it |
Which seven parts do you still write?
- Policy gate. Tag every tool as read, write or money. Allow reads, send writes to an approval queue, block anything unlisted. A
tool_policies(tool, risk, requires_approval)table is enough to start. - Token vault. Encrypt refresh tokens per user with a KMS or libsodium, refresh them in a worker, and never put them in the prompt.
- Memory. Postgres with pgvector, a
memoriestable that stores the source and the timestamp of each fact, and a settings page where users can read and delete entries. - Jobs. A queue such as BullMQ on Redis or a managed queue such as SQS, plus a scheduler for requests like "in five days".
- Audit log. An append-only
agent_eventstable with task, tool, input hash, result summary, approver and time. - Injection defenses. Mark tool results that come from the web as untrusted. After the agent reads untrusted content, allow only read tools until the user approves the next write.
- Browser sandbox. One container per user session running Playwright, with a live view and a kill switch.
What does a minimal agent loop look like?
This is the core of the runtime in TypeScript. The model and the MCP servers do most of the work. The loop adds the policy check, the log and the pause.
import Anthropic from "@anthropic-ai/sdk";
const claude = new Anthropic();
export async function runTask(task: Task) {
const messages = await memory.load(task.userId, task.goal);
const tools = await mcp.listTools(task.userId); // tools from every connected MCP server
for (let step = 0; step < 40; step++) {
const res = await claude.messages.create({
model: "claude-sonnet-5-5",
max_tokens: 4096,
tools,
messages,
});
messages.push({ role: "assistant", content: res.content });
if (res.stop_reason !== "tool_use") return finish(task, res);
const results = [];
for (const call of res.content.filter((b) => b.type === "tool_use")) {
const verdict = policy.check(task.userId, call); // "allow" | "ask" | "deny"
if (verdict === "ask") return queue.pauseForApproval(task, messages, call);
const output = verdict === "allow"
? await mcp.callTool(task.userId, call.name, call.input)
: "Blocked by policy.";
await audit.log({ task: task.id, tool: call.name, verdict, output });
results.push({ type: "tool_result", tool_use_id: call.id, content: output });
}
messages.push({ role: "user", content: results });
await memory.checkpoint(task.id, messages);
}
}
pauseForApproval saves the conversation and the pending call. When the user taps approve, a worker runs that call, appends the result and starts runTask again from the checkpoint. That small function is how "keeps working after you close the app" happens.
How does the job-hunt task look in the audit log?
Here is the example request "find remote React jobs, check my calendar, prepare applications, follow up in five days" as the log would record it:
09:02 plan search remote React roles, posted in last 7 days
09:02 browser navigate jobs board allow 42 listings (untrusted)
09:03 postgres read prefs allow remote, React, salary floor
09:03 calendar find_free_slots(14d) allow 6 slots
09:05 drive create_doc x5 allow 5 cover letters
09:05 gmail create_draft x5 ask waiting for user
09:41 approval user approved 4, rejected 1
09:42 browser submit application x4 allow 4 submitted
09:42 scheduler follow_up in 5 days allow job #8812
day 5 gmail search replies allow 1 reply, 3 silent
day 5 gmail create_draft x3 ask waiting for user
Each line marked ask is a point where Muse's Sentinel would also pause. Each line with a job number is something a chat window cannot do on its own.
Should you use Muse or your own stack?
Use Muse if you are one person who wants results and you are fine with Meta's data terms, its US-only rollout and its current integration list.
Build on Claude and MCP if the agent has to live inside your product, reach internal systems, follow your retention rules, or switch models later without rewriting every integration.
Verdict
Sonnet 5.5 gives you a planner that sits close to the top model at a mid-tier price. MCP turns the tool layer into configuration. Muse's real lead is its runtime: the VM, the policy agent, the vault, the scheduler and the log. That layer is where agent products will compete next, and it is also the layer you can build yourself.