Browser automation#Browser automation
Playwright MCP: browser control through page structure
Microsoft's official browser MCP server: agents act on the page's accessibility tree instead of screenshots, so a plain text model can drive the web reliably.
Project and installation docs
View projecthttps://github.com/microsoft/playwright-mcp
The usual way to let an agent use a website is to screenshot it and ask a vision model where the button is. That’s slow and misclicks are common. Playwright MCP hands the model the page’s accessibility tree instead, and the model acts on element references, with no vision needed. It’s maintained by Microsoft’s Playwright team and had about 38k stars as of 2026-10-06.
What it does
- Structured snapshots:
browser_snapshotreturns the accessibility tree, andbrowser_clickorbrowser_typeact on the element references in it. - The full set of browser actions: navigation, back and forward, forms, dropdowns, file uploads, dialogs and tabs.
- Opt-in extras:
--capsturns on coordinate clicks (vision), PDF output, DevTools, network and storage tools. They’re off by default so they don’t eat context. - Three session modes: a persistent profile by default,
--isolatedfor a throwaway in-memory profile, or--extensionto attach to your signed-in Chrome or Edge. - Browser and device choice:
--browserpicks Chrome, Firefox, WebKit or the Edge channel, and--deviceor--mobileemulate a specific or generic mobile device, which also trims token use since mobile pages are lighter.
Who it’s for
- Developers who want Claude Code, Codex or Cursor to open a page, check it, or walk through a sign-up flow without screenshotting and guessing where the button is.
- Teams running exploratory or self-healing tests, or long autonomous web tasks where browser state has to persist.
Setup
Needs Node.js 18 or newer. In Claude Code it’s one command:
claude mcp add playwright npx @playwright/mcp@latest
A Docker image is also available, though it currently supports only headless Chromium:
{
"mcpServers": {
"playwright": {
"command": "docker",
"args": [
"run",
"-i",
"--rm",
"--init",
"--pull=always",
"mcr.microsoft.com/playwright/mcp"
]
}
}
}
Other clients use the standard config:
{
"mcpServers": {
"playwright": {
"command": "npx",
"args": ["@playwright/mcp@latest"]
}
}
}
Our take
There are plenty of browser MCP servers, and this is the baseline: officially maintained, frequently updated, with clear tool descriptions, and acting on element references is far steadier than clicking coordinates. If you also need to drive mobile apps or Electron rather than just the browser, WebdriverIO MCP covers that in one server. Two caveats. The project itself says coding agents short on context may do better with the team’s Playwright CLI plus Skills, which uses fewer tokens; the MCP server suits long tasks that need browser state. And the default profile saves sign-ins to disk, so two clients sharing it will collide; use --isolated for testing. The project is explicit that it is not a security boundary, so treat any page it visits as untrusted input. Licensed Apache-2.0.