If you're using an AI coding agent to drive a browser, the Playwright CLI is where I'd start today. It does the same core jobs as the Playwright MCP server, and it keeps your agent's context window lean while doing it.
I've been digging into both this month for browser automation and testing work. Here's what I found:
- What each tool actually is
- Why the CLI uses fewer tokens (and an honest update on how big that gap is now)
- Headed vs. headless mode
- How to install it in about two minutes
- When MCP is still the better call
Two ways to hand an agent a browser
Microsoft ships both tools. They sit on the same Playwright engine, but they plug into your agent differently.
Playwright MCP is a Model Context Protocol server (@playwright/mcp). Your agent gets a set of browser tools like navigate, click, and fill form, each with its own schema loaded into context.
Playwright CLI (@playwright/cli) is a plain command-line tool. Your agent runs shell commands like playwright-cli click e15 and reads short results back. It ships with a "skill" file that teaches agents like Claude Code or GitHub Copilot how to use it.
The Playwright team's own guidance is clear: if you're working with a coding agent, the CLI is the better fit. MCP is aimed at long-running autonomous loops that need persistent browser state and deep page introspection.
Why the CLI uses fewer tokens
Both tools read the page the same way: through its accessibility tree, a structured list of every button, link, and field with a short ref like e21. That's what lets an agent click the right thing without guessing at pixels. The difference is where that tree goes.
- MCP (historically): sent the full accessibility snapshot back into the agent's context after almost every action. A content-heavy page could cost thousands of tokens per click.
- CLI: saves the snapshot to a YAML file on disk and returns a one-line summary. The agent only reads the slice it needs, or searches it with playwright-cli find "Add to cart".
- Startup cost: MCP loads every tool definition into context up front. The CLI loads a short skill description and discovers the rest through playwright-cli --help.
| Measurement | Playwright MCP | Playwright CLI |
| --- | --- | --- |
| Multi-step UI test, early 2026 benchmark | ~114,000 tokens | ~27,000 tokens |
| Tool/skill cost loaded up front | ~3,600 tokens | ~68 tokens |
| Same test, re-run July 2026 (MCP 0.0.78 vs CLI 0.1.17) | ~3,580 tokens | ~3,620 tokens |
The honest update: newer MCP versions also write snapshots to disk now, so the per-step gap has largely closed. The CLI still wins on fixed overhead and on flexibility, since it's just shell commands your agent can chain, script, and drop into CI. For me, that's the real reason to switch.
Sources: Better Stack, TestCollab, DEV Community re-benchmark
Headed vs. headless: watch your agent work
The CLI runs headless by default, meaning the browser works invisibly in the background. That's what you want for CI and speed.
Add --headed and a real browser window opens so you can watch every click and keystroke:
bash
playwright-cli open https://playwright.dev --headed
This is my favorite part. When you're building a new automation, seeing it happen live beats reading logs. You catch the modal that blocked a click, or the page that loaded slower than expected.
There's also playwright-cli show, a dashboard with live previews of every running browser session. You can click into one and take over the mouse and keyboard if your agent gets stuck.
Install it in two minutes
You need Node.js 18 or newer and a coding agent (Claude Code, GitHub Copilot, or similar).
- Install the CLI globally
bash
npm install -g @playwright/cli@latest
- Install the skills so your agent knows how to use it
bash
playwright-cli install --skills
- Pick a browser if you want something specific, e.g. playwright-cli open --browser=chrome, or emulate a phone with --device="iPhone 15".
- Try it by hand
bash
playwright-cli open https://demo.playwright.dev/todomvc/ --headed
playwright-cli type "Buy groceries"
playwright-cli press Enter
playwright-cli screenshot
Or skip the typing and just tell your agent: "Use playwright-cli to test the add-todo flow on this site. Check playwright-cli --help for commands." Everything else lives in the repo: github.com/microsoft/playwright-cli.
When to use which
Reach for the CLI when:
Your agent has shell access (Claude Code, Copilot, Codex)
It's juggling browser work alongside a real codebase and needs context room
You want scriptable, repeatable runs you can move into CI
You want to watch it headed, record video, or generate Playwright test code
Stick with MCP when:Your agent is sandboxed with no shell or filesystem access (e.g. a desktop chat client)
You're running long, exploratory loops where continuous browser state matters more than tokens
For most developers building automations or tests with a coding agent, that points to the CLI.
What I'm building next
I'm testing this right now on Snapie (snapie.io). The goal: a script that uses the Playwright CLI to work through the first 10 posts on the platform, running headed so I can watch every step.
I'll record the demo and share it in a follow-up post, including what worked, what broke, and how many tokens it actually used.
If you're doing browser automation or testing with AI agents, install the CLI this week and try one small workflow. Then tell me what you built. I'd love to compare notes.