AIHawk: Web Automation MCP Client for Claude Code & Gemini CLI
AIHawk: Web Automation MCP Client for Claude Code & Gemini CLI
AIHawk is an open-source AI browser agent designed to automate web interactions using plain English commands. It functions as a web browsing and computer-use agent, capable of navigating, clicking, typing, and reading the actual web to accomplish tasks. With over 30,000 GitHub stars, AIHawk has been featured in publications like Business Insider and TechCrunch for its capabilities in areas such as automating job applications.
Integrating AIHawk as an MCP Client
The core utility of AIHawk for developers leveraging MCP lies in its direct integration with AI assistants like Claude Code and Gemini CLI. This allows you to expose AIHawk's web automation capabilities directly to your assistant, turning natural language instructions into real browser actions.
To set up AIHawk for use with your assistant over MCP, first install the uvx invisible-playwright dependencies.
For Windows (PowerShell):
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
$env:Path = "$env:USERPROFILE\.local\bin;$env:Path"
uvx invisible-playwright fetchFor Linux:
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env
uvx invisible-playwright fetchOnce uvx invisible-playwright is fetched, you can add AIHawk as an available tool to your chosen assistant. This registers uvx aihawk as a command your assistant can invoke via MCP.
For Claude Code:
claude mcp add --scope user stealth -- uvx aihawkFor Codex:
codex mcp add stealth -- uvx aihawkFor Gemini CLI:
gemini mcp add --scope user stealth uvx aihawkAfter executing these commands, you can instruct your assistant in plain English, and it will be able to utilize AIHawk to perform web-based tasks. The --scope user stealth flag indicates how the tool should be exposed and invoked by the assistant.
Web Browsing and Computer-Use Agent Capabilities
AIHawk operates as a browser agent with a real browser instance. Its primary function is to interpret plain language instructions and translate them into actions on the web. This includes:
- Web Browsing: Navigating to specified URLs.
- Clicking: Interacting with elements on a webpage.
- Typing: Inputting text into forms or search bars.
- Reading: Extracting information from web pages.
This enables automation of complex, multi-step web processes that would typically require manual interaction. The agent's ability to act as a general computer-use agent suggests broader interaction capabilities beyond just the browser, though the source material emphasizes its web browsing features.