Nuphus Mcp

Desktop automation MCP server — computer use for any AI agent: control screen, windows, mouse/keyboard, and Chrome via Model Context Protocol (stdio)

LLM Mart
17 views 156 listing impressions

Desktop automation MCP server — computer use for any AI agent. See the screen, control windows/mouse/keyboard, and drive Chrome over the Model Context Protocol (stdio). Desktop & browser automation need no API key; OCR runs locally; vision plugs into your own vision LLM (OpenAI-compatible or Anthropic native, BYOK).

nuphus-mcp is a lightweight, cross-platform desktop automation MCP server that exposes desktop + browser automation as standard MCP tools. It speaks JSON-RPC 2.0 over stdio — no daemon, no network service, one binary. Claude Desktop, Cursor, VS Code, Copilot, or any MCP client can connect and immediately control the screen, windows, keyboard/mouse, and Chrome — computer use for any AI agent — desktop & browser automation need no API key; local OCR is built in; vision works with your own vision LLM (OpenAI-compatible, BYOK).

🇨🇳 Mainland China mirror: this repo is mirrored on Gitee for fast in-China access (Chinese docs served by default there). 中文文档

Part of the Nuphus ecosystem: Nuphus — a local-first AI agent with real desktop execution and dual-device (phone-as-second-screen) sync. nuphus-mcp exposes the same desktop + browser automation to any MCP client.

Using DeepSeek Harness (DSH)? nuphus-mcp itself is a plain stdio MCP server — not a DSH/Cordis plugin. Install the dedicated dsh-nuphus-mcp plugin instead (see DeepSeek Harness (DSH)).

┌──────────────────┐   stdio JSON-RPC   ┌──────────────────────┐
│  Any MCP Client  │  ───────────────►  │      nuphus-mcp      │
│  (Claude/Cursor/ │  ◄───────────────  │  desktop-api crate   │──► screen/window/mouse/keyboard
│   Nuphus itself) │  single-line JSON  │  nuphus-browser crate│──► Chrome (CDP)
└──────────────────┘                    └──────────────────────┘

Features

  • 38 MCP tools (15 desktop + 23 browser) — screenshots, window control, mouse/keyboard, Chrome CDP automation, and more — see TOOLS.md / TOOLS.zh-CN.md for the full reference.
  • Desktop automation: screen size, screenshot (PNG/base64), window list, window activate/screenshot/move/resize/info, mouse click/drag/scroll/position, keyboard input/hotkey, clipboard write/clean — implemented on the desktop-api crate (xcap + Win32, no Tauri dependency).
  • Computer vision pair: desktop_vision (BYOK — send a screenshot to your own vision model via an OpenAI-compatible or Anthropic native API) + desktop_perceive (local OCR with PaddleOCR, models auto-downloaded on first run; optional YOLO icon detection). Used together they give AI agents both semantic understanding and pixel-precise coordinates — the battle-tested vision→perceive flow from the Nuphus desktop app. See TOOLS.md for BYOK env vars, model setup, and the recommended flow.
  • Browser automation: navigate, snapshot (accessibility tree with @N refs), click, type, exec, scroll, extract, screenshot, evaluate, back/forward, wait_for, cookies get/set/import, upload, tabs, downloads — implemented on nuphus-browser (chromiumoxide CDP).
  • Zero-cost stdio: no HTTP server, no daemon. The process reads single-line JSON from stdin and writes responses to stdout.
  • Safety-first: destructive tools are annotated per the MCP spec; optional strict-confirm mode; path validation for screenshots, uploads, and file drags.

Repository Layout

nuphus-mcp/
├── Cargo.toml                  # workspace root
├── TOOLS.md / TOOLS.zh-CN.md   # 38-tool reference
├── crates/
│   ├── nuphus-mcp/             # MCP Server (this repo's product)
│   ├── nuphus-browser/         # Browser automation core (CDP)
│   └── desktop-api/            # Desktop control core (vendored)
└── ...

Prerequisites

From the project's README.

Comments (0)

Sign in to join the conversation.

No comments yet.