Claude Skill

agent-browser

Browse the web for any task — research topics, read articles, interact with web apps, fill forms, take screenshots, extract data, and test web pages. Use whenever a browser would be useful, not just when the user explicitly asks.

LLM Mart · 0 points · 16 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download nanocoai-nanoclaw-container_skills_agent-browser-ad8837c.zip · 2 KB
Part of nanocoai/nanoclaw — 49 skills

Install

skills CLI npx skills add https://github.com/nanocoai/nanoclaw/tree/main/container/skills/agent-browser
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install nanocoai-nanoclaw@llmmart
Git git clone https://github.com/nanocoai/nanoclaw.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole nanocoai/nanoclaw collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Browser Automation with agent-browser

Quick start

agent-browser open <url>        # Navigate to page
agent-browser snapshot -i       # Get interactive elements with refs
agent-browser click @e1         # Click element by ref
agent-browser fill @e2 "text"   # Fill input by ref
agent-browser close             # Close browser

Core workflow

  1. Navigate: agent-browser open <url>
  2. Snapshot: agent-browser snapshot -i (returns elements with refs like @e1, @e2)
  3. Interact using refs from the snapshot
  4. Re-snapshot after navigation or significant DOM changes

Commands

Navigation

agent-browser open <url>      # Navigate to URL
agent-browser back            # Go back
agent-browser forward         # Go forward
agent-browser reload          # Reload page
agent-browser close           # Close browser

Snapshot (page analysis)

agent-browser snapshot            # Full accessibility tree
agent-browser snapshot -i         # Interactive elements only (recommended)
agent-browser snapshot -c         # Compact output
agent-browser snapshot -d 3       # Limit depth to 3
agent-browser snapshot -s "#main" # Scope to CSS selector

Interactions (use @refs from snapshot)

agent-browser click @e1           # Click
agent-browser dblclick @e1        # Double-click
agent-browser fill @e2 "text"     # Clear and type
agent-browser type @e2 "text"     # Type without clearing
agent-browser press Enter         # Press key
agent-browser hover @e1           # Hover
agent-browser check @e1           # Check checkbox
agent-browser uncheck @e1         # Uncheck checkbox
agent-browser select @e1 "value"  # Select dropdown option
agent-browser scroll down 500     # Scroll page
agent-browser upload @e1 file.pdf # Upload files

Get information

agent-browser get text @e1        # Get element text
agent-browser get html @e1        # Get innerHTML
agent-browser get value @e1       # Get input value
agent-browser get attr @e1 href   # Get attribute
agent-browser get title           # Get page title
agent-browser get url             # Get current URL
agent-browser get count ".item"   # Count matching elements

Screenshots & PDF

agent-browser screenshot          # Save to temp directory
agent-browser screenshot path.png # Save to specific path
agent-browser screenshot --full   # Full page
agent-browser pdf output.pdf      # Save as PDF

Wait

agent-browser wait @e1                     # Wait for element
agent-browser wait 2000                    # Wait milliseconds
agent-browser wait --text "Success"        # Wait for text
agent-browser wait --url "**/dashboard"    # Wait for URL pattern
agent-browser wait --load networkidle      # Wait for network idle

Waiting for a custom condition — ALWAYS bound it

Prefer the built-in wait subcommands above. Only fall back to eval-polling when you must wait on a custom JS condition (e.g. a spinner disappearing or a "Send" button re-enabling in a chat UI).

Never write an unbounded wait loop. A bare until … do sleep; done that polls a page condition will loop forever if the condition never becomes true (page failed to load, selector changed, network stalled). That does not just fail the command — it wedges the entire agent turn: the runner keeps the model stream open, later messages get silently swallowed, and the container can hang for hours without the host's stuck-detection firing.

Always cap the wait with BOTH a wall-clock timeout and a max-attempts counter, and always exit the loop (never leave a sleep loop as the last thing running):

# Bounded wait: succeeds when the condition is met, gives up after ~90s.
timeout 90 bash -c '
  for i in $(seq 1 30); do
    if agent-browser eval "document.querySelector(\".loading\") === null" 2>/dev/null | grep -q true; then
      echo READY; exit 0
    fi
    sleep 3
  done
  echo TIMEOUT; exit 1
'
# Check the exit status / output: on TIMEOUT, snapshot the page and decide —
# do NOT re-enter another unbounded wait.

If the wait times out, treat it as a real failure: take a snapshot -i or screenshot to see the actual page state, report what you found, and move on. Retrying the same unbounded wait is what causes the hang.

Semantic locators (alternative to refs)

agent-browser find role button click --name "Submit"
agent-browser find text "Sign In" click
agent-browser find label "Email" fill "user@test.com"
agent-browser find placeholder "Search" type "query"

Authentication with saved state

# Login once
agent-browser open https://app.example.com/login
agent-browser snapshot -i
agent-browser fill @e1 "username"
agent-browser fill @e2 "password"
agent-browser click @e3
agent-browser wait --url "**/dashboard"
agent-browser state save auth.json

# Later: load saved state
agent-browser state load auth.json
agent-browser open https://app.example.com/dashboard

Cookies & Storage

agent-browser cookies                     # Get all cookies
agent-browser cookies set name value      # Set cookie
agent-browser cookies clear               # Clear cookies
agent-browser storage local               # Get localStorage
agent-browser storage local set k v       # Set value

JavaScript

agent-browser eval "document.title"   # Run JavaScript

Example: Form submission

agent-browser open https://example.com/form
agent-browser snapshot -i
# Output shows: textbox "Email" [ref=e1], textbox "Password" [ref=e2], button "Submit" [ref=e3]

agent-browser fill @e1 "user@example.com"
agent-browser fill @e2 "password123"
agent-browser click @e3
agent-browser wait --load networkidle
agent-browser snapshot -i  # Check result

Example: Data extraction

agent-browser open https://example.com/products
agent-browser snapshot -i
agent-browser get text @e1  # Get product title
agent-browser get attr @e2 href  # Get link URL
agent-browser screenshot products.png
Files (nanoclaw)
  • SKILL.md 6.2 KB
    ---
    name: agent-browser
    description: Browse the web for any task — research topics, read articles, interact with web apps, fill forms, take screenshots, extract data, and test web pages. Use whenever a browser would be useful, not just when the user explicitly asks.
    allowed-tools: Bash(agent-browser:*)
    ---
    
    # Browser Automation with agent-browser
    
    ## Quick start
    
    ```bash
    agent-browser open <url>        # Navigate to page
    agent-browser snapshot -i       # Get interactive elements with refs
    agent-browser click @e1         # Click element by ref
    agent-browser fill @e2 "text"   # Fill input by ref
    agent-browser close             # Close browser
    ```
    
    ## Core workflow
    
    1. Navigate: `agent-browser open <url>`
    2. Snapshot: `agent-browser snapshot -i` (returns elements with refs like `@e1`, `@e2`)
    3. Interact using refs from the snapshot
    4. Re-snapshot after navigation or significant DOM changes
    
    ## Commands
    
    ### Navigation
    
    ```bash
    agent-browser open <url>      # Navigate to URL
    agent-browser back            # Go back
    agent-browser forward         # Go forward
    agent-browser reload          # Reload page
    agent-browser close           # Close browser
    ```
    
    ### Snapshot (page analysis)
    
    ```bash
    agent-browser snapshot            # Full accessibility tree
    agent-browser snapshot -i         # Interactive elements only (recommended)
    agent-browser snapshot -c         # Compact output
    agent-browser snapshot -d 3       # Limit depth to 3
    agent-browser snapshot -s "#main" # Scope to CSS selector
    ```
    
    ### Interactions (use @refs from snapshot)
    
    ```bash
    agent-browser click @e1           # Click
    agent-browser dblclick @e1        # Double-click
    agent-browser fill @e2 "text"     # Clear and type
    agent-browser type @e2 "text"     # Type without clearing
    agent-browser press Enter         # Press key
    agent-browser hover @e1           # Hover
    agent-browser check @e1           # Check checkbox
    agent-browser uncheck @e1         # Uncheck checkbox
    agent-browser select @e1 "value"  # Select dropdown option
    agent-browser scroll down 500     # Scroll page
    agent-browser upload @e1 file.pdf # Upload files
    ```
    
    ### Get information
    
    ```bash
    agent-browser get text @e1        # Get element text
    agent-browser get html @e1        # Get innerHTML
    agent-browser get value @e1       # Get input value
    agent-browser get attr @e1 href   # Get attribute
    agent-browser get title           # Get page title
    agent-browser get url             # Get current URL
    agent-browser get count ".item"   # Count matching elements
    ```
    
    ### Screenshots & PDF
    
    ```bash
    agent-browser screenshot          # Save to temp directory
    agent-browser screenshot path.png # Save to specific path
    agent-browser screenshot --full   # Full page
    agent-browser pdf output.pdf      # Save as PDF
    ```
    
    ### Wait
    
    ```bash
    agent-browser wait @e1                     # Wait for element
    agent-browser wait 2000                    # Wait milliseconds
    agent-browser wait --text "Success"        # Wait for text
    agent-browser wait --url "**/dashboard"    # Wait for URL pattern
    agent-browser wait --load networkidle      # Wait for network idle
    ```
    
    ### Waiting for a custom condition — ALWAYS bound it
    
    Prefer the built-in `wait` subcommands above. Only fall back to `eval`-polling
    when you must wait on a custom JS condition (e.g. a spinner disappearing or a
    "Send" button re-enabling in a chat UI).
    
    **Never write an unbounded wait loop.** A bare `until … do sleep; done` that
    polls a page condition will loop *forever* if the condition never becomes true
    (page failed to load, selector changed, network stalled). That does not just
    fail the command — it wedges the entire agent turn: the runner keeps the model
    stream open, later messages get silently swallowed, and the container can hang
    for hours without the host's stuck-detection firing.
    
    Always cap the wait with BOTH a wall-clock `timeout` and a max-attempts counter,
    and always exit the loop (never leave a `sleep` loop as the last thing running):
    
    ```bash
    # Bounded wait: succeeds when the condition is met, gives up after ~90s.
    timeout 90 bash -c '
      for i in $(seq 1 30); do
        if agent-browser eval "document.querySelector(\".loading\") === null" 2>/dev/null | grep -q true; then
          echo READY; exit 0
        fi
        sleep 3
      done
      echo TIMEOUT; exit 1
    '
    # Check the exit status / output: on TIMEOUT, snapshot the page and decide —
    # do NOT re-enter another unbounded wait.
    ```
    
    If the wait times out, treat it as a real failure: take a `snapshot -i` or
    `screenshot` to see the actual page state, report what you found, and move on.
    Retrying the same unbounded wait is what causes the hang.
    
    ### Semantic locators (alternative to refs)
    
    ```bash
    agent-browser find role button click --name "Submit"
    agent-browser find text "Sign In" click
    agent-browser find label "Email" fill "user@test.com"
    agent-browser find placeholder "Search" type "query"
    ```
    
    ### Authentication with saved state
    
    ```bash
    # Login once
    agent-browser open https://app.example.com/login
    agent-browser snapshot -i
    agent-browser fill @e1 "username"
    agent-browser fill @e2 "password"
    agent-browser click @e3
    agent-browser wait --url "**/dashboard"
    agent-browser state save auth.json
    
    # Later: load saved state
    agent-browser state load auth.json
    agent-browser open https://app.example.com/dashboard
    ```
    
    ### Cookies & Storage
    
    ```bash
    agent-browser cookies                     # Get all cookies
    agent-browser cookies set name value      # Set cookie
    agent-browser cookies clear               # Clear cookies
    agent-browser storage local               # Get localStorage
    agent-browser storage local set k v       # Set value
    ```
    
    ### JavaScript
    
    ```bash
    agent-browser eval "document.title"   # Run JavaScript
    ```
    
    ## Example: Form submission
    
    ```bash
    agent-browser open https://example.com/form
    agent-browser snapshot -i
    # Output shows: textbox "Email" [ref=e1], textbox "Password" [ref=e2], button "Submit" [ref=e3]
    
    agent-browser fill @e1 "user@example.com"
    agent-browser fill @e2 "password123"
    agent-browser click @e3
    agent-browser wait --load networkidle
    agent-browser snapshot -i  # Check result
    ```
    
    ## Example: Data extraction
    
    ```bash
    agent-browser open https://example.com/products
    agent-browser snapshot -i
    agent-browser get text @e1  # Get product title
    agent-browser get attr @e2 href  # Get link URL
    agent-browser screenshot products.png
    ```
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related