canghe-url-to-markdown
Fetch any URL and convert to markdown using Chrome CDP. Supports two modes - auto-capture on page load, or wait for user signal (for pages requiring login). Use when user wants to save a webpage as markdown.
By freestylefly · 499 installs
npx skills add freestylefly/canghe-skills --skill canghe-url-to-markdown
Source repository · Upstream listing
URL to Markdown
Fetches any URL via Chrome CDP and converts HTML to clean markdown.
Script Directory
Important : All scripts are located in the scripts/ subdirectory of this skill.
Agent Execution Instructions :
1. Determine this SKILL.md file's directory path as SKILL DIR
2. Script path = ${SKILL DIR}/scripts/<script name .ts
3. Replace all ${SKILL DIR} in this document with the actual path
Script Reference :
Script Purpose
scripts/main.ts CLI entry point for URL fetching
Preferences (EXTEND.md)
Use Bash to check EXTEND.md existence (priority order):
┌────────────────────────────────────────────────────────┬───────────────────┐
│ Path │ Location │
├────────────────────────────────────────────────────────┼───────────────────┤
│ .canghe skills/canghe url to markdown/EXTEND.md │ Project directory │
├────────────────────────────────────────────────────────┼───────────────────┤
│ $HOME/.canghe skills/canghe url to markdown/EXTEND.md │ User home │
└────────────────────────────────────────────────────────┴───────────────────┘
┌───────────┬───────────────────────────────────────────────────────────────────────────┐
│ Result │ Action │
├───────────┼───────────────────────────────────────────────────────────────────────────┤
│ Found │ Read, parse, apply settings │
├───────────┼───────────────────────────────────────────────────────────────────────────┤
│ Not found │ Use defaults │
└───────────┴───────────────────────────────────────────────────────────────────────────┘
EXTEND.md Supports : Default output directory Default capture mode Timeout settings
Features
Chrome CDP for full JavaScript rendering
Two capture modes: auto or wait for user
Clean markdown output with metadata
Handles login required pages via wait mode
Usage
Options
Option Description
<url URL to fetch
o <path Output file path (default: auto generated)
wait Wait for user signal before capturing
timeout <ms Page load timeout (default: 30000)
Capture Modes
Mode Behavior Use When
Auto (default) Capture on network idle Public pages, static content
Wait ( wait ) User signals when ready Login required, lazy loading, paywalls
Wait mode workflow :
1. Run with wait → script outputs "Press Enter when ready"
2. Ask user to confirm page is ready
3. Send newline to stdin to trigger capture
Output Format
YAML front matter with url , title , description , author , published , captured at fields, followed by converted markdown content.
Output Directory
<slug : From page title or URL path (kebab case, 2 6 words)
Conflict resolution: Append timestamp <slug YYYYMMDD HHMMSS.md
Environment Variables
Variable Description
URL CHROME PATH Custom Chrome executable path
URL DATA DIR Custom data directory
URL CHROME PROFILE DIR Custom Chrome profile directory
Troubleshooting : Chrome not found → set URL CHROME PATH . Timeout → increase timeout . Complex pages → try wait mode.
Extension Support
Custom configurations via EXTEND.md. See Preferences section for paths and supported options.