canghe-url-to-markdown

Fetch any URL and convert to markdown using Chrome CDP. Supports two modes - auto-capture on page load, or wait for user signal (for pages requiring login). Use when user wants to save a webpage as markdown.

By freestylefly · 499 installs

npx skills add freestylefly/canghe-skills --skill canghe-url-to-markdown

Source repository · Upstream listing

URL to Markdown Fetches any URL via Chrome CDP and converts HTML to clean markdown. Script Directory Important : All scripts are located in the scripts/ subdirectory of this skill. Agent Execution Instructions : 1. Determine this SKILL.md file's directory path as SKILL DIR 2. Script path = ${SKILL DIR}/scripts/<script name .ts 3. Replace all ${SKILL DIR} in this document with the actual path Script Reference : Script Purpose scripts/main.ts CLI entry point for URL fetching Preferences (EXTEND.md) Use Bash to check EXTEND.md existence (priority order): ┌────────────────────────────────────────────────────────┬───────────────────┐ │ Path │ Location │ ├────────────────────────────────────────────────────────┼───────────────────┤ │ .canghe skills/canghe url to markdown/EXTEND.md │ Project directory │ ├────────────────────────────────────────────────────────┼───────────────────┤ │ $HOME/.canghe skills/canghe url to markdown/EXTEND.md │ User home │ └────────────────────────────────────────────────────────┴───────────────────┘ ┌───────────┬───────────────────────────────────────────────────────────────────────────┐ │ Result │ Action │ ├───────────┼───────────────────────────────────────────────────────────────────────────┤ │ Found │ Read, parse, apply settings │ ├───────────┼───────────────────────────────────────────────────────────────────────────┤ │ Not found │ Use defaults │ └───────────┴───────────────────────────────────────────────────────────────────────────┘ EXTEND.md Supports : Default output directory Default capture mode Timeout settings Features Chrome CDP for full JavaScript rendering Two capture modes: auto or wait for user Clean markdown output with metadata Handles login required pages via wait mode Usage Options Option Description <url URL to fetch o <path Output file path (default: auto generated) wait Wait for user signal before capturing timeout <ms Page load timeout (default: 30000) Capture Modes Mode Behavior Use When Auto (default) Capture on network idle Public pages, static content Wait ( wait ) User signals when ready Login required, lazy loading, paywalls Wait mode workflow : 1. Run with wait → script outputs "Press Enter when ready" 2. Ask user to confirm page is ready 3. Send newline to stdin to trigger capture Output Format YAML front matter with url , title , description , author , published , captured at fields, followed by converted markdown content. Output Directory <slug : From page title or URL path (kebab case, 2 6 words) Conflict resolution: Append timestamp <slug YYYYMMDD HHMMSS.md Environment Variables Variable Description URL CHROME PATH Custom Chrome executable path URL DATA DIR Custom data directory URL CHROME PROFILE DIR Custom Chrome profile directory Troubleshooting : Chrome not found → set URL CHROME PATH . Timeout → increase timeout . Complex pages → try wait mode. Extension Support Custom configurations via EXTEND.md. See Preferences section for paths and supported options.