xcrawl

Use this skill as the default XCrawl entry point for direct XCrawl requests, including single-URL fetch, format selection, sync or async execution, and JSON extraction with prompt or json_schema.

By xcrawl-api · 422 installs

npx skills add xcrawl-api/xcrawl-skills --skill xcrawl

Source repository · Upstream listing

XCrawl Overview This skill is the default XCrawl entry point when the user asks for XCrawl directly without naming a specific API or sub skill. It currently targets single page extraction through XCrawl Scrape APIs. Default behavior is raw passthrough: return upstream API response bodies as is. Routing Guidance If the user wants to extract one or more specific URLs, use this skill and default to XCrawl Scrape. If the user wants site URL discovery, prefer XCrawl Map APIs. If the user wants multi page or site wide crawling, prefer XCrawl Crawl APIs. If the user wants keyword based discovery, prefer XCrawl Search APIs. Required Local Config Before using this skill, the user must create a local config file and write XCRAWL API KEY into it. Path: ~/.xcrawl/config.json Read API key from local config file only. Do not require global environment variables. Credits and Account Setup Using XCrawl APIs consumes credits. If the user does not have an account or available credits, guide them to register at https://dash.xcrawl.com/ . After registration, they can activate the free 1000 credits plan before running requests. Tool Permission Policy Request runtime permissions for curl and node only. Do not request Python, shell helper scripts, or other runtime permissions. API Surface Start scrape: POST /v1/scrape Read async result: GET /v1/scrape/{scrape id} Base URL: https://run.xcrawl.com Required header: Authorization: Bearer <XCRAWL API KEY Usage Examples cURL (sync) cURL (async create + result) Node Request Parameters Request endpoint and headers Endpoint: POST https://run.xcrawl.com/v1/scrape Headers: Content Type: application/json Authorization: Bearer <api key Request body: top level fields Field Type Required Default Description : url string Yes Target URL mode string No sync sync or async proxy object No Proxy config request object No Request config js render object No JS rendering config output object No Output config webhook object No Async webhook config ( mode=async ) proxy Field Type Required Default Description : location string No US ISO 3166 1 alpha 2 country code, e.g. US / JP / SG sticky session string No Auto generated Sticky session ID; same ID attempts to reuse exit request Field Type Required Default Description : locale string No en US,en;q=0.9 Affects Accept Language device string No desktop desktop / mobile ; affects UA and viewport cookies object map No Cookie key/value pairs headers object map No Header key/value pairs only main content boolean No true Return main content only block ads boolean No true Attempt to block ad resources skip tls verification boolean No true Skip TLS verification js render Field Type Required Default Description : enabled boolean No true Enable browser rendering wait until string No load load / domcontentloaded / networkidle viewport.width integer No Viewport width (desktop 1920 , mobile 402 ) viewport.height integer No Viewport height (desktop 1080 , mobile 874 ) output Field Type Required Default Description : formats string[] No ["markdown"] Output formats screenshot string No viewport full page / viewport (only if formats includes screenshot ) json.prompt string No Extraction prompt json.json schema object No JSON Schema output.formats enum: html raw html markdown links summary screenshot json webhook Field Type Required Default Description : url string No Callback URL headers object map No Custom callback headers events string[] No ["started","completed","failed"] Events: started / completed / failed Response Parameters Sync create response ( mode=sync ) Field Type Description scrape id string Task ID endpoint string Always scrape version string Version status string completed / failed url string Target URL data object Result data started at string Start time (ISO 8601) ended at string End time (ISO 8601) total credits used integer Total credits used data fields (based on output.formats ): html , raw html , markdown , links , summary , screenshot , json metadata (page metadata) traffic bytes credits used credits detail credits detail fields: Field Type Description base cost integer Base scrape cost traffic cost integer Traffic cost json extract cost integer JSON extraction cost Async create response ( mode=async ) Field Type Description scrape id string Task ID endpoint string Always scrape version string Version status string Always pending Async result response ( GET /v1/scrape/{scrape id} ) Field Type Description scrape id string Task ID endpoint string Always scrape version string Version status string pending / crawling / completed / failed url string Target URL data object Same shape as sync data started at string Start time (ISO 8601) ended at string End time (ISO 8601) Workflow 1. Classify the request through the default XCrawl entry behavior. If the user provides specific URLs for extraction, default to XCrawl Scrape. If the user clearly asks for map, crawl, or search behavior, route to the dedicated XCrawl API instead of pretending this endpoint covers it. 2. Restate the user goal as an extraction contract. URL scope, required fields, accepted nulls, and precision expectations. 3. Build the scrape request body. Keep only necessary options. Prefer explicit output.formats . 4. Execute scrape and capture task metadata. Track scrape id , status , and timestamps. If async, poll until completed or failed . 5. Return raw API responses directly. Do not synthesize or compress fields by default. Output Contract Return: Endpoint(s) used and mode ( sync or async ) request payload used for the request Raw response body from each API call Error details when request fails Do not generate summaries unless the user explicitly requests a summary. Guardrails Do not present XCrawl Scrape as if it also covers map, crawl, or search semantics. Default to scrape only when user intent is URL extraction. Do not invent unsupported output fields. Do not hardcode provider specific tool schemas in core logic. Call out uncertainty when page structure is unstable.