xcrawl
Use this skill as the default XCrawl entry point for direct XCrawl requests, including single-URL fetch, format selection, sync or async execution, and JSON extraction with prompt or json_schema.
By xcrawl-api · 422 installs
npx skills add xcrawl-api/xcrawl-skills --skill xcrawl
Source repository · Upstream listing
XCrawl
Overview
This skill is the default XCrawl entry point when the user asks for XCrawl directly without naming a specific API or sub skill.
It currently targets single page extraction through XCrawl Scrape APIs.
Default behavior is raw passthrough: return upstream API response bodies as is.
Routing Guidance
If the user wants to extract one or more specific URLs, use this skill and default to XCrawl Scrape.
If the user wants site URL discovery, prefer XCrawl Map APIs.
If the user wants multi page or site wide crawling, prefer XCrawl Crawl APIs.
If the user wants keyword based discovery, prefer XCrawl Search APIs.
Required Local Config
Before using this skill, the user must create a local config file and write XCRAWL API KEY into it.
Path: ~/.xcrawl/config.json
Read API key from local config file only. Do not require global environment variables.
Credits and Account Setup
Using XCrawl APIs consumes credits.
If the user does not have an account or available credits, guide them to register at https://dash.xcrawl.com/ .
After registration, they can activate the free 1000 credits plan before running requests.
Tool Permission Policy
Request runtime permissions for curl and node only.
Do not request Python, shell helper scripts, or other runtime permissions.
API Surface
Start scrape: POST /v1/scrape
Read async result: GET /v1/scrape/{scrape id}
Base URL: https://run.xcrawl.com
Required header: Authorization: Bearer <XCRAWL API KEY
Usage Examples
cURL (sync)
cURL (async create + result)
Node
Request Parameters
Request endpoint and headers
Endpoint: POST https://run.xcrawl.com/v1/scrape
Headers:
Content Type: application/json
Authorization: Bearer <api key
Request body: top level fields
Field Type Required Default Description
:
url string Yes Target URL
mode string No sync sync or async
proxy object No Proxy config
request object No Request config
js render object No JS rendering config
output object No Output config
webhook object No Async webhook config ( mode=async )
proxy
Field Type Required Default Description
:
location string No US ISO 3166 1 alpha 2 country code, e.g. US / JP / SG
sticky session string No Auto generated Sticky session ID; same ID attempts to reuse exit
request
Field Type Required Default Description
:
locale string No en US,en;q=0.9 Affects Accept Language
device string No desktop desktop / mobile ; affects UA and viewport
cookies object map No Cookie key/value pairs
headers object map No Header key/value pairs
only main content boolean No true Return main content only
block ads boolean No true Attempt to block ad resources
skip tls verification boolean No true Skip TLS verification
js render
Field Type Required Default Description
:
enabled boolean No true Enable browser rendering
wait until string No load load / domcontentloaded / networkidle
viewport.width integer No Viewport width (desktop 1920 , mobile 402 )
viewport.height integer No Viewport height (desktop 1080 , mobile 874 )
output
Field Type Required Default Description
:
formats string[] No ["markdown"] Output formats
screenshot string No viewport full page / viewport (only if formats includes screenshot )
json.prompt string No Extraction prompt
json.json schema object No JSON Schema
output.formats enum:
html
raw html
markdown
links
summary
screenshot
json
webhook
Field Type Required Default Description
:
url string No Callback URL
headers object map No Custom callback headers
events string[] No ["started","completed","failed"] Events: started / completed / failed
Response Parameters
Sync create response ( mode=sync )
Field Type Description
scrape id string Task ID
endpoint string Always scrape
version string Version
status string completed / failed
url string Target URL
data object Result data
started at string Start time (ISO 8601)
ended at string End time (ISO 8601)
total credits used integer Total credits used
data fields (based on output.formats ):
html , raw html , markdown , links , summary , screenshot , json
metadata (page metadata)
traffic bytes
credits used
credits detail
credits detail fields:
Field Type Description
base cost integer Base scrape cost
traffic cost integer Traffic cost
json extract cost integer JSON extraction cost
Async create response ( mode=async )
Field Type Description
scrape id string Task ID
endpoint string Always scrape
version string Version
status string Always pending
Async result response ( GET /v1/scrape/{scrape id} )
Field Type Description
scrape id string Task ID
endpoint string Always scrape
version string Version
status string pending / crawling / completed / failed
url string Target URL
data object Same shape as sync data
started at string Start time (ISO 8601)
ended at string End time (ISO 8601)
Workflow
1. Classify the request through the default XCrawl entry behavior.
If the user provides specific URLs for extraction, default to XCrawl Scrape.
If the user clearly asks for map, crawl, or search behavior, route to the dedicated XCrawl API instead of pretending this endpoint covers it.
2. Restate the user goal as an extraction contract.
URL scope, required fields, accepted nulls, and precision expectations.
3. Build the scrape request body.
Keep only necessary options.
Prefer explicit output.formats .
4. Execute scrape and capture task metadata.
Track scrape id , status , and timestamps.
If async, poll until completed or failed .
5. Return raw API responses directly.
Do not synthesize or compress fields by default.
Output Contract
Return:
Endpoint(s) used and mode ( sync or async )
request payload used for the request
Raw response body from each API call
Error details when request fails
Do not generate summaries unless the user explicitly requests a summary.
Guardrails
Do not present XCrawl Scrape as if it also covers map, crawl, or search semantics.
Default to scrape only when user intent is URL extraction.
Do not invent unsupported output fields.
Do not hardcode provider specific tool schemas in core logic.
Call out uncertainty when page structure is unstable.