exa-contents

Call Exa Contents directly with cURL or raw HTTP. Use when an agent already has URLs and needs POST /contents without an SDK for extracted text, highlights, summaries, links, image links, subpages, freshness-controlled crawling, or per-URL status handling.

By exa-labs · 552 installs

npx skills add exa-labs/agent-skills --skill exa-contents

Source repository · Upstream listing

Exa Contents Requires API key: Get one at https://dashboard.exa.ai/api keys Header: x api key: $EXA API KEY Use POST https://api.exa.ai/contents when the agent already knows the URLs and needs clean, LLM ready extraction without running a new search. Start with one content mode: highlights for compact agent context, text for broad page context, or summary for Exa side compression. Quick Start (cURL) Basic text extraction Highlights with freshness control Endpoint Authentication: x api key: <API KEY header. Exa also accepts Authorization: Bearer <API KEY , but prefer x api key in cURL examples for consistency. Use this endpoint for known URL extraction. If the agent needs discovery or ranking first, use POST /search . Parameters Core request parameters Parameter Type Required Default Description urls string[] Yes URLs to extract content from. Use this for known URLs. text boolean or object No Return full page text as markdown. Object form supports maxCharacters , includeHtmlTags , verbosity , includeSections , and excludeSections . highlights boolean or object No Return key excerpts. Prefer true for agent workflows unless a custom focus is needed. summary boolean or object No Return per page LLM summaries. Use when the caller wants Exa side compression or structured extraction. maxAgeHours integer No Freshness control. 0 always live crawls; 1 uses cache only; omit for default cache first behavior with crawl fallback. livecrawlTimeout integer No 10000 Timeout for live crawling in milliseconds. Use 10000 to 15000 for most freshness sensitive calls. subpages integer No 0 Number of linked subpages to crawl from each URL. subpageTarget string or string[] No Terms used to prioritize which subpages matter, such as ["api", "reference", "pricing"] . extras.links integer No 0 Number of links to extract from each page. extras.imageLinks integer No 0 Number of image URLs to extract from each page. compliance string No Enterprise only compliance mode, such as hipaa , when enabled for the account. Text object options Parameter Type Default Description maxCharacters integer Character limit for returned text. Use this instead of tokensNum . includeHtmlTags boolean false Preserve HTML tags in output. verbosity string compact compact , standard , or full . Pair fresh section aware extraction with maxAgeHours: 0 . includeSections string[] Only include selected sections: header , navigation , banner , body , sidebar , footer , metadata . excludeSections string[] Exclude selected sections from the same section list. Highlights object options Prefer highlights: true for the highest quality default. Only use object form when the agent needs a custom focus or budget. Parameter Type Default Description query string Custom query guiding which excerpts are returned. maxCharacters integer Cap highlight characters per URL. Omit unless the caller has a strict budget. Summary object options Parameter Type Default Description query string Custom query for the summary. schema object JSON Schema for structured per page summaries. Content Modes On /contents , text , highlights , and summary are top level request fields. Mode Best for Notes text Deep analysis and broad page context Use maxCharacters to keep payloads bounded. highlights Agent workflows and factual lookups Most token efficient default. Excerpts are grounded in the source page. summary Compression or structured per page extraction Adds Exa side synthesis per page. Avoid requesting multiple modes unless the caller truly needs multiple views of the same page. Freshness and Crawling Use maxAgeHours as the normative freshness control. Value Behavior omitted Use default cache first behavior with crawl fallback when needed. positive integer Use cache if it is less than N hours old, otherwise live crawl. 0 Always live crawl. Highest freshness, higher latency. 1 Cache only. Fastest, but fails if no cached content exists. Set livecrawlTimeout when live crawling should not block past a fixed budget. Subpages and extras Use subpages and subpageTarget when linked pages matter. Start with subpages around 5 to 10 , then increase only when the caller needs broader site coverage. Response Fields and Statuses Inspect statuses even when the HTTP status is 200. The endpoint can succeed for one URL and fail for another in the same request. Field Type Description requestId string Unique request identifier. results array Extracted content result objects. results[].title string Page title. results[].url string Page URL. results[].publishedDate string or null Estimated publication date when available. results[].author string or null Author when available. results[].text string Returned when text is requested. results[].highlights string[] Returned when highlights is requested. results[].highlightScores number[] Similarity scores for highlights. results[].summary string Returned when summary is requested. results[].subpages array Nested result objects from subpage crawling. results[].extras.links string[] Extracted links when requested. statuses array Per URL success or error states. Always inspect this field. statuses[].id string Requested URL. statuses[].status string success or error . statuses[].error.tag string Error type for failed URLs. statuses[].error.httpStatusCode integer or null HTTP code associated with a per URL failure. costDollars.total number Total request cost when returned. Common per URL error tags include CRAWL NOT FOUND , CRAWL TIMEOUT , CRAWL LIVECRAWL TIMEOUT , SOURCE NOT AVAILABLE , UNSUPPORTED URL , and CRAWL UNKNOWN ERROR . Critical Pitfalls Keep text , highlights , and summary at the top level on /contents . Do not wrap extraction options in a contents object; that nesting belongs to /search . Do not assume HTTP 200 means every URL succeeded; inspect statuses . Do not send stream: true ; /contents is not a streaming endpoint. Do not send tokensNum ; use text.maxCharacters to cap extracted text. Do not use useAutoprompt , numSentences , highlightsPerUrl , or older livecrawl string values in new requests. Prefer maxAgeHours for freshness and pair it with livecrawlTimeout when crawl latency matters. Use subpageTarget with subpages ; otherwise subpage selection is best effort. Pick one of highlights , text , or summary by default. Stack modes only when the caller truly needs multiple views of each page.