Extract
Up to 20 URLs (HTML or PDF) per call as clean Markdown, text or HTML, with query-ranked highlights, links and images. Failed URLs are listed with a reason and cost nothing.
POST
https://api.serpkite.com/v1/extract
Credits: 1 per URL that comes back (0.5 from cache); failed URLs are free
Extract reads many pages in one call: the same fetch and Markdown conversion as /v1/webpage, PDFs included, for up to 20 URLs at once, fetched in parallel. Duplicate URLs are dropped.
With query and highlights, each page also gets its passages most relevant to the query, ranked with BM25 (a lexical ranker: deterministic, no model involved). Use max_tokens to cap each page instead.
The response is 200 even when some or all URLs fail: each failure is listed in failed with a code (invalid_request, upstream_error, upstream_timeout, empty_page) and costs nothing.
Request body
| Parameter | Type | Default | Description |
|---|---|---|---|
urls required | string[] | 1–20 public http(s) URLs of HTML pages or PDFs. | |
format | string | markdown | What each result carries: markdown, text or the raw html. One of: markdown, text, html. |
query | string | What highlights are ranked against (up to 4,096 characters). |
|
highlights | integer | 0 | Passages per page ranked against query, 0–10 (about 600 characters each). |
max_tokens | integer | Trim each page's Markdown or text to about this many tokens (100–100,000). | |
include_links | boolean | false | Webpage: also return the page's outbound links (url, text). |
include_images | boolean | false | Webpage: also return the page's image URLs. |
max_age | integer | Accept a cached page up to this many seconds old, for 0.5 credit. | |
country | string | Fetch through an exit in this country (two-letter ISO code). | |
timeout | integer | Seconds to wait for the pages, 1–90 (default 50). A page still loading is listed in failed (upstream_timeout) and not charged. |
Example request
Authenticate with Authorization: Bearer $SERPKITE_API_KEY (GET requests may pass
?api_key= instead). The official SDKs read
SERPKITE_API_KEY for you. See Authentication.
curl https://api.serpkite.com/v1/extract \
-H "Authorization: Bearer $SERPKITE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"urls":["https://en.wikipedia.org/wiki/Web_crawler","https://www.rfc-editor.org/rfc/rfc9309.html"],"query":"robots.txt crawl delay","highlights":3}'Response fields
| Field | Type | Description |
|---|---|---|
request | object | urls (deduplicated), format, query. |
results[] | array | One entry per page read: url (final, after redirects), status_code, title, published_at, metadata, cached, markdown / text / html, and links, images, highlights (text, score, heading) when asked for. |
failed[] | array | URLs that could not be read: url, error (code, message). Not charged. |
meta | object | request_id, credits_used, latency_ms, succeeded, failed. |
Every billed response also carries credit and latency headers
(X-Credits-Used, X-Credits-Remaining, X-Request-Id…).
Example response
{
"request": {
"endpoint": "extract",
"urls": [
"https://en.wikipedia.org/wiki/Web_crawler",
"https://www.rfc-editor.org/rfc/rfc9309.html"
],
"format": "markdown",
"query": "robots.txt crawl delay"
},
"results": [
{
"url": "https://www.rfc-editor.org/rfc/rfc9309.html",
"status_code": 200,
"title": "RFC 9309: Robots Exclusion Protocol",
"metadata": {
"title": "RFC 9309: Robots Exclusion Protocol",
"content_type": "text/html"
},
"cached": false,
"markdown": "# Robots Exclusion Protocol\n\n…",
"highlights": [
{
"text": "Crawlers SHOULD NOT use the cached version for more than 24 hours, unless the robots.txt is unreachable…",
"score": 6.214,
"heading": "Robots Exclusion Protocol > Caching"
}
]
}
],
"failed": [
{
"url": "https://en.wikipedia.org/wiki/Web_crawler",
"error": {
"code": "upstream_timeout",
"message": "the page did not load in time"
}
}
],
"meta": {
"request_id": "req_01J8ZK4M6Q2V7",
"credits_used": 1,
"latency_ms": 2210,
"succeeded": 1,
"failed": 1
}
}Errors
Errors use one shape: {"error":{"code","message","request_id"}}. Errors are never billed.
Full list in Errors.
| Status | Code | Meaning |
|---|---|---|
| 400 | invalid_request | A parameter is missing or invalid. |
| 401 | unauthorized | The API key is missing, invalid or revoked. |
| 402 | insufficient_credits | Your balance is too low. Buy a pack or wait for the monthly free grant. |
| 429 | rate_limited | Too many requests per second for your plan. Retry after the Retry-After header. |
| 503 | upstream_error | Google could not be fetched or parsed. Not billed; retry after Retry-After. |
Notes
- Pages are fetched in parallel, 5 at a time; the call takes as long as the slowest page or
timeout. - The remote MCP server has an
extracttool.