Ingest a website or a set of URLs
Discover URLs with map, read selected HTML pages and PDFs with extract, or crawl a site asynchronously. SDK examples, partial failures, polling, cancellation and credit reservations.
Use SerpKite’s site APIs to turn public web content into documents for a knowledge base, a research pipeline or an agent. Choose the endpoint by how much of the site you need:
| Starting point | Endpoint | Result |
|---|---|---|
| One known URL | /v1/webpage |
One page’s Markdown and metadata |
| Up to 20 known URLs | /v1/extract |
Page content, optional ranked passages and individual failures |
| A site whose URLs you want to discover | /v1/map |
URLs from sitemaps and the start page’s links |
| A site section you want to read | /v1/crawl |
An asynchronous task containing the pages it read |
All requests use the same Bearer API key. The site endpoints take JSON objects over POST; their options differ from the search endpoints. For example, format: "markdown" on extract or crawl selects the content field inside a JSON response. It does not turn the whole response into a Markdown string.
Discover URLs, then extract selected pages
Map discovers URLs without extracting every page. Its search option ranks URL paths and known titles against your terms; it does not search the pages’ full contents. include_paths and exclude_paths are regular expressions matched against the URL path.
The following Python example discovers Python documentation about asyncio, reads at most 20 matching pages and prints passages relevant to a question. Install the Python SDK and set SERPKITE_API_KEY first.
from serpkite import SerpKite
sk = SerpKite()
discovered = sk.map(
"https://docs.python.org/3/",
search="asyncio",
include_paths=[r"^/3/library/"],
limit=20,
)
urls = [row.url for row in discovered.results]
if urls:
extracted = sk.extract(
urls,
query="How do asyncio task groups handle cancellation?",
highlights=3,
max_tokens=4000,
include_links=True,
timeout=50,
)
for page in extracted.results:
print(page.url, page.title)
for passage in page.highlights or []:
print(passage.text)
for failure in extracted.failed:
print("Could not read:", failure.url, failure.error.code)
else:
print("No matching URLs were discovered.")
The same workflow in TypeScript:
import { SerpKite } from "serpkite";
const sk = new SerpKite();
const discovered = await sk.map({
url: "https://docs.python.org/3/",
search: "asyncio",
include_paths: ["^/3/library/"],
limit: 20,
});
const urls = discovered.results.map((row) => row.url);
if (urls.length) {
const extracted = await sk.extract({
urls,
query: "How do asyncio task groups handle cancellation?",
highlights: 3,
max_tokens: 4000,
include_links: true,
timeout: 50,
});
for (const page of extracted.results) {
console.log(page.url, page.highlights);
}
for (const failure of extracted.failed) {
console.error(failure.url, failure.error.code);
}
}
An extract response is 200 even when some or all URLs fail. Check both results and failed; only successful pages are charged. URLs can redirect, so store the final results[].url with each document. For more than 20 URLs, divide your list into groups of 20 and respect your rate limit.
highlights are passages ranked with BM25, a lexical ranker, rather than generated summaries. Extract adds them alongside the page content. max_tokens caps Markdown or text per page. PDFs use their text layer; scanned PDFs without readable text produce no content. See Extract for formats and limits.
Crawl a section in the background
Crawl follows links and sitemap URLs and reads pages for you. Start with a small limit, then poll the task. The start URL is always read; path filters scope the pages discovered from it. Set the start URL to the section you intend to ingest.
from serpkite import SerpKite
sk = SerpKite()
task = sk.crawl(
"https://docs.python.org/3/library/",
limit=25,
max_depth=2,
include_paths=[r"^/3/library/"],
query="asyncio cancellation",
max_tokens=4000,
)
finished = sk.wait_for_crawl(task.id)
print(finished.status, finished.credits_used)
if finished.result:
for page in finished.result.pages:
print(page.url, page.markdown)
print("Stopped because:", finished.result.stats.stopped)
if finished.error:
print(finished.error.code, finished.error.message)
query prioritizes promising URLs and link text; it does not exclude every irrelevant page. max_depth is the number of link hops, up to 10; sitemap URLs count as one hop. sitemap: "only" reads the start page and sitemap URLs without following page links, while sitemap: "skip" follows links without sitemap discovery.
For direct HTTP use:
curl https://api.serpkite.com/v1/crawl \
-H "Authorization: Bearer $SERPKITE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url":"https://docs.python.org/3/library/",
"limit":25,
"max_depth":2,
"include_paths":["^/3/library/"],
"max_tokens":4000
}'
The 202 response contains id and poll_url. Poll GET /v1/crawl/{id} until status is completed, failed or canceled. While running, progress reports counts; after it ends, result.pages, result.failed and result.stats describe what was read. A completed crawl can contain fewer than limit pages or individual failures, so inspect the result before indexing it.
Polling is free. Results are kept for 24 hours: save the documents in your own store before they expire. A webhook_url on the create request delivers a signed crawl.completed event; large results are omitted from the webhook with result_omitted: true, so fetch its poll_url. See Webhooks.
Cancel with DELETE /v1/crawl/{id} or sk.cancel_crawl(task.id). A queued crawl is refunded; a running crawl stops at its next checkpoint and charges for pages already read. Canceling an unknown task returns 404; a finished task returns 409.
Budget and retries
| Operation | Credits |
|---|---|
| Map | 1 per nonempty call |
| Extract | 1 per successfully read URL, or 0.5 from cache |
| Crawl | 1 per successfully read page, or 0.5 from cache |
| Crawl polling and cancellation | 0 |
Crawl reserves limit credits when queued and refunds the unused part when it ends. For example, limit: 25 reserves 25 credits; reading 18 live pages settles to 18 and refunds 7. Reservations count against your balance, key limit and spend cap.
max_age lets extract and crawl accept cached pages. Failed and empty pages cost nothing. Map does not accept max_age. The batch lane accepts search verticals and webpage requests; it does not accept extract, map or crawl.
The SDKs retry extract and crawl creation only on 429. After a lost response, automatically repeating either request can charge for another extraction or start another crawl. The batch API’s Idempotency-Key support does not extend to these endpoints.
Crawls honor robots.txt and Crawl-delay, which can reduce coverage or slow a task. SerpKite reads public, logged-out content; private addresses and authenticated pages are unavailable.
Related
- RAG pipeline: retrieve passages and build a prompt with citations.
- Scheduled monitoring: watch for new results or changed pages.
- Crawl reference: all options, task states and response fields.
- MCP server:
map,extract,crawlandcrawl_resulttools.
Last updated: 2026-10-04