# Ingest a website or a set of URLs

> Discover URLs with map, read selected HTML pages and PDFs with extract, or crawl a site asynchronously. SDK examples, partial failures, polling, cancellation and credit reservations.

Use SerpKite's site APIs to turn public web content into documents for a knowledge base, a research pipeline or an agent. Choose the endpoint by how much of the site you need:

| Starting point | Endpoint | Result |
| --- | --- | --- |
| One known URL | [`/v1/webpage`](https://serpkite.com/docs/endpoints/webpage) | One page's Markdown and metadata |
| Up to 20 known URLs | [`/v1/extract`](https://serpkite.com/docs/endpoints/extract) | Page content, optional ranked passages and individual failures |
| A site whose URLs you want to discover | [`/v1/map`](https://serpkite.com/docs/endpoints/map) | URLs from sitemaps and the start page's links |
| A site section you want to read | [`/v1/crawl`](https://serpkite.com/docs/endpoints/crawl) | An asynchronous task containing the pages it read |

All requests use the same [Bearer API key](https://serpkite.com/docs/authentication). The site endpoints take JSON objects over `POST`; their options differ from the search endpoints. For example, `format: "markdown"` on extract or crawl selects the content field inside a JSON response. It does not turn the whole response into a Markdown string.

## Discover URLs, then extract selected pages

Map discovers URLs without extracting every page. Its `search` option ranks URL paths and known titles against your terms; it does not search the pages' full contents. `include_paths` and `exclude_paths` are regular expressions matched against the URL path.

The following Python example discovers Python documentation about asyncio, reads at most 20 matching pages and prints passages relevant to a question. Install the [Python SDK](https://serpkite.com/docs/sdks#python) and set `SERPKITE_API_KEY` first.

```python
from serpkite import SerpKite

sk = SerpKite()
discovered = sk.map(
    "https://docs.python.org/3/",
    search="asyncio",
    include_paths=[r"^/3/library/"],
    limit=20,
)
urls = [row.url for row in discovered.results]

if urls:
    extracted = sk.extract(
        urls,
        query="How do asyncio task groups handle cancellation?",
        highlights=3,
        max_tokens=4000,
        include_links=True,
        timeout=50,
    )
    for page in extracted.results:
        print(page.url, page.title)
        for passage in page.highlights or []:
            print(passage.text)
    for failure in extracted.failed:
        print("Could not read:", failure.url, failure.error.code)
else:
    print("No matching URLs were discovered.")
```

The same workflow in TypeScript:

```ts
import { SerpKite } from "serpkite";

const sk = new SerpKite();
const discovered = await sk.map({
  url: "https://docs.python.org/3/",
  search: "asyncio",
  include_paths: ["^/3/library/"],
  limit: 20,
});
const urls = discovered.results.map((row) => row.url);

if (urls.length) {
  const extracted = await sk.extract({
    urls,
    query: "How do asyncio task groups handle cancellation?",
    highlights: 3,
    max_tokens: 4000,
    include_links: true,
    timeout: 50,
  });
  for (const page of extracted.results) {
    console.log(page.url, page.highlights);
  }
  for (const failure of extracted.failed) {
    console.error(failure.url, failure.error.code);
  }
}
```

An extract response is `200` even when some or all URLs fail. Check both `results` and `failed`; only successful pages are charged. URLs can redirect, so store the final `results[].url` with each document. For more than 20 URLs, divide your list into groups of 20 and respect your [rate limit](https://serpkite.com/docs/rate-limits).

`highlights` are passages ranked with BM25, a lexical ranker, rather than generated summaries. Extract adds them alongside the page content. `max_tokens` caps Markdown or text per page. PDFs use their text layer; scanned PDFs without readable text produce no content. See [Extract](https://serpkite.com/docs/endpoints/extract) for formats and limits.

## Crawl a section in the background

Crawl follows links and sitemap URLs and reads pages for you. Start with a small `limit`, then poll the task. The start URL is always read; path filters scope the pages discovered from it. Set the start URL to the section you intend to ingest.

```python
from serpkite import SerpKite

sk = SerpKite()
task = sk.crawl(
    "https://docs.python.org/3/library/",
    limit=25,
    max_depth=2,
    include_paths=[r"^/3/library/"],
    query="asyncio cancellation",
    max_tokens=4000,
)
finished = sk.wait_for_crawl(task.id)

print(finished.status, finished.credits_used)
if finished.result:
    for page in finished.result.pages:
        print(page.url, page.markdown)
    print("Stopped because:", finished.result.stats.stopped)
if finished.error:
    print(finished.error.code, finished.error.message)
```

`query` prioritizes promising URLs and link text; it does not exclude every irrelevant page. `max_depth` is the number of link hops, up to 10; sitemap URLs count as one hop. `sitemap: "only"` reads the start page and sitemap URLs without following page links, while `sitemap: "skip"` follows links without sitemap discovery.

For direct HTTP use:

```bash
curl https://api.serpkite.com/v1/crawl \
  -H "Authorization: Bearer $SERPKITE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url":"https://docs.python.org/3/library/",
    "limit":25,
    "max_depth":2,
    "include_paths":["^/3/library/"],
    "max_tokens":4000
  }'
```

The `202` response contains `id` and `poll_url`. Poll `GET /v1/crawl/{id}` until `status` is `completed`, `failed` or `canceled`. While running, `progress` reports counts; after it ends, `result.pages`, `result.failed` and `result.stats` describe what was read. A completed crawl can contain fewer than `limit` pages or individual failures, so inspect the result before indexing it.

Polling is free. Results are kept for 24 hours: save the documents in your own store before they expire. A `webhook_url` on the create request delivers a signed `crawl.completed` event; large results are omitted from the webhook with `result_omitted: true`, so fetch its `poll_url`. See [Webhooks](https://serpkite.com/docs/webhooks).

Cancel with `DELETE /v1/crawl/{id}` or `sk.cancel_crawl(task.id)`. A queued crawl is refunded; a running crawl stops at its next checkpoint and charges for pages already read. Canceling an unknown task returns `404`; a finished task returns `409`.

## Budget and retries

| Operation | Credits |
| --- | --- |
| Map | 1 per nonempty call |
| Extract | 1 per successfully read URL, or 0.5 from cache |
| Crawl | 1 per successfully read page, or 0.5 from cache |
| Crawl polling and cancellation | 0 |

Crawl reserves `limit` credits when queued and refunds the unused part when it ends. For example, `limit: 25` reserves 25 credits; reading 18 live pages settles to 18 and refunds 7. Reservations count against your balance, key limit and spend cap.

`max_age` lets extract and crawl accept cached pages. Failed and empty pages cost nothing. Map does not accept `max_age`. The [batch lane](https://serpkite.com/docs/batch) accepts search verticals and webpage requests; it does not accept extract, map or crawl.

The SDKs retry extract and crawl creation only on `429`. After a lost response, automatically repeating either request can charge for another extraction or start another crawl. The batch API's `Idempotency-Key` support does not extend to these endpoints.

Crawls honor `robots.txt` and `Crawl-delay`, which can reduce coverage or slow a task. SerpKite reads public, logged-out content; private addresses and authenticated pages are unavailable.

## Related

- [RAG pipeline](https://serpkite.com/docs/guides/rag-pipeline): retrieve passages and build a prompt with citations.
- [Scheduled monitoring](https://serpkite.com/docs/guides/scheduled-monitoring): watch for new results or changed pages.
- [Crawl reference](https://serpkite.com/docs/endpoints/crawl): all options, task states and response fields.
- [MCP server](https://serpkite.com/docs/mcp): `map`, `extract`, `crawl` and `crawl_result` tools.