# Map

`POST https://api.serpkite.com/v1/map` · Credits: 1 per call that found URLs (free when nothing is found)

A site's URLs from robots.txt, its sitemaps (XML, gzip, RSS/Atom, plain text and indexes) and the links on the start page, cleaned and de-duplicated, filtered by path regexes and optionally ranked by a search phrase.

Map lists the pages of a website without reading them: it reads the sitemaps named in `robots.txt` (or `/sitemap.xml`), follows sitemap indexes, understands XML, gzip, RSS/Atom and plain-text sitemaps, and adds the links of the start page. Use it to pick what to read next with [`/v1/webpage`](https://serpkite.com/docs/endpoints/webpage) or to scope a crawl.

URLs are cleaned (fragments and tracking parameters such as `utm_*`, `gclid` and `fbclid` dropped), kept on the site's host, filtered by your path patterns and de-duplicated (`www.`, scheme and trailing slash folded). With `search`, the URLs are ranked by how well their path words and titles match (BM25) and the ones that don't match at all are dropped.

## Request body

| Parameter | Type | Default | Description |
| --- | --- | --- | --- |
| `url` **required** | string |  | Any page of the site. Sitemaps are read from its origin. |
| `search` | string |  | Keep only URLs relevant to these words, most relevant first (up to 512 characters). |
| `limit` | integer | `100` | URLs to return, 1–5,000. |
| `sitemap` | string | `include` | `include` sitemaps and the start page's links, `only` sitemaps, `skip` sitemaps (links only). One of: `include`, `only`, `skip`. |
| `include_subdomains` | boolean | `false` | Also keep URLs on subdomains of the start page's host. |
| `include_paths` | string[] |  | Up to 20 regular expressions matched against the URL path; keep only paths matching one, e.g. `^/docs/`. |
| `exclude_paths` | string[] |  | Up to 20 regular expressions matched against the URL path; never keep paths matching one. |
| `ignore_query_parameters` | boolean | `false` | Treat URLs that differ only in their query string as one (the first one found is kept). |

## Example request

cURL:

```bash
curl https://api.serpkite.com/v1/map \
  -H "Authorization: Bearer $SERPKITE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://docs.python.org/3/","search":"asyncio","limit":50}'
```

TypeScript:

```ts
import { SerpKite } from "serpkite";

const sk = new SerpKite(); // reads SERPKITE_API_KEY
const res = await sk.map({ url: "https://docs.python.org/3/", search: "asyncio", limit: 50 });
console.log(res.results.map((u) => u.url));
```

Python:

```python
from serpkite import SerpKite

sk = SerpKite()  # reads SERPKITE_API_KEY
res = sk.map("https://docs.python.org/3/", search="asyncio", limit=50)
print([u.url for u in res.results])
```

## Response fields

| Field | Type | Description |
| --- | --- | --- |
| `request` | object | The normalised request: `url`, `limit`, `sitemap`, `include_subdomains`, and the filters you sent. |
| `results[]` | array | `url`, `title` (link text or page title, when known), `lastmod` (the sitemap's, as written), `source` (`sitemap` or `page`). |
| `meta` | object | `request_id`, `credits_used`, `latency_ms`, `count`. |

## Example response

```json
{
  "request": {
    "endpoint": "map",
    "url": "https://docs.python.org/3/",
    "limit": 50,
    "sitemap": "include",
    "include_subdomains": false,
    "search": "asyncio"
  },
  "results": [
    {
      "url": "https://docs.python.org/3/library/asyncio.html",
      "title": "asyncio — Asynchronous I/O",
      "source": "page"
    },
    {
      "url": "https://docs.python.org/3/library/asyncio-task.html",
      "lastmod": "2026-09-01",
      "source": "sitemap"
    }
  ],
  "meta": {
    "request_id": "req_01J8ZK4M6Q2V7",
    "credits_used": 1,
    "latency_ms": 840,
    "count": 2
  }
}
```

## Errors

| Status | Code | Meaning |
| --- | --- | --- |
| 400 | `invalid_request` | A parameter is missing or invalid. |
| 401 | `unauthorized` | The API key is missing, invalid or revoked. |
| 402 | `insufficient_credits` | Your balance is too low. Buy a pack or wait for the monthly free grant. |
| 429 | `rate_limited` | Too many requests per second for your plan. Retry after the Retry-After header. |
| 503 | `upstream_error` | Google could not be fetched or parsed. Not billed; retry after Retry-After. |

All errors: https://serpkite.com/docs/errors

## Notes

- A call that finds no URL is free. A call takes at most about 30 seconds.
- Sitemaps are fetched like any page: logged out, through SerpKite's proxies.
- The remote MCP server has a `map` tool with the same parameters.

## Related

- [Site ingestion guide](https://serpkite.com/docs/guides/site-ingestion)
- [Webpage](https://serpkite.com/docs/endpoints/webpage)
- [MCP server](https://serpkite.com/docs/mcp)