# Build a RAG pipeline on live search

> Answer questions from the live web. Search Google with SerpKite, get the top pages as clean Markdown in the same call, chunk and rank them, and have an LLM answer with citations. Python and TypeScript code, plus the credit math.

Retrieval-augmented generation (RAG) on a fixed document set goes stale. For questions about current events, prices, releases or anything on the open web, retrieve from Google instead. The pipeline is short:

1. **Search** Google for the question.
2. **Fetch** the top result pages as Markdown.
3. **Chunk and rank** the text against the question.
4. **Answer** with an LLM, citing the source URLs.

SerpKite does steps 1 and 2 in one request with [`include_content`](https://serpkite.com/docs/include-content).

## Step 1 and 2: search and fetch in one call

`include_content: N` fetches the top `N` organic results (0–5) and adds each page's main content as Markdown in `results[].content`:

cURL:

```bash
curl https://api.serpkite.com/v1/search \
  -H "Authorization: Bearer $SERPKITE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"q":"what changed in python 3.14 asyncio","country":"us","num":10,"include_content":3}'
```

TypeScript:

```ts
import { SerpKite } from "serpkite";

const sk = new SerpKite(); // reads SERPKITE_API_KEY
const res = await sk.search({ q: "what changed in python 3.14 asyncio", country: "us", num: 10, include_content: 3 });
console.log(res.results[0].title, res.meta.credits_used);
```

Python:

```python
from serpkite import SerpKite

sk = SerpKite()  # reads SERPKITE_API_KEY
res = sk.search("what changed in python 3.14 asyncio", country="us", num=10, include_content=3)
print(res.results[0].title, res.meta.credits_used)
```

Go:

```go
package main

import (
	"context"
	"fmt"
	"log"

	serpkite "github.com/serpkite/serpkite-go"
)

func main() {
	ctx := context.Background()
	c := serpkite.NewClient() // reads SERPKITE_API_KEY
	res, err := c.Search(ctx, serpkite.SearchParams{Q: "what changed in python 3.14 asyncio", Country: "us", Num: 10, IncludeContent: 3})
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Results[0].Title, res.Meta.CreditsUsed)
}
```

Cost: 1 credit for the search plus 1 per fetched page, so this call is **4 credits**. Pages that can't be fetched aren't billed; their `content` is simply missing.

If you want to choose which pages to read (say, skip forums or pick by domain), search first and then fetch selected URLs with [`POST /v1/webpage`](https://serpkite.com/docs/endpoints/webpage) (`sk.webpage(url)` in the SDKs) at 1 credit each.

## Python

This version uses the [SerpKite Python SDK](https://serpkite.com/docs/sdks#python), a simple lexical ranker so there's no vector database to set up, and an LLM for the answer. Swap in embeddings when the corpus grows (see below).

```python
import re
import anthropic
from serpkite import SerpKite

sk = SerpKite()  # reads SERPKITE_API_KEY
llm = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY

def retrieve(question: str, pages: int = 3) -> list[dict]:
    res = sk.search(
        question,
        include_content=pages,
        fields="results.title,results.link,results.snippet,results.content",
    )
    return [r.model_dump() for r in res.results]  # pydantic models -> dicts

def chunk(text: str, size: int = 1200, overlap: int = 200) -> list[str]:
    # Split on paragraphs, then pack into ~size-character chunks.
    paras, chunks, cur = re.split(r"\n{2,}", text), [], ""
    for p in paras:
        if len(cur) + len(p) > size and cur:
            chunks.append(cur)
            cur = cur[-overlap:]
        cur += "\n\n" + p
    if cur.strip():
        chunks.append(cur)
    return chunks

def score(question: str, text: str) -> float:
    terms = {t for t in re.findall(r"\w+", question.lower()) if len(t) > 2}
    words = re.findall(r"\w+", text.lower())
    return sum(w in terms for w in words) / (len(words) ** 0.5 + 1)

def answer(question: str) -> str:
    results = retrieve(question)
    passages = []
    for i, r in enumerate(results, start=1):
        body = r.get("content") or r.get("snippet", "")
        for c in chunk(body):
            passages.append((score(question, c), i, r["link"], c))
    top = sorted(passages, reverse=True)[:8]

    context = "\n\n".join(f"[{i}] {url}\n{text}" for _, i, url, text in top)
    sources = "\n".join(f"[{i}] {r['title']}: {r['link']}" for i, r in enumerate(results, start=1))
    msg = llm.messages.create(
        model="claude-opus-5-5",
        max_tokens=16000,
        system="Answer only from the provided sources. Cite them inline as [n]. If the sources don't answer the question, say so.",
        messages=[{"role": "user", "content": f"Sources:\n{context}\n\nQuestion: {question}"}],
    )
    text = "".join(b.text for b in msg.content if b.type == "text")
    return f"{text}\n\nSources:\n{sources}"

print(answer("What changed in Python 3.14 asyncio?"))
```

## TypeScript

```ts
import Anthropic from "@anthropic-ai/sdk";
import { SerpKite } from "serpkite";

const sk = new SerpKite(); // reads SERPKITE_API_KEY
const llm = new Anthropic(); // reads ANTHROPIC_API_KEY

type Result = { title: string; link: string; snippet?: string; content?: string };

async function retrieve(question: string, pages = 3): Promise<Result[]> {
  const res = await sk.search({
    q: question,
    include_content: pages,
    fields: "results.title,results.link,results.snippet,results.content",
  });
  return res.results;
}

function chunk(text: string, size = 1200): string[] {
  const out: string[] = [];
  let cur = "";
  for (const p of text.split(/\n{2,}/)) {
    if (cur.length + p.length > size && cur) {
      out.push(cur);
      cur = "";
    }
    cur += `\n\n${p}`;
  }
  if (cur.trim()) out.push(cur);
  return out;
}

function score(question: string, text: string): number {
  const terms = new Set(question.toLowerCase().match(/\w{3,}/g) ?? []);
  const words = text.toLowerCase().match(/\w+/g) ?? [];
  return words.filter((w) => terms.has(w)).length / (Math.sqrt(words.length) + 1);
}

export async function answer(question: string): Promise<string> {
  const results = await retrieve(question);
  const passages = results.flatMap((r, i) =>
    chunk(r.content ?? r.snippet ?? "").map((text) => ({ n: i + 1, url: r.link, text, s: score(question, text) })),
  );
  const top = passages.sort((a, b) => b.s - a.s).slice(0, 8);
  const context = top.map((p) => `[${p.n}] ${p.url}\n${p.text}`).join("\n\n");

  const msg = await llm.messages.create({
    model: "claude-opus-5-5",
    max_tokens: 16000,
    system: "Answer only from the provided sources. Cite them inline as [n]. If the sources don't answer the question, say so.",
    messages: [{ role: "user", content: `Sources:\n${context}\n\nQuestion: ${question}` }],
  });
  const text = msg.content.flatMap((b) => (b.type === "text" ? [b.text] : [])).join("");
  const sources = results.map((r, i) => `[${i + 1}] ${r.title}: ${r.link}`).join("\n");
  return `${text}\n\nSources:\n${sources}`;
}
```

## LangChain

With [`langchain-serpkite`](https://serpkite.com/docs/sdks#langchain), steps 1 and 2 are a retriever that returns LangChain `Document`s, with the page Markdown as content:

```python
from langchain_serpkite import SerpKiteRetriever

retriever = SerpKiteRetriever(k=5, include_content=2)
docs = retriever.invoke("What changed in Python 3.14 asyncio?")
```

Plug it into any LangChain chain or LangGraph node. To load specific URLs, use `SerpKiteWebpageLoader(["https://…"]).load()`.

## Using embeddings instead of lexical ranking

The lexical ranker above is fine for a handful of pages. For better recall, especially with paraphrased questions:

1. Embed each chunk with the embedding model you already use.
2. Embed the question with the same model.
3. Keep the top 5–10 chunks by cosine similarity, optionally re-ranked with a cross-encoder.

With `include_content: 3` you typically get a few dozen chunks per question, small enough to embed on the fly and keep in memory. Persist embeddings only if you answer many questions over the same pages; add `max_age` to the search so repeated questions reuse cached results.

## Credit math

| Step | Credits |
| --- | --- |
| Search (`/v1/search`, 10 results) | 1 |
| Page content (`include_content: 3`) | +3 |
| **Per question** | **4** |
| Same question again within `max_age` (cache hit) | 2 (half of 4) |

In general a question costs **1 + N** credits, where N is the number of pages you fetch. At the Pro pack price ($300 for 500,000 credits, $0.60 per 1,000), 4 credits is $0.0024 per question. The `X-Credits-Used` header on each response confirms the actual charge.

## Tips

- **Ground the prompt.** Tell the model to answer only from the sources and to say when they don't cover the question. Include the URLs so citations are checkable.
- **Localize.** Pass `country` and `language` for country- and language-specific answers. See [Localization](https://serpkite.com/docs/localization).
- **Freshness.** Add `time: "week"` (or `day`, `month`) for news-like questions.
- **Batch offline work.** If you pre-compute answers for many questions, [`POST /v1/batches`](https://serpkite.com/docs/batch) runs searches at half price.

## Related

- [Page content](https://serpkite.com/docs/include-content): How include_content fetches and bills pages.
- [Webpage to Markdown](https://serpkite.com/docs/endpoints/webpage): Fetch any public URL as Markdown.
- [Tool calling for agents](https://serpkite.com/docs/guides/agents-tool-calling): Let the model decide when to search.
- [LlamaIndex](https://serpkite.com/integrations/llamaindex): Use SerpKite as a LlamaIndex retriever.