Build a RAG pipeline on live search
Answer questions from the live web. Search Google with SerpKite, get the top pages as clean Markdown in the same call, chunk and rank them, and have an LLM answer with citations. Python and TypeScript code, plus the credit math.
Retrieval-augmented generation (RAG) on a fixed document set goes stale. For questions about current events, prices, releases or anything on the open web, retrieve from Google instead. The pipeline is short:
- Search Google for the question.
- Fetch the top result pages as Markdown.
- Chunk and rank the text against the question.
- Answer with an LLM, citing the source URLs.
SerpKite does steps 1 and 2 in one request with include_content.
Step 1 and 2: search and fetch in one call
include_content: N fetches the top N organic results (0–5) and adds each page’s main content as Markdown in results[].content:
curl https://api.serpkite.com/v1/search \
-H "Authorization: Bearer $SERPKITE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"q":"what changed in python 3.14 asyncio","country":"us","num":10,"include_content":3}'Cost: 1 credit for the search plus 1 per fetched page, so this call is 4 credits. Pages that can’t be fetched aren’t billed; their content is simply missing.
If you want to choose which pages to read (say, skip forums or pick by domain), search first and then fetch selected URLs with POST /v1/webpage (sk.webpage(url) in the SDKs) at 1 credit each.
Python
This version uses the SerpKite Python SDK, a simple lexical ranker so there’s no vector database to set up, and an LLM for the answer. Swap in embeddings when the corpus grows (see below).
import re
import anthropic
from serpkite import SerpKite
sk = SerpKite() # reads SERPKITE_API_KEY
llm = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
def retrieve(question: str, pages: int = 3) -> list[dict]:
res = sk.search(
question,
include_content=pages,
fields="results.title,results.link,results.snippet,results.content",
)
return [r.model_dump() for r in res.results] # pydantic models -> dicts
def chunk(text: str, size: int = 1200, overlap: int = 200) -> list[str]:
# Split on paragraphs, then pack into ~size-character chunks.
paras, chunks, cur = re.split(r"\n{2,}", text), [], ""
for p in paras:
if len(cur) + len(p) > size and cur:
chunks.append(cur)
cur = cur[-overlap:]
cur += "\n\n" + p
if cur.strip():
chunks.append(cur)
return chunks
def score(question: str, text: str) -> float:
terms = {t for t in re.findall(r"\w+", question.lower()) if len(t) > 2}
words = re.findall(r"\w+", text.lower())
return sum(w in terms for w in words) / (len(words) ** 0.5 + 1)
def answer(question: str) -> str:
results = retrieve(question)
passages = []
for i, r in enumerate(results, start=1):
body = r.get("content") or r.get("snippet", "")
for c in chunk(body):
passages.append((score(question, c), i, r["link"], c))
top = sorted(passages, reverse=True)[:8]
context = "\n\n".join(f"[{i}] {url}\n{text}" for _, i, url, text in top)
sources = "\n".join(f"[{i}] {r['title']}: {r['link']}" for i, r in enumerate(results, start=1))
msg = llm.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
system="Answer only from the provided sources. Cite them inline as [n]. If the sources don't answer the question, say so.",
messages=[{"role": "user", "content": f"Sources:\n{context}\n\nQuestion: {question}"}],
)
text = "".join(b.text for b in msg.content if b.type == "text")
return f"{text}\n\nSources:\n{sources}"
print(answer("What changed in Python 3.14 asyncio?"))
TypeScript
import Anthropic from "@anthropic-ai/sdk";
import { SerpKite } from "serpkite";
const sk = new SerpKite(); // reads SERPKITE_API_KEY
const llm = new Anthropic(); // reads ANTHROPIC_API_KEY
type Result = { title: string; link: string; snippet?: string; content?: string };
async function retrieve(question: string, pages = 3): Promise<Result[]> {
const res = await sk.search({
q: question,
include_content: pages,
fields: "results.title,results.link,results.snippet,results.content",
});
return res.results;
}
function chunk(text: string, size = 1200): string[] {
const out: string[] = [];
let cur = "";
for (const p of text.split(/\n{2,}/)) {
if (cur.length + p.length > size && cur) {
out.push(cur);
cur = "";
}
cur += `\n\n${p}`;
}
if (cur.trim()) out.push(cur);
return out;
}
function score(question: string, text: string): number {
const terms = new Set(question.toLowerCase().match(/\w{3,}/g) ?? []);
const words = text.toLowerCase().match(/\w+/g) ?? [];
return words.filter((w) => terms.has(w)).length / (Math.sqrt(words.length) + 1);
}
export async function answer(question: string): Promise<string> {
const results = await retrieve(question);
const passages = results.flatMap((r, i) =>
chunk(r.content ?? r.snippet ?? "").map((text) => ({ n: i + 1, url: r.link, text, s: score(question, text) })),
);
const top = passages.sort((a, b) => b.s - a.s).slice(0, 8);
const context = top.map((p) => `[${p.n}] ${p.url}\n${p.text}`).join("\n\n");
const msg = await llm.messages.create({
model: "claude-opus-5-5",
max_tokens: 16000,
system: "Answer only from the provided sources. Cite them inline as [n]. If the sources don't answer the question, say so.",
messages: [{ role: "user", content: `Sources:\n${context}\n\nQuestion: ${question}` }],
});
const text = msg.content.flatMap((b) => (b.type === "text" ? [b.text] : [])).join("");
const sources = results.map((r, i) => `[${i + 1}] ${r.title}: ${r.link}`).join("\n");
return `${text}\n\nSources:\n${sources}`;
}
LangChain
With langchain-serpkite, steps 1 and 2 are a retriever that returns LangChain Documents, with the page Markdown as content:
from langchain_serpkite import SerpKiteRetriever
retriever = SerpKiteRetriever(k=5, include_content=2)
docs = retriever.invoke("What changed in Python 3.14 asyncio?")
Plug it into any LangChain chain or LangGraph node. To load specific URLs, use SerpKiteWebpageLoader(["https://…"]).load().
Using embeddings instead of lexical ranking
The lexical ranker above is fine for a handful of pages. For better recall, especially with paraphrased questions:
- Embed each chunk with the embedding model you already use.
- Embed the question with the same model.
- Keep the top 5–10 chunks by cosine similarity, optionally re-ranked with a cross-encoder.
With include_content: 3 you typically get a few dozen chunks per question, small enough to embed on the fly and keep in memory. Persist embeddings only if you answer many questions over the same pages; add max_age to the search so repeated questions reuse cached results.
Credit math
| Step | Credits |
|---|---|
Search (/v1/search, 10 results) |
1 |
Page content (include_content: 3) |
+3 |
| Per question | 4 |
Same question again within max_age (cache hit) |
2 (half of 4) |
In general a question costs 1 + N credits, where N is the number of pages you fetch. At the Pro pack price ($300 for 500,000 credits, $0.60 per 1,000), 4 credits is $0.0024 per question. The X-Credits-Used header on each response confirms the actual charge.
Tips
- Ground the prompt. Tell the model to answer only from the sources and to say when they don’t cover the question. Include the URLs so citations are checkable.
- Localize. Pass
countryandlanguagefor country- and language-specific answers. See Localization. - Freshness. Add
time: "week"(orday,month) for news-like questions. - Batch offline work. If you pre-compute answers for many questions,
POST /v1/batchesruns searches at half price.
Related
Last updated: 2026-09-29