Skip to content

Latest commit

 

History

History
386 lines (270 loc) · 12.6 KB

File metadata and controls

386 lines (270 loc) · 12.6 KB

openserp

PyPI version Python versions License

pip install openserp

Cloud:

import os
from openserp import OpenSERP

client = OpenSERP(api_key=os.environ["OPENSERP_API_KEY"])
resp = client.search(engine="google", text="openserp")

print(resp.results[0].title, resp.results[0].url)

Self-hosted:

from openserp import OpenSERP

client = OpenSERP(base_url="http://localhost:7000")
resp = client.search(engine="bing", text="openserp")

print(resp.results[0].title, resp.results[0].url)

Python SDK for the OpenSERP multi-engine SERP API - Google, Bing, Yandex, Baidu, DuckDuckGo, and Ecosia results in a single call. Works against the self-hosted open-source server and against OpenSERP Cloud with the same code.

Use it for AI grounding, RAG pipelines, LLM tool use, agent tool use, LangChain / LlamaIndex integrations, SEO rank tracking, competitor analysis, and search-powered automations. Open-source alternative to SerpAPI, DataForSEO, ScrapingBee, Bright Data SERP, Oxylabs SERP, and Zenserp.

Also available for TypeScript / JavaScript: @openserp/sdk.

Alpha - the API may change before 1.0.0. Pin a version in production.

Contents

Install

pip install openserp

DataFrame export is an optional extra:

pip install "openserp[pandas]"

Requires Python 3.10+.

Why OpenSERP

Need OpenSERP fit
Local development Run the OSS server and use the same SDK surface as Cloud.
AI grounding Pull fresh SERP snippets and optional extracted page text for prompts, RAG, or agents.
SEO checks Query Google, Bing, Yandex, Baidu, DuckDuckGo, and Ecosia with typed models.
Migration path Start self-hosted, then switch to Cloud by adding OPENSERP_API_KEY.

Compared with hosted-only SERP APIs, OpenSERP keeps the client contract portable. You can test locally without an API key, then use the hosted API when you want managed infrastructure.

Quickstart - OSS (self-hosted)

Run the open-source server locally, no API key required:

docker run -p 7000:7000 karust/openserp serve
from openserp import OpenSERP

client = OpenSERP(base_url="http://localhost:7000")

resp = client.search(
    engine="google",
    text="openserp",
    limit=10,
    region="US",
)

print(resp.results[0].title, resp.results[0].url)

If you pass no options, the client defaults to http://localhost:7000.

Quickstart - Cloud

Get an API key from the API keys section in the dashboard. When api_key is set, the SDK defaults base_url to https://api.openserp.org/v1 and sends Authorization: Bearer ... for you.

import os
from openserp import OpenSERP

client = OpenSERP(api_key=os.environ["OPENSERP_API_KEY"])

resp = client.search(engine="google", text="openserp")

print(resp.results[0].title)
print(client.last_response.credits)  # CreditInfo(used=..., remaining=...)

If both base_url and api_key are set, base_url wins and the key is still sent. Use this for an authenticated self-hosted deployment. Add backend="oss" when you also need OSS-only methods such as stats() or health().

Why two backends?

OpenSERP Cloud uses the same public HTTP contract as the OSS server, with a /v1/ prefix and bearer auth. The same SDK call works on both; you only change base_url / api_key. Start with OSS locally, then move to Cloud when you want the hosted API. See openserp.org/docs/oss-vs-cloud for the full comparison.

Search

single = client.search(engine="bing", text="golang", limit=10, region="US")

mega = client.mega_search(
    text="golang",
    engines=["google", "bing", "yandex"],
    mode="balanced",
    limit=20,
)

fast = client.fast_search(text="golang", engines=["google", "bing"])
any_ = client.any_search(text="golang", engines=["google", "yandex"])

mega_search aggregates multiple engines. mode is "balanced" (default, merged and deduplicated), "any" (first successful engine wins), or "fast" (engines reordered by recent health). fast_search / any_search are sugar for the matching mode.

To enrich top search results with cleaned page content, pass the extraction flags:

grounded = client.search(
    engine="google",
    text="openserp docs",
    extract=3,
    extract_mode="auto",
    min_runes=500,
)

print(grounded.results[0].extracted.content)

Extract

page = client.extract(
    url="https://openserp.org/docs",
    mode="auto",
    clean=True,
)

print(page.markdown)

Use min_runes to set the auto-mode escalation floor, clean=False for whole-page readable extraction, and use_llms_txt=True to prefer /llms-full.txt or /llms.txt for site-root URLs. Non-JSON formats are returned as strings:

markdown = client.extract(url="https://openserp.org", format="markdown")

Batch extract

batch_extract takes up to 20 URLs in one request. A URL that fails becomes an item with an error instead of failing the whole call, so one dead link never costs you the other results:

batch = client.batch_extract(
    urls=[
        "https://openserp.org/docs",
        "https://openserp.org/blog",
    ],
    mode="auto",
)

for item in batch.results:
    if item.error:
        print(item.url, "failed:", item.error)
    else:
        print(item.url, item.page_content[:120])

On the hosted API, billing is per URL and matches calling extract that many times - successful extractions bill their mode, failed and empty ones are free.

Regions

Pass region (a two-letter country code) to extract as a visitor from that country - useful for geo-fenced or localized pages. On the hosted API this adds 1 credit per successfully extracted URL:

page = client.extract(url="https://example.com/pricing", region="DE")

Images

images = client.image(engine="bing", text="golang logo", limit=20)

mega_images = client.mega_image(text="golang logo", engines=["bing", "google"])

Async

import asyncio, os
from openserp import AsyncOpenSERP


async def main() -> None:
    async with AsyncOpenSERP(api_key=os.environ["OPENSERP_API_KEY"]) as client:
        resp = await client.search(engine="google", text="openserp")
        print(resp.results[0].title)


asyncio.run(main())

Run hundreds of queries concurrently with a semaphore:

import asyncio
from openserp import AsyncOpenSERP


async def main() -> None:
    sem = asyncio.Semaphore(20)
    queries = [f"keyword {i}" for i in range(500)]

    async with AsyncOpenSERP() as client:
        async def run(query: str):
            async with sem:
                return await client.search(engine="google", text=query, limit=10)

        responses = await asyncio.gather(*(run(q) for q in queries))
        print(len(responses))


asyncio.run(main())

Endpoint availability

OSS-only operational methods raise OssOnlyError when the client is configured for Cloud:

client.parse_google(html="<html>...</html>")
client.stats()
client.health()

Cloud-only account methods raise CloudOnlyError when the client is configured for OSS:

client.me()
client.pricing()
client.engines_status()
client.engines_capabilities()

The backend is inferred from base_url and api_key. Pass backend="oss" or backend="cloud" to the constructor to override.

Cloud engine status

Authenticated Cloud engine status contains engines; anonymous status contains only overall. Engine states are operational, loaded, and down; overall states are operational, degraded, and down. Idle engines retain their last known state.

Cloud paging and filters

For later pages, keep limit=10 and pass pagination.next_start while pagination.has_more is true. Google, Bing, and Yandex take offsets in multiples of 10; Baidu supports early pages, Ecosia any offset, and DuckDuckGo only the first page. Unsupported offsets return 400 invalid_request without charge.

with OpenSERP(api_key=os.environ["OPENSERP_API_KEY"]) as client:
    first = client.search(engine="google", text="openserp", limit=10)
    if first.pagination and first.pagination.has_more:
        next_page = client.search(
            engine="google", text="openserp", limit=10, start=first.pagination.next_start,
        )
        print(next_page.results)

Cloud web search accepts publication-date ranges (date="20250101..20251231") on Google and Ecosia. Malformed ranges and unsupported filters return 400 invalid_request; change the request before retrying.

mega_search in balanced mode rejects start > 0. any_search and fast_search forward the offset to compatible engines. Any starts engines in your order, overlapping slow attempts; Fast prioritizes recent health and latency. Use response.meta.engine_used or client.last_response.engine_used for the winner. engines_tried and engines_skipped may be absent. These rules also apply to the async client.

Telemetry

client.last_response is updated after every HTTP response:

client.last_response.credits          # Cloud - CreditInfo(used, remaining)
client.last_response.engine_used      # both - X-Engine-Used
client.last_response.request_id       # X-Request-Id, also meta.request_id
client.last_response.fallback_engine  # OSS only
client.last_response.cache            # OSS only
client.last_response.headers          # raw response headers (lower-cased)

Some self-hosted operational headers are not part of the Cloud response contract, so expect those fields to be None against api.openserp.org. credits is Cloud-specific.

Error handling

from openserp import OpenSERP, RateLimitError, CaptchaError, SERPError

client = OpenSERP(api_key="osk_live_xxx")

try:
    client.search(engine="google", text="openserp")
except RateLimitError:
    # slow down or queue the request
    ...
except CaptchaError:
    # inspect the upstream search failure and retry later
    ...
except SERPError as err:
    print(err.status, err.code, err.reason, err.request_id, err.retry_after)

Retry hook

The SDK does not apply a retry policy. Provide a hook when you want one:

import os, random, time
from openserp import OpenSERP, SERPError

RETRYABLE = {408, 429, 500, 502, 503, 504}
client: OpenSERP


def should_retry(err: Exception, attempt: int) -> bool:
    if attempt >= 3 or not isinstance(err, SERPError) or err.status not in RETRYABLE:
        return False
    wait = err.retry_after if err.retry_after is not None else min(2 ** attempt * 0.25, 8.0)
    time.sleep(wait + random.random() * 0.25)
    return True


client = OpenSERP(api_key=os.environ["OPENSERP_API_KEY"], retry=should_retry)
client.search(engine="google", text="openserp")

SERPError.retry_after is in seconds, read from Retry-After or retry_after. Cloud 503 engine_unavailable carries a 60-second delay. Structured errors remain available even when you request Markdown, text, or NDJSON. The SDK makes no automatic retries.

Use cases

  • AI grounding / RAG - feed top-N results into an LLM prompt (OpenAI, Anthropic, Ollama) for up-to-date answers.
  • LLM tool use - expose client.search as a tool to your agent.
  • SEO monitoring - daily rank tracking across multiple engines and regions, export to a DataFrame or Sheets.
  • Competitor analysis - weekly diff of top-10 results for a keyword set.
  • Data pipelines - stream SERPs to ClickHouse, BigQuery, or a DataFrame for NLP on snippets.

Quick SEO rank report with pandas:

import pandas as pd
from openserp import OpenSERP

client = OpenSERP()
keywords = ["openserp", "serp api", "google search api"]
frames = []

for keyword in keywords:
    resp = client.search(engine="google", text=keyword, region="US", limit=10)
    frame = resp.to_pandas()
    frame["keyword"] = keyword
    frames.append(frame)

pd.concat(frames, ignore_index=True).to_csv("rank-report.csv", index=False)