Configure catalog search tools for agents and LLMs
This page is how to wire search tools into Cursor, ChatGPT, Claude, and similar agents so they can (1) look up catalogs already in this registry and (2) discover installations that are not registered yet.
Query recipes (Google operators, Censys CenQL, FOFA as a Censys alternative, Shodan filters): discovery-search-tools.md. Agent checklist: agents/discover.md. This repository does not ship a production search API or MCP server — see when-to-use.md.
Two jobs, two tool stacks
| Job | Data the agent needs | Typical tools |
|---|---|---|
| Query catalogs already in the registry | DuckDB / Parquet / JSONL in this repo, or dateno-api | Local files, SQL, optional HTTP API |
| Discover catalogs not yet registered | Web search + internet maps + a live GET to confirm the site | Google (or CSE/Brave), Censys, Shodan, FOFA, GitHub (gh api forks and code search), URLScan, crt.sh, browser |
Do not mix them up. Searching Google for “CKAN Portugal” does not tell you whether dados.gov.pt is already in data/datasets/datasets.duckdb. Duplicate-check exports before adding YAML (agents/query.md).
What to install (minimum vs full)
| Setup | Query this registry | Discover new catalogs |
|---|---|---|
| Minimum (Cursor in this repo) | Files + DuckDB; follow llms.txt | Cursor web search + browser; paste queries from discovery-search-tools.md |
| Minimum (ChatGPT / Claude in the browser) | Upload llms.txt + a DuckDB/CSV extract, or use dateno-api | Built-in web search / browsing; paste the shared instructions |
| Full agent stack | Same as minimum, plus dateno-api if you need HTTP | Official Censys MCP, or FOFA API when Censys search is unavailable; Google Programmable Search or Brave/Tavily; optional Shodan / URLScan keys |
Start with the minimum. Add Censys MCP when Google stops listing sites (no inbound links, IP-only GeoServer, certificate names). If Censys search is not on the plan or OAuth is blocked, use FOFA with the same title / body / country filters (translation table).
Shared system instructions
Paste this into a Custom GPT, ChatGPT Project, Claude Project, Gemini Gem, Cursor user rule, or any other agent instruction field. Point the model at the live docs, not a stale copy.
You help maintain dataportals-registry (https://github.com/datenoio/dataportals-registry).
Read https://datenoio.github.io/dataportals-registry/llms.txt first.
Query existing catalogs from exports (DuckDB/Parquet/JSONL), never by walking data/entities/**/*.yaml.
If datasets.duckdb is locked, use full.parquet.
Discover missing catalogs with docs/agents/discover.md and docs/discovery-search-tools.md.
FOFA is the Censys alternative (title= / body= / host= / country=).
For open-source catalog software deployed by forking, search GitHub
forks and code search (code search skips forks): docs/discovery-search-tools.md#github.
Hunt patterns (harvest sources, university IRs, country indicators, named directories,
software tenant lists, subnational municipal GIS):
docs/discovery.md#hunt-patterns and docs/agents/improve.md.
Scope every hunt: one country, city, TLD, software.id, or named list URL. No internet-wide scans.
A hunt that finds 0 missing catalogs is complete — report that and stop.
For each candidate: duplicate-check hostname; confirm a public catalog UI or harvestable API;
set software.id only with two matching signals (else custom); do not invent uid.
Stop on HTTP 401/403. Do not follow login forms or guess API keys.
Register the catalog homepage, not a single dataset URL.
Prefer add-single --scheduled, then assign and validate-yaml --id {id}.
Do not mix ArcGIS Hub, Experience Builder, Web AppBuilder, Instant Apps, Dashboards,
StoryMaps, and Server on the same org unless they are distinct public catalogs.
Search-tool config: docs/discovery-agent-tools.md
Platform fingerprints: docs/discovery-opendata.md, discovery-geoportals.md,
discovery-scientific.md, discovery-metadata.md, discovery-indicators.md.
Software ID map: docs/software-index.md.
Keep secrets out of these instructions. Put API keys in the client’s env / MCP config / GPT Action auth, never in git.
Secrets and environment variables
Do not commit keys. Use the user environment, Cursor MCP env, ChatGPT Action authentication, or a local .env that is gitignored.
| Variable | Tool | Where to get it |
|---|---|---|
GOOGLE_CSE_API_KEY | Google Custom Search JSON API | Google Cloud Console → APIs & Services → Credentials |
GOOGLE_CSE_CX | Same (search engine id) | Programmable Search Engine → Setup → Basics |
BRAVE_API_KEY | Brave Search API (agent-friendly Google alternative) | Brave Search API |
TAVILY_API_KEY | Tavily Search (LLM-oriented) | Tavily |
CENSYS_PAT | Censys Platform API (header auth) | Censys Platform → Account → Personal access token |
CENSYS_ORG_ID | Censys org header | Censys Platform organization settings |
SHODAN_API_KEY | Shodan | account.shodan.io |
FOFA_EMAIL / FOFA_KEY | FOFA | en.fofa.info account |
URLSCAN_API_KEY | urlscan.io Search API | urlscan.io/user/signup |
crt.sh needs no key. Cursor’s built-in web search and ChatGPT browsing need no extra key.
Cursor
Cursor already has this repository’s AGENTS.md and docs/agents/. Open the dataportals-registry folder as the workspace so the agent can read exports and write YAML.
Built-in tools (no API key)
- Web search — ask the agent to run the Google queries from discovery-search-tools.md (or Bing/DuckDuckGo if Google is blocked).
- Browser — open a candidate URL, confirm it is a catalog, copy the homepage
link. - Terminal — duplicate-check DuckDB or
full.parquetif DuckDB is locked;add-single --scheduled;validate-yaml --id. - Rules — project rule
.cursor/rules/catalog-discovery.mdcattaches when the task is discovery. User rules can paste the shared instructions if you want them in every chat.
Reuse a prior hunt transcript for the same software.id before starting a second pass. Example prompt:
Discover CKAN open data portals in Portugal that are not in this registry.
Follow docs/agents/discover.md. Duplicate-check datasets.duckdb first
(or full.parquet if DuckDB is locked).
Use Google queries from docs/discovery-opendata.md, then confirm
/api/3/action/status_show. If Censys MCP is not connected, use FOFA
body="name=\"generator\" content=\"ckan" && country="PT"
(if body= is not on the plan: title="CKAN" && country="PT").
See docs/discovery-opendata.md#ckan.
Add verified finds with add-single --scheduled.
MCP in Cursor
Cursor Settings → Tools & MCP → New MCP server, or edit ~/.cursor/mcp.json (user-wide) or .cursor/mcp.json in a project.
- Prefer user-wide
~/.cursor/mcp.jsonfor API keys so they are not committed. - If you add a project
.cursor/mcp.json, store only public URLs (for example the Censys OAuth MCP). Never put a personal access token in the repo. - After saving, reload MCP. CLI:
agent mcp list(see Cursor MCP docs).
Censys (official, recommended for discovery beyond Google)
OAuth (preferred). Cursor opens a consent page; pick your Censys organization. Calls count against Censys credits.
{
"mcpServers": {
"censys-platform": {
"url": "https://mcp.platform.censys.io/platform/mcp/"
}
}
}
Header auth (only in user mcp.json, not in git):
{
"mcpServers": {
"censys-platform": {
"url": "https://mcp.platform.censys.io/platform/mcp/",
"headers": {
"X-Organization-ID": "your-organization-id",
"Authorization": "Bearer your-censys-personal-access-token"
}
}
}
}
Official docs: Platform MCP Server. For catalog discovery, ask the agent to use web property search (generate_and_search_query / CenQL on web.endpoints.http.html_title and web.names), not discover_attack_surface or the Adversary Investigation MCP. Those tools are for security hunting, not registry work.
Shodan (optional, third-party MCP)
There is no official Shodan MCP. Community servers exist (search GitHub for shodan-mcp). Review the code before installing. Typical pattern in user mcp.json:
{
"mcpServers": {
"shodan": {
"command": "npx",
"args": ["-y", "mcp-shodan"],
"env": {
"SHODAN_API_KEY": "your-shodan-api-key"
}
}
}
}
Package names vary. If you do not want a third-party MCP, have Cursor run the Shodan CLI in the terminal instead.
FOFA (Censys alternative; no official MCP)
There is no official FOFA MCP. Do not add an unreviewed third-party FOFA MCP for this registry. Have Cursor run the FOFA API in the terminal with FOFA_EMAIL / FOFA_KEY from the user environment, then probe live hosts the same way as Censys hits.
Do not add a “Google scrape” MCP. Use Cursor web search, or configure Programmable Search and let the agent curl it with keys from the environment.
Cursor Cloud Agents
Cloud agents do not automatically inherit your laptop ~/.cursor/mcp.json or local API keys. Give them the GitHub repo, llms.txt, and tell them to use public web search. Attach Censys only if you add a project MCP with OAuth (no PAT in the repo) or run discovery locally. FOFA keys stay in the user environment — run FOFA hunts locally, not in cloud agents.
ChatGPT app
ChatGPT (chatgpt.com and the desktop/mobile app) cannot see your local data/datasets/ tree. Give it either (a) web search + this project’s published docs, or (b) an HTTP API / uploaded extract.
Everyday chat (web search)
On a paid plan, enable web search (or browsing) in the composer. Start the thread with:
Read https://datenoio.github.io/dataportals-registry/llms.txt
and https://datenoio.github.io/dataportals-registry/docs/discovery-search-tools
Find GeoNetwork catalogs in Czechia not obviously on the GeoNetwork gallery.
Return name, URL, proposed software.id, and why you think it is a catalog.
Do not invent registry uids. I will duplicate-check in the repo.
Paste DuckDB hits if you already searched locally, so ChatGPT does not propose duplicates.
ChatGPT Projects
Create a Project. Add the shared instructions above. Upload (or link):
llms.txtdocs/agents/discover.mddocs/discovery-search-tools.md- Optionally a CSV/JSONL slice of catalogs for one country (
id,name,link,software)
Projects persist that context across chats. Still use web search for live discovery.
Custom GPT
- chatgpt.com/gpts/editor → Create.
- Instructions: paste the shared system instructions.
- Knowledge: upload the same files as for Projects. Keep them updated when docs change.
- Capabilities: enable Web Search. Do not enable image generation for this job.
- Actions (optional): add an OpenAPI action for Google CSE, Censys, Shodan, or FOFA (examples below). Store keys in the GPT’s authentication, not in the spec. For FOFA, the HTTP API is simpler from Cursor’s terminal than as a GPT Action (email + key query params).
- Conversation starters: “Find CKAN portals in Kenya”, “GeoNetwork in the Balkans”, “Dataverse installations missing from the registry”, “Use FOFA instead of Censys for ODWeb catalogs in China”.
Custom GPTs cannot run builder.py. The GPT should return a candidate table; you (or Cursor) write YAML in the repo.
Developer Mode and MCP connectors
ChatGPT can attach remote MCP servers (HTTPS). Local stdio MCP (a command on your laptop) is not supported.
- Enable Developer Mode (Settings → Apps / Advanced — paid plans; write-capable MCP is limited on some consumer plans).
- Add a connector with URL
https://mcp.platform.censys.io/platform/mcp/and complete Censys OAuth. - In a chat, attach the Censys app/connector from the + menu and ask for web-property searches scoped by country.
OpenAI help: Developer mode and MCP apps. ChatGPT Apps in the directory are vendor-reviewed; Censys is added as a custom connector, not as an OpenAI-built plugin (plugins were retired).
Custom GPT Actions (OpenAPI)
Example Google Custom Search action (replace {cx} or pass cx as a parameter). Authentication: API key query param key.
openapi: 3.0.1
info:
title: Google Custom Search
version: 1.0.0
servers:
- url: https://www.googleapis.com/customsearch
paths:
/v1:
get:
operationId: googleCustomSearch
summary: Search the web with a Programmable Search Engine
parameters:
- name: q
in: query
required: true
schema: { type: string }
- name: cx
in: query
required: true
schema: { type: string }
- name: num
in: query
schema: { type: integer, default: 10, maximum: 10 }
responses:
"200":
description: Search results
In the GPT Action auth UI, choose API key, header or query key, and paste GOOGLE_CSE_API_KEY. Put your engine id in instructions: “Always pass cx=YOUR_CX.”
For Censys, prefer the official MCP connector over hand-written OpenAPI; the Platform schema is large and versioned at api.platform.censys.io.
Claude (Desktop, Claude Code, Projects)
Claude.ai Projects: same pattern as ChatGPT Projects — paste shared instructions, add knowledge files / links to llms.txt and the discovery guides, enable web search.
Claude Desktop MCP (~/Library/Application Support/Claude/claude_desktop_config.json on macOS):
{
"mcpServers": {
"censys-platform": {
"url": "https://mcp.platform.censys.io/platform/mcp/"
}
}
}
Claude Code:
claude mcp add censys-platform https://mcp.platform.censys.io/platform/mcp/
Then open this repository so Claude Code can duplicate-check DuckDB and add YAML.
Other LLM clients
| Client | How to attach this registry | How to attach search tools |
|---|---|---|
| VS Code Copilot | Open the repo; Copilot reads AGENTS.md if present | .vscode/mcp.json with the same Censys URL |
| Continue.dev | Index the repo | mcpServers in Continue config (Censys example) |
| Windsurf / Cline | Open the repo | User MCP JSON, same schema as Cursor |
| Gemini Gems | Paste shared instructions; link llms.txt | Built-in Google Search grounding (strong for Google dorks); no Censys/FOFA unless you add an API call in a larger app |
| Perplexity | Link the docs site | Built-in web search; paste candidate URLs back into Cursor for YAML |
| Open WebUI / local models | RAG over docs/ + llms.txt | Add CSE/Brave/Censys/FOFA as HTTP tools in the tool registry |
Configure each search tool
Query syntax lives in discovery-search-tools.md. This section is accounts, APIs, and how agents should call them.
Google (and Bing / DuckDuckGo)
For Cursor and ChatGPT, built-in web search is enough for most hunts. Tell the agent the exact query string ("Powered by CKAN" inurl:/dataset site:.pt).
Do not have agents scrape google.com/search HTML (ToS, captchas, brittle). If you need stable JSON for automation, use Programmable Search, Brave, or Tavily.
Google Programmable Search JSON API
- Create a Programmable Search Engine.
- Under Setup → Basics, turn Search the entire web on (otherwise you only search sites you listed).
- Copy the Search engine ID (
cx). - In Google Cloud, enable Custom Search API and create an API key. Restrict the key to that API.
- Free quota is small (on the order of 100 queries/day unless you enable billing). Scope hunts tightly.
curl -sS "https://www.googleapis.com/customsearch/v1" \
--get \
--data-urlencode "key=$GOOGLE_CSE_API_KEY" \
--data-urlencode "cx=$GOOGLE_CSE_CX" \
--data-urlencode "q=Powered by CKAN inurl:/dataset site:.pt" \
--data-urlencode "num=10"
Agent rule: one query per country×software, num≤10, then a human or browser confirms the catalog root. Parse items[].link and items[].title only.
Brave Search API
Useful when you want JSON without building a CSE. Header X-Subscription-Token: $BRAVE_API_KEY.
curl -sS "https://api.search.brave.com/res/v1/web/search" \
-H "X-Subscription-Token: $BRAVE_API_KEY" \
--get \
--data-urlencode "q=intitle:GeoNetwork site:.cz"
Tavily
LLM-oriented search (POST https://api.tavily.com/search with api_key and query). Good inside custom agents; still duplicate-check this registry afterward.
Censys
- Create a Censys Platform account. API search is a paid/credit capability on most plans; Free may only allow lookups, not search — check Get started with Censys APIs.
- Grant your user the API Access role for the organization.
- Prefer OAuth MCP in Cursor / Claude / ChatGPT Developer Mode.
- For raw HTTP, create a personal access token. Base URL:
https://api.platform.censys.io/v3/global/.
curl -sS "https://api.platform.censys.io/v3/global/search/query" \
-H "Authorization: Bearer $CENSYS_PAT" \
-H "X-Organization-ID: $CENSYS_ORG_ID" \
-H "Content-Type: application/json" \
-d '{
"query": "web.location.country_code = \"PT\" and web.endpoints.http.body: \"ckan-footer-logo\"",
"page_size": 25
}'
Agent rules:
- Search web properties for catalog UIs; use hosts for GeoServer/ArcGIS product hits; use certificates for
opendata.names. page_size≤ 25 per call. Do not page through the entire internet.- Convert hostnames to
https://{name}/, then probe. Never setlinkto a bare IP. - MCP helpers
generate_queryandvalidate_censys_queryare useful;investigate_hostis optional and expensive — skip it unless you need to disambiguate one IP. - If Censys search returns no credits / plan error, switch to FOFA instead of inventing hosts.
Legacy Search (search.censys.io / v1 API) is a different product. New integrations should use Platform CenQL (query language).
Shodan
- Create an account and copy the API key from account.shodan.io.
- Search filters consume query credits on most plans. Use
shodan countbeforeshodan search.
pip install -U --user shodan
shodan init "$SHODAN_API_KEY"
shodan count 'http.title:"CKAN" country:PT'
shodan search --limit 20 'http.title:"CKAN" country:PT'
REST:
curl -sS "https://api.shodan.io/shodan/host/search" \
--get \
--data-urlencode "key=$SHODAN_API_KEY" \
--data-urlencode "query=http.title:\"GeoNetwork\" country:CZ"
Agent rules: prefer hostname / SSL CN in the banner over ip_str for link. Skip results with no HTTP title. Do not run shodan scan (active scanning) for this registry.
FOFA
FOFA is the Censys alternative for this registry: same job (title / body / country / hostname internet map), different syntax and API. Query translation: discovery-search-tools.md. Prefer Censys MCP when it is connected; use FOFA when Censys search is unavailable, or when the hunt is East Asia.
- Create an account at en.fofa.info. Copy the API email and key from the account page. Do not commit them.
- There is no official FOFA MCP. Run the HTTP API from Cursor’s terminal (same pattern as Shodan). Do not install an unreviewed third-party FOFA MCP for this work.
- Encode the query as Base64 (
qbase64). Agent must not log the key. - Request
fields=host,ip,port,protocol,title,domain. Preferhost/domainforlink.size≤ 20 per call (plan caps are often 20–100). Do not page through the entire internet. - Free and low plans often cannot search
body=orheader=over the API. Fall back totitle=,host=,domain=,app=, andcert=, or paste results from the FOFA web UI.
python - <<'PY'
import base64, json, os, urllib.parse, urllib.request
q = 'body="name=\\"generator\\" content=\\"ckan" && country="JP"'
qb = base64.b64encode(q.encode()).decode()
url = (
"https://fofa.info/api/v1/search/all?"
+ urllib.parse.urlencode({
"email": os.environ["FOFA_EMAIL"],
"key": os.environ["FOFA_KEY"],
"qbase64": qb,
"fields": "host,ip,port,protocol,title,domain",
"size": 20,
})
)
req = urllib.request.Request(
url,
headers={"User-Agent": "dataportals-registry-hunt/1.0"},
)
data = json.loads(urllib.request.urlopen(req, timeout=30).read().decode())
if data.get("error"):
raise SystemExit(data)
for row in data.get("results") or []:
print(row)
PY
Example queries an agent can paste into the FOFA UI or the script above:
body="name=\"generator\" content=\"ckan" && country="PT"
body="ckan-footer-logo" && country="JP"
body="OpenDataSoft" && country="BE"
title="GeoNetwork" && country="CZ"
domain="opendatasoft.com"
app="GeoServer" && country="ID"
body="/odweb/" && title="数据开放"
title="Dataverse"
body="hyrax"
host="www2.wagmap.jp"
Agent rules:
- Search titles and body snippets for catalog UIs; use
app=for GeoServer-class products; usehost=/domain=/cert=for SaaS hostname patterns. - Convert
hosttohttps://{host}/, then probe. Never setlinkto a bare IP. Deduplicate 80/443 before probing. - A FOFA hit is a lead, not a catalog. Duplicate-check exports, then confirm the public UI.
- If
erroris true orresultsis empty becausebody=is not on the plan, retry withtitle=/host=and say so in the hunt report.
ChatGPT / Claude without the API: open en.fofa.info, paste the translated query, and return host + title for duplicate-check in the repo. Do not invent hosts from memory.
urlscan.io
Good for recently crawled Hub and OpenDataSoft sites.
curl -sS "https://urlscan.io/api/v1/search/" \
-H "api-key: $URLSCAN_API_KEY" \
--get \
--data-urlencode "q=page.title:CKAN" \
--data-urlencode "size=20"
Use page.url / page.domain from results. Search-only is enough; do not mass-submit live scans of government sites.
crt.sh (no key)
curl -sS "https://crt.sh/?q=%.opendata.pt&output=json"
Treat names as leads. Resolve HTTPS and confirm a catalog UI.
Which tool the agent should call
User names a country / city / software
│
├─► DuckDB / Parquet duplicate-check
│
├─► Vendor lists (CKAN ecosystem, GeoNetwork gallery, Dataverse JSON, …)
│
├─► Web search (Cursor/ChatGPT/CSE/Brave) with platform queries
│
├─► If thin results: Censys web properties (MCP or API)
│ or FOFA title/body/country (API) when Censys is not configured
│ optional Shodan / URLScan / crt.sh
│
└─► Browser or GET: confirm catalog + software probe
then add-single --scheduled (in the git workspace)
If the agent has no Censys/FOFA/Shodan credentials, it must still finish the Google + vendor-list path and report “Censys/FOFA not configured” rather than inventing hosts.
Example: Cursor vs ChatGPT for the same hunt
Cursor (repo open, Censys MCP connected)
Find unregistered OpenDataSoft portals in Belgium.
1) SQL duplicate-check on link like '%opendatasoft%' and '%belgium%' / '.be'
2) Google: inurl:/explore site:.be OpenDataSoft
3) Censys MCP: web.names: ".be" and web.endpoints.http.body: "OpenDataSoft"
(FOFA alternative: body="OpenDataSoft" && country="BE")
4) Confirm /api/explore/v2.1/catalog/datasets
5) add-single --scheduled for new hosts only
Cursor (repo open, FOFA API instead of Censys)
Find unregistered CKAN portals in Japan. Censys is not configured; use FOFA.
1) SQL duplicate-check software.id = 'ckan' and link like '%.jp%'
2) Google: "ckan-footer-logo" site:.jp
3) FOFA: body="name=\"generator\" content=\"ckan" && country="JP"
(if body= is not on the plan: title="CKAN" && country="JP")
GitHub: "ckan.site_url = https://" — see docs/discovery-opendata.md#ckan
4) Confirm /api/3/action/status_show
5) add-single --scheduled for new hosts only
ChatGPT (no local repo, web search + docs)
Follow https://datenoio.github.io/dataportals-registry/docs/discovery-opendata
Search for OpenDataSoft portals in Belgium. Return a markdown table:
name | url | evidence | likely duplicate of data.gov.be?
I will check the registry myself. Do not invent uids.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Agent walks thousands of YAML files | Ignored llms.txt | Point it at exports; repeat agents/query.md |
| Censys MCP does nothing | No API role / OAuth not completed / Free plan has no search | Finish consent; check credits; fall back to FOFA or Google |
401 from Censys API | Missing PAT or org header | Use MCP OAuth, or set both Authorization and X-Organization-ID |
FOFA HTTP 403 | Request has no User-Agent | Send one; the API rejects a bare urllib client |
FOFA error: true / empty results | Missing key, body= not on the plan, or query too broad | Check FOFA_EMAIL/FOFA_KEY; retry title=/host=; add country= |
Agent sets link to an IP | Censys/Shodan/FOFA host record | Use certificate / web-property / FOFA host name; confirm HTTPS vhost |
| Google CSE returns only your site | Entire-web toggle off | Enable “Search the entire web” on the engine |
CSE 403 | API not enabled or key restricted | Enable Custom Search API; relax key restrictions for the agent runtime |
| Shodan empty / error | No credits or query too broad | shodan info; add country: and a title filter |
| ChatGPT proposes zenodo.org / ckan.org | No duplicate-check | Paste existing link values or a country CSV into the thread |
| Keys leaked in a PR | mcp.json or GPT spec committed | Rotate the key; move secrets to user config |