Skip to main content

Discovering domain scientific repositories

Biodiversity, facility, crop, chemistry, and earth-system repositories (catalog_type: Scientific data repository). Institutional IRs: discovery-scientific.md. Search-engine syntax (Google, Censys, and FOFA as a Censys alternative): discovery-search-tools.md. Harvest: harvest-scientific-domain.md.

High-count domain stacks with their own recipes: IPT, Symbiota, Biodiv, BIMS, THREDDS, ERDDAP, FROST-Server, Breedbase, Tripal, VEuPathDB, MassBank, ioChem-BD, ESGF, ALA, BirdMap Africa, SciCat, InterMine, GRIN-Global, PlutoF, JGI Genome Portal, cBioPortal, ESA Science Archive, CLLD, TalkBank, Pathway Tools, IBDC.

One portal / node = one registry record. Do not add gene pages, occurrences, or ESGF data nodes as extra catalogs.

Ocean and earth directories (not software): ODIS, PANGAEA harvest sources, WMO WIS2 GDC, DataONE member nodes, CLARIN / VLO. Treat each as a named-list hunt (discovery.md); skip org homepages and hijacked hosts.

GBIF IPT (ipt)​

Integrated Publishing Toolkit for biodiversity data. List: gbif.org/ipt.

Confirm: /rss.do, /inventory/dataset, or the IPT homepage with installation name.

footer.ftl serves GBIF-2015-standard-ipt.png (431 hosts in September 2026, including ipt.nature.ca). body="Integrated Publishing Toolkit" matched 485. Pages that name the toolkit without the logo still match the phrase, so keep both.

ToolQuery
Google"Integrated Publishing Toolkit" IPT GBIF
Googleinurl:/ipt "GBIF"
Censysweb.endpoints.http.body: "GBIF-2015-standard-ipt.png"
FOFAbody="GBIF-2015-standard-ipt.png"
Censysweb.endpoints.http.body: "Integrated Publishing Toolkit"
FOFAbody="Integrated Publishing Toolkit"

Prefer GBIF’s official installation list, then fill gaps with search.

Symbiota (symbiota)​

Open-source biodiversity collections CMS. Official portal directory: symbiota.org/symbiota-portals. Docs: docs.symbiota.org.

Theme-based portals (SEINet, MyCoPortal, CCH2, Ecdysis, and others) publish specimen occurrences, images, checklists, and Darwin Core datasets. Register one catalog per portal, not per collection or GBIF IPT mirror. Use software.id: symbiota.

Confirm: public collection search (/collections/index.php or /portal/collections/) and/or dataset RSS at /collections/datasets/rsshandler.php. Page signals include “Symbiota”, collid=, and “Search Collections”. Skip login-only portals and the vendor homepage.

head_template.php links symbiota/header.css (92 hosts in September 2026, including mycoportal.org). body="Symbiota" matched 256. Themed portals can drop that stylesheet path, so keep both.

ToolQuery
Google"Powered by Symbiota" OR "Symbiota portal" (collections OR occurrences) -site:symbiota.org -site:github.com
Googleinurl:/collections/datasets/rsshandler.php
Censysweb.endpoints.http.body: "symbiota/header.css"
FOFAbody="symbiota/header.css"
Censysweb.endpoints.http.body: "Symbiota"
FOFAbody="Symbiota"

Biodiv (biodiv)​

Strand Life Sciences Biodiversity Informatics Platform. Product: strandls.com/biodiversity-informatics. UI source: strandls/biodiv-ui. Flagship: indiabiodiversity.org. Other installations include the Bhutan Biodiversity Portal, WIKTROP, and related regional portals. Distinct from Symbiota (symbiota), Specify (specify), and a GeoServer that only publishes a Biodiv workspace.

Signals: footer text Biodiversity Informatics Platform; technology partner Strand Life Sciences; public paths /dataset/list, /observation/list, /species/list, /document/list.

Confirm: GET the portal home and match the platform footer. One record per portal installation, not per species page, observation, or user group. Do not retag a /geoserver endpoint on the same host as biodiv. Skip the Strand marketing page.

ToolQuery
Google"Biodiversity Informatics Platform" (portal OR observations OR species) -site:github.com
Google"Technology Partner" "Strand Life Sciences" biodiversity
Censysweb.endpoints.http.body: "Biodiversity Informatics Platform"
FOFAbody="Biodiversity Informatics Platform"

BIMS (bims)​

Kartoza Biodiversity Information Management System. Product page: bims.kartoza.com. Source: kartoza/django-bims. Published portals include FBIS South Africa, FBIS Africa, SANParks BIMS, RBIS Rwanda, Kafue Flats, and FADA. Distinct from Symbiota (symbiota), Biodiv (biodiv), and the Jordan ArcGIS host bims.rscn.org.jo.

Signals: inline githubRepo = 'kartoza/django-bims'; source-reference list at /source-references/; JSON at /api/module-summary/ and /api/layer/.

Confirm: GET the portal home and match kartoza/django-bims. One record per portal, not per occurrence, taxon, or map tile. Skip the vendor homepage bims.kartoza.com. Skip FIPbio while the site says it is under development, and skip ORBIS while it is offline. Login-only downloads still count as one catalog when the home page is the BIMS portal.

ToolQuery
Google"kartoza/django-bims" (biodiversity OR freshwater OR occurrence)
Googleinurl:/source-references "Occurrence Records"
Censysweb.endpoints.http.body: "kartoza/django-bims"
FOFAbody="kartoza/django-bims"

THREDDS (thredds)​

Scientific data servers (often climate/ocean). Confirm: /thredds/catalog.html or /thredds/catalog.xml.

catalog.html links tds.css (356 hosts in September 2026, including thredds-su.ipsl.fr). body="THREDDS" matched 878, including documentation that only mentions the server. Keep both.

ToolQuery
Googleinurl:/thredds/catalog.html
Google"THREDDS Data Server" catalog
Censysweb.endpoints.http.body: "tds.css"
FOFAbody="tds.css"
Censysweb.endpoints.http.body: "THREDDS"
FOFAbody="THREDDS"
Shodanhttp.html:"THREDDS Data Server"
Shodanhttp.html:"THREDDS Data Server"

ERDDAP (erddap)​

NOAA-style tabular/gridded data server. Confirm: /erddap/index.html or /erddap/info/index.json.

ToolQuery
Googleinurl:/erddap "ERDDAP"
Censysweb.endpoints.http.body: "ERDDAP"
FOFAbody="ERDDAP"

FROST-Server (frostserver)​

Fraunhofer IOSB open-source OGC SensorThings API server. Product: FROST-Server; source: github.com/FraunhoferIOSB/FROST-Server. Register one public SensorThings catalog per independently operated instance. Typical catalog_type is Scientific data repository (urban IoT and groundwater stations are still a Things/Datastreams catalog, not a map viewer). Distinct from THREDDS, ERDDAP, and a SensorThings endpoint that is only a download option on another catalog.

Confirm: JSON at /v1.1/, /v1.0/, or /FROST-Server/v1.1/ listing Things / Datastreams / Locations, or the default HTML start page titled Start Page with heading FROST-Server. Things?$top=1&$count=true returns @iot.count. Skip Fraunhofer k8s demos, SensorUp scratchpads, login-walled hydrometry, and hosts with only a single Thing.

index.html uses the heading FROST-Server (101 hosts in September 2026, including dgw.ve.ismar.cnr.it).

ToolQuery
Google"FROST-Server" (SensorThings OR Things OR Datastreams) -site:github.com
Googleinurl:/FROST-Server/ "Start Page"
Censysweb.endpoints.http.body: "FROST-Server"
FOFAbody="FROST-Server"

52°North SOS (52northsos)​

52°North open-source OGC Sensor Observation Service (SOS 1.0/2.0). Product: 52°North SOS; source: github.com/52North/SOS. Register one public SOS catalog per independently operated instance. Typical catalog_type is Scientific data repository (observation offerings are a sensor data catalog, not a map viewer — same rule as FROST-Server). Distinct from FROST-Server (SensorThings JSON) and from an SOS endpoint that is only a download option on another catalog.

Confirm: XML Capabilities at /52n-sos-webapp/sos?service=SOS&request=GetCapabilities (or a custom mount) listing sos:Contents offerings, or the default webapp page titled 52°North Sensor Observation Service (HTML-escaped as 52°North). Page bodies commonly link /52n-sos-webapp/. Skip *.52north.org vendor demos, bare-IP test boxes, and login-walled instances.

header.jsp loads static/css/52n.css (2 hosts in September 2026, including a page titled 52°North Sensor Observation Service). body="52n-sos-webapp" matched 3, including geomon.geologie.ac.at. Custom mounts drop the stylesheet, so keep both.

ToolQuery
Google"52n-sos-webapp" -site:github.com -site:52north.org
Googleintitle:"52°North Sensor Observation Service"
Censysweb.endpoints.http.body: "static/css/52n.css"
FOFAbody="static/css/52n.css"
Censysweb.endpoints.http.body: "52n-sos-webapp"
FOFAbody="52n-sos-webapp"

OPeNDAP (opendap)​

Remote subsetting protocol and server ecosystem. Site: opendap.org. Use opendap for a public OPeNDAP catalog whose server implementation is not identified as Hyrax, Pydap, THREDDS, or ERDDAP. Do not register OPeNDAP only as a download option on a THREDDS (thredds) or ERDDAP (erddap) catalog.

Confirm: GET a DAP catalog or directory listing and verify that it exposes multiple datasets.

ToolQuery
Google"OPeNDAP" ("catalog.xml" OR DODS) -Hyrax -site:opendap.org -site:github.com
Censysweb.endpoints.http.body: "OPeNDAP"
FOFAbody="OPeNDAP"

OPeNDAP Hyrax (opendaphyrax)​

The OPeNDAP 4 Data Server, unrelated to the Samvera repository product that uses software.id: hyrax. Register one public Hyrax server per independently operated dataset catalog. Prefer thredds or erddap when Hyrax is only an alternate access service for one of those catalogs.

Confirm: the directory page title starts with OPeNDAP Hyrax: Contents of, the footer reports Hyrax (version), or /opendap/catalog.xml returns the server catalog. threddsCatalogPresentation.xsl writes OPeNDAP Hyrax into that directory page (52 hosts in September 2026, including opendap.mars.le.isac.cnr.it).

ToolQuery
Google"OPeNDAP Hyrax: Contents of" -site:opendap.org -site:github.com
Googleinurl:/opendap/ "Hyrax development sponsored by"
Censysweb.endpoints.http.body: "OPeNDAP Hyrax"
FOFAbody="OPeNDAP Hyrax"

DataONE (dataone)​

Earth-science member-node network. Site: dataone.org. Prefer the member node catalog URL, not every harvested dataset.

Confirm: GET the member-node home. Duplicate-check before adding nodes already in re3data / this registry.

MetacatUI’s src/index.html comments “configuration file for MetacatUI” (105 hosts in September 2026, including opc.dataone.org). body="DataONE" matched 2,274 hosts, including pages that only mention the project. Keep both.

ToolQuery
Google"DataONE" ("member node" OR MN) repository
Censysweb.endpoints.http.body: "configuration file for MetacatUI"
FOFAbody="configuration file for MetacatUI"
Censysweb.endpoints.http.body: "DataONE"
FOFAbody="DataONE"

MOLGENIS (molgenis)​

Configurable FAIR scientific data platform used for research catalogues, biobank directories, and registries. Site and public-instance list: molgenis.org/tools.

Signals: current EMX2 catalogues say “Created with MOLGENIS” and expose /api/graphql, /api/rdf, or /<database>/api/csv/<table>. Legacy installations use molgenis.do, often below a project path. A generic University of Groningen page is not sufficient evidence.

Confirm: GET the public catalogue/search UI and one read-only API surface when available. Register one public catalogue or registry per installation, not each database table, cohort, biobank, or variable.

FooterComponent.vue renders the heading Created with MOLGENIS (6 hosts in September 2026, including catalogue.hdsu.nl). body="MOLGENIS" matched 205 hosts, including directory.canserv.eu. Themed catalogues drop the heading, so keep both.

ToolQuery
Google"Created with MOLGENIS" (catalogue OR registry OR collections)
Googleinurl:molgenis.do (data OR database OR repository)
Censysweb.endpoints.http.body: "Created with MOLGENIS"
FOFAbody="Created with MOLGENIS"
Censysweb.endpoints.http.body: "MOLGENIS"
FOFAbody="MOLGENIS"

BEXIS2 (bexis2)​

Open-source research data management and repository platform for structured and unstructured data. Site: bexis2.uni-jena.de. Public instances can be heavily themed, so confirm both the application route and assets.

Signals: root redirects to /home/Start; BEXIS2 name or footer; /Content/ assets and ASP.NET application; read-only /api/dataset, /api/metadata/{id}, or /api/data/{id} endpoints.

Confirm: GET the public search and a read-only dataset API. Register one BEXIS2 installation, not each project, metadata schema, or dataset.

_Layout.cshtml renders bundles/bexis (35 hosts in September 2026, including bexis.ufz.de). body="BEXIS2" matched 64. Themed shells can drop the bundle path, so keep both.

ToolQuery
Google"BEXIS2" (repository OR "research data") -site:github.com
Googleinurl:/home/Start BEXIS
Censysweb.endpoints.http.body: "bundles/bexis"
FOFAbody="bundles/bexis"
Censysweb.endpoints.http.body: "BEXIS2"
FOFAbody="BEXIS2"

Diversity Workbench (diversityworkbench)​

Modular bio- and geodiversity research-data environment maintained by the SNSB IT Center and partners. Site: diversityworkbench.net. Deployments often publish through project-specific web interfaces or the SNSB BioCASe/RDF pipeline rather than a uniform DWB homepage.

Signals: explicit “Diversity Workbench” or “DWB” attribution; modules such as DiversityCollection, DiversityDescriptions, DiversityTaxonNames, or DiversityProjects; BioCASe/ABCD publication backed by a DWB cache database; id.snsb.info RDF identifiers.

Confirm: require an explicit DWB attribution from the portal or its operator. One public catalog or publication pipeline = one registry record; do not register every DWB module, project database, BioCASe datasource, or occurrence record separately.

ToolQuery
Google"Diversity Workbench" (database OR repository OR data)
Google"DiversityCollection" (BioCASe OR RDF OR dataset)
Censysweb.endpoints.http.body: "Diversity Workbench"
FOFAbody="Diversity Workbench"

Greenstone (greenstone)​

Open-source digital-library collection software from the University of Waikato. The official examples page lists independent public libraries built with Greenstone 2 and Greenstone 3.

Signals: Greenstone 3 uses /greenstone3/<library>/collection/<collection>/..., xmlns:gs3, greenstone.org/gs3, or an /greenstone3/oaiserver endpoint. Legacy Greenstone 2 installations commonly use /greenstone/cgi-bin/library.cgi.

Confirm: GET the library home and, when enabled, the OAI Identify response. Register one independently operated library/catalog, not every collection inside it.

ToolQuery
Googleinurl:/greenstone3/library/collection
Googleinurl:/greenstone/cgi-bin/library.cgi (collection OR library)
Censysweb.endpoints.http.body: "greenstone.org/gs3"
FOFAbody="greenstone.org/gs3"

VIVO (vivo)​

Open-source semantic web platform for research discovery. Site: vivo.lyrasis.org. VIVO normally catalogs people and research activity, but some deployments also index datasets and repository records.

Signals: VIVO attribution together with vitro/vivo assets; RDF entity pages; faceted classes for datasets or data records; /api/sparqlQuery, /reconcile, or a configured Data Distribution API.

Confirm: the installation must expose a public dataset or research-object catalog. Do not register a profiles-only VIVO deployment, individual researcher pages, or the project website itself.

footer.ftl links http://vivoweb.org (386 hosts in September 2026, including experts.colorado.edu). body="vitro" && body="VIVO" is not usable: it matched 21,978 hosts, including www.jennio-bio.com.

ToolQuery
Google"Powered by VIVO" (dataset OR repository OR data)
Google"VIVO" "research data" (search OR repository)
Censysweb.endpoints.http.body: "vivoweb.org"
FOFAbody="vivoweb.org"

CWIS (cwis)​

The Collection Workflow Integration System is an open-source metadata collection and digital-library platform from Internet Scout. Site: scout.wisc.edu/cwis.

Signals: “Powered by CWIS” or the expanded product name; CWIS PHP assets; a root OAI-PMH response using ?verb=Identify; qualified Dublin Core and RSS links.

Confirm: GET the public resource search and OAI Identify response. Avoid the unrelated Chest Wall Injury Society acronym. One CWIS collection site = one registry record. INFOMED tesis.sld.cu and the Artemisa provincial node are verified CWIS catalogs (CWIS JavaScript and a scout.wisc.edu/cwis credit).

ToolQuery
Google"Powered by CWIS" (repository OR collection OR resources)
Google"Collection Workflow Integration System" -site:scout.wisc.edu
Censysweb.endpoints.http.body: "Powered by CWIS"
FOFAbody="Powered by CWIS"

Galaxy (galaxy)​

Usable-analysis platform that sometimes publishes public data libraries. Site: usegalaxy.org. Register public Galaxy instances with a data library / shared histories catalog, not every private analysis server.

Confirm: GET the instance and a public data-library or toolshed-adjacent dataset listing.

templates/js-app.mako includes the noscript heading “Javascript Required for Galaxy” (987 hosts in September 2026, including www.mutarget.com). body="usegalaxy" only matches the public usegalaxy.org family. Forks of galaxyproject/galaxy are the software, not a portal list.

ToolQuery
Google"Galaxy" ("data libraries" OR usegalaxy) -site:galaxyproject.org
Censysweb.endpoints.http.body: "Javascript Required for Galaxy"
FOFAbody="Javascript Required for Galaxy"
Censysweb.endpoints.http.body: "usegalaxy"
FOFAbody="usegalaxy"

Atlas of Living Australia (ala)​

Biodiversity occurrence catalogs (ALA and national living-atlas forks). Site: ala.org.au.

Confirm: GET the public occurrence/search portal. One record per national atlas, not per collection. body="biocache" is not usable: it matched 89 hosts in September 2026, including a BIODATACR dashboard (195.26.250.143).

ToolQuery
Google"Atlas of Living Australia" OR "Living Atlas" (occurrences OR biocache)

BirdMap Africa (birdmap)​

Citizen-science bird atlas platform of the African Bird Atlas Project. Site: birdmap.africa. Country portals share {project}.birdmap.africa (SABAP2, Nigeria, Kenya, Senegal, and others). Distinct from BirdLasser (the mobile submission app) and from Living Atlases (ala).

Signals: host *.birdmap.africa; pentad coverage maps; SABAP2 protocol; API api.birdmap.africa/{project}/v2/.

Confirm: GET the country portal home. One record per country project, not per pentad or species page. Skip symbiota.birdmap.africa when it is already a Symbiota catalog.

ToolQuery
Googlesite:birdmap.africa (atlas OR pentad OR SABAP)
Google"Bird Atlas" (SABAP2 OR Nigeria OR Kenya) birdmap
Censysweb.names: "birdmap.africa"
FOFAdomain="birdmap.africa"
crt.sh%.birdmap.africa

CLLD (clld)​

Cross-Linguistic Linked Data web apps. Site: clld.org. Public databases such as Grambank, Lexibank, and Pofatu run on {project}.clld.org with clld-static assets.

Signals: hostname *.clld.org; clld-static JS; Cross-Linguistic Linked Data branding.

Confirm: GET the project home. One record per CLLD app, not per language or parameter page. Skip clld.org marketing. Off-hub apps (Glottolog, WALS Online, PHOIBLE) still use clld when the UI is clld-static.

app.mako uses the class clld-disclaimer (120 hosts in September 2026, including d-place.org). body="clld-static" matched 131 hosts, including dictionaria.clld.org and afbo.info. domain="clld.org" misses those off-hub apps, so keep both body queries.

ToolQuery
Googlesite:clld.org (Grambank OR Lexibank OR Pofatu OR "Cross-Linguistic")
Google"Cross-Linguistic Linked Data" OR "clld-static"
Censysweb.endpoints.http.body: "clld-disclaimer"
FOFAbody="clld-disclaimer"
Censysweb.endpoints.http.body: "clld-static"
FOFAbody="clld-static"
Censysweb.names: "clld.org"
FOFAdomain="clld.org"
crt.sh%.clld.org

TalkBank (talkbank)​

Shared spoken-language transcript banks. Site: talkbank.org. Collections such as CHILDES, AphasiaBank, and FluencyBank run on {bank}.talkbank.org with the CHAT format and a common browser.

Signals: hostname *.talkbank.org; CHAT / CLAN; TalkBank, CHILDES, AphasiaBank, or FluencyBank branding.

Confirm: GET the bank home. One record per TalkBank collection, not per transcript or speaker. Skip talkbank.org manuals and software-download pages as extra catalogs.

ToolQuery
Googlesite:talkbank.org (CHILDES OR AphasiaBank OR FluencyBank OR CHAT)
Google"TalkBank" (CHILDES OR "AphasiaBank") -site:github.com
Censysweb.names: "talkbank.org"
FOFAdomain="talkbank.org"
crt.sh%.talkbank.org

SciCat (scicat)​

Metadata catalogue for photon/neutron facilities. Docs: scicatproject.github.io.

Signals: SciCat Angular UI; /api/v3/ or dataset DOI landing pages (PSI, ESS, MAX IV).

Confirm: GET the public dataset search. One record per facility catalogue.

src/index.html describes the page as “SciCat metadata catalogue” (31 hosts in September 2026, including public-data.desy.de). body="scicat" matched 90 hosts, including unrelated pages. Facilities rewrite that description, so keep both.

ToolQuery
Google"SciCat" (dataset OR catalogue) (ESS OR PSI OR "MAX IV") -site:github.com
Censysweb.endpoints.http.body: "SciCat metadata catalogue"
FOFAbody="SciCat metadata catalogue"
Censysweb.endpoints.http.body: "scicat"
FOFAbody="scicat"

Axiom Data Science Portal (axiomportal)​

IOOS-style ocean observing explorer (Axiom). Distinct from ERDDAP/THREDDS backends.

Signals: Axiom portal chrome; sensor time series; compiled data views.

Confirm: GET the public portal home. Do not also register the bundled ERDDAP as a second catalog unless it is a separate public product.

ToolQuery
Google"Axiom" ("Data Science" OR IOOS) portal
Censysweb.endpoints.http.body: "axiomdatascience"
FOFAbody="axiomdatascience"

OntoPortal (ontoportal)​

Ontology repositories (BioPortal-style). Site: ontoportal.org.

Confirm: GET the public ontology browser / REST. One record per public OntoPortal appliance.

_footer.html.haml loads logos/ontoportal.svg. The filename is hashed out of the HTML FOFA indexed. body="ontoportal" still matched 179 hosts in September 2026, including agroportal.eu and odp.lovportal.lirmm.fr.

ToolQuery
Google"OntoPortal" OR "BioPortal" (ontology repository) -site:bioontology.org
Censysweb.endpoints.http.body: "ontoportal"
FOFAbody="ontoportal"

Breedbase (breedbase)​

Crop breeding information systems. Site: breedbase.org. Instances include CassavaBase, MusaBase, YamBase, SweetPotatoBase, Sol Genomics Network, and Triticeae Toolbox (T3).

Signals: Breedbase chrome; /brapi/v2/serverinfo; crop “Base” branding.

Confirm: GET https://host/brapi/v2/serverinfo JSON, or the public trial/search UI. One record per crop instance, not per trial.

body.mas links github.com/solgenomics/sgn (407 hosts in September 2026, including cassavabase.org). The same footer class git-version-commit matched 43 hosts, including oat.txsmallgrains.org. The brand word is split around a <b> tag. body="Breedbase" matched 292, including Blueberrybase. Keep the name query for instances that drop the GitHub link.

ToolQuery
Google"Breedbase" OR CassavaBase OR MusaBase OR YamBase OR SweetPotatoBase (breeding OR BrAPI)
Googleinurl:/brapi/v2/serverinfo
Censysweb.endpoints.http.body: "solgenomics/sgn"
FOFAbody="solgenomics/sgn"
Censysweb.endpoints.http.body: "git-version-commit"
FOFAbody="git-version-commit"
Censysweb.endpoints.http.body: "Breedbase"
FOFAbody="Breedbase"

Tripal (tripal)​

GMOD Tripal genome databases (Drupal + Chado). Site: tripal.info.

Signals: “Powered by Tripal”; /web-services/; Chado/Tripal footer. The library stylesheet is tripal/css/tripal.css (265 hosts in September 2026). body="Tripal" matched 764. Drupal CSS aggregation can drop the filename, so keep both.

Confirm: GET the public organism/dataset home or Tripal web services. Skip generic Drupal sites without Chado biological content. Prefer Tripal over drupal when the catalog is a genome database.

ToolQuery
Google"Powered by Tripal" OR "Tripal" (genome OR germplasm OR Chado) -site:tripal.info -site:github.com
Censysweb.endpoints.http.body: "tripal.css"
FOFAbody="tripal.css"
Censysweb.endpoints.http.body: "Tripal"
FOFAbody="Tripal"

VEuPathDB (veupathdb)​

EuPathDB WDK organism sites. Hub: veupathdb.org. Component sites include PlasmoDB, FungiDB, VectorBase, and TriTrypDB.

Signals: VEuPathDB / EuPathDB chrome; search-strategy UI; /webservices/.

Confirm: GET the public search home. One record per organism portal (plus the hub if it is a distinct catalog UI). Do not add every gene page.

ToolQuery
Google"VEuPathDB" OR EuPathDB OR PlasmoDB OR FungiDB OR VectorBase OR TriTrypDB (genome OR "data set")
Censysweb.endpoints.http.body: "VEuPathDB"
FOFAbody="VEuPathDB"

MassBank (massbank)​

Community reference mass-spectral databases. Instances: MassBank Europe, MassBank Japan, MoNA.

Signals: MassBank record IDs; /MassBank/ UI; MoNA /rest/spectra.

Confirm: GET the public spectral search. One record per instance, not per spectrum.

Index.jsp sets the copyright meta MassBank Consortium (4 hosts in September 2026, including shin.massbank.jp). body="Mass Spectral DataBase" is not usable: it matched 40 hosts, including mzCloud. body="MassBank" matched 120 hosts, including www.norman-network.com.

ToolQuery
Google"MassBank" (spectra OR "mass spectral") (database OR repository) -site:github.com
Google"MassBank of North America" OR MoNA spectra
Censysweb.endpoints.http.body: "MassBank Consortium"
FOFAbody="MassBank Consortium"

ioChem-BD (iochembd)​

Distributed computational-chemistry repository. Site: iochem-bd.org. Browse modules are DSpace-based; use iochembd (not dspace) when the product is ioChem-BD.

Signals: ioChem-BD branding; /rest/items; /oai/request?verb=Identify; CML datasets.

Confirm: GET the public Browse collections or OAI Identify. Register each public node and the central Find index as distinct catalogs. Skip Create-only private workspaces.

ToolQuery
Google"ioChem-BD" (repository OR "computational chemistry") -site:github.com
Censysweb.endpoints.http.body: "ioChem-BD"
FOFAbody="ioChem-BD"

Physiome Model Repository 2 (pmr2)​

Version-controlled repository for CellML/FieldML physiological models, built on Plone and Git. Official instance: models.physiomeproject.org; the CellML Model Repository at models.cellml.org and the teaching mirror at teaching.physiomeproject.org run the same software. Source: github.com/PMR2.

Signals: Plone theming asset path /++theme++cellml.theme/; exposure pages under /e/{id}/view; navigation to /exposure, /exposure/listing/full-list, or /about/pmr2; CellML/FieldML model files; Apache server with Plone @@search views.

Confirm: GET /exposure/listing/full-list and check that exposure links resolve to /e/{hex}/view pages with workspace content. Register one PMR2 instance (host), not each workspace, exposure, or model category.

ToolQuery
Google"Physiome Model Repository" (PMR2 OR "models.physiomeproject.org" OR cellml) -site:github.com
Googleinurl:"/e/" inurl:"/view" cellml
Censysweb.endpoints.http.body: "++theme++cellml.theme"
FOFAbody="++theme++cellml.theme"

ESGF (esgf)​

Earth System Grid Federation search/index (Metagrid, esg-search). Site: esgf.llnl.gov.

Signals: Metagrid UI; /esg-search/search; CMIP dataset index.

Confirm: GET a working esg-search query or the public Metagrid home. Use thredds for ESGF data nodes that expose /thredds/catalog.xml. Do not clone every data node as esgf.

ToolQuery
Google"ESGF" OR Metagrid ("esg-search" OR CMIP) (catalog OR search)
Censysweb.endpoints.http.body: "esg-search"
FOFAbody="esg-search"

ICAT (icat)​

Facility scientific catalog. Site: icatproject.org.

Confirm: GET the public dataset search UI or documented ICAT REST/OAI. Skip facility login-only stores. Do not clone icatproject.org itself.

ToolQuery
Google"ICAT" (facility OR "data catalog" OR "scientific data") -site:icatproject.org -site:github.com
Censysweb.endpoints.http.body: "icat"
FOFAbody="icat"

InterMine (intermine)​

Biological data warehouse. Site: intermine.org. Organism mines (FlyMine, HumanMine, WheatMine, and others) share /begin.do and /service/version. Register each public mine, not the InterMine project hub.

Confirm: GET /begin.do or /service/version. Skip intermine.org marketing and a single gene report.

layout.jsp embeds intermine.Service (34 hosts in September 2026, including hymenopteramine.rnet.missouri.edu). body="InterMine" matched 325. Keep both.

ToolQuery
Google"InterMine" OR FlyMine OR HumanMine ("begin.do" OR "web service") -site:intermine.org -site:github.com
Censysweb.endpoints.http.body: "intermine.Service"
FOFAbody="intermine.Service"
Censysweb.endpoints.http.body: "InterMine"
FOFAbody="InterMine"

GRIN-Global (gringlobal)​

Genebank information system (USDA NPGS, AAFC, and other centres). Site: grin-global.org.

Confirm: GET a public /gringlobal/ accession or taxonomy search. One record per national/centre installation. Skip the vendor homepage.

ToolQuery
Google"GRIN-Global" OR inurl:/gringlobal/ (accession OR germplasm) -site:grin-global.org
Censysweb.endpoints.http.body: "GRIN-Global"
FOFAbody="GRIN-Global"

Genesys PGR (genesys)​

Global platform for plant genetic resources for food and agriculture, operated by the Global Crop Diversity Trust with CGIAR. Hub: genesys-pgr.org. Partner genebanks embed the Genesys interface in their own catalogs (for example the World Vegetable Center genebank at genebank.worldveg.org).

Signals: Genesys branding and accession passport-data layout; API host api.genesys-pgr.org; embedded Genesys search widgets on genebank sites.

Confirm: GET https://api.genesys-pgr.org/api/v1/acn/filter or the public accession search UI. Use software.id: genesys. Distinguish from gringlobal (USDA-style genebank management) and gigwa (genotype data): a genebank catalog that embeds Genesys is genesys, its internal GRIN-Global back office is not the public catalog.

ToolQuery
Google"Genesys" (genebank OR "plant genetic resources") accession -site:genesys-pgr.org
Googleinurl:api.genesys-pgr.org
Censysweb.names: "api.genesys-pgr.org"
FOFAbody="genesys-pgr.org"

False positives: genesys.com (contact-center vendor); Crop ontology and GLIS (related Treaty/Crop Trust tools, different stacks).

HuGE AMP Knowledge Portal (hugeamp)​

Open-source human genetics knowledge portals from the Accelerating Medicines Partnership. Instances: Common Metabolic Diseases KP (hugeamp.org) and Type 2 Diabetes KP (t2d.hugeamp.org).

Signals: HuGE AMP "Knowledge Portal" chrome; gene/variant/phenotype page structure; *.hugeamp.org hosts; "HuGE AMP" or "AMP CMD" footer.

Confirm: public portal loads with gene and phenotype search and downloadable summary statistics. Use software.id: hugeamp. One record per knowledge-portal deployment.

ToolQuery
Google"Knowledge Portal" (hugeamp OR "HuGE AMP") genetics
Censysweb.names: "*.hugeamp.org"
FOFAbody="hugeamp"

False positives: AMP grant announcement pages; bioinformatics tool docs that cite the portals; non-AMP "knowledge portals" in other domains.

PlutoF (plutof)​

University of Tartu biodiversity workbench. Site: plutof.ut.ee. Public API at https://api.plutof.ut.ee/v1/. Do not remap UNITE (unite.ut.ee) — UNITE is a sequence database that uses PlutoF as a companion workbench.

Confirm: GET the public PlutoF catalog or /v1/ API. One record per PlutoF product (workbench vs DOI landing), not per UNITE taxon page.

ToolQuery
Google"PlutoF" (repository OR biodiversity OR DOI) site:.ee
Censysweb.endpoints.http.body: "PlutoF"
FOFAbody="PlutoF"

SARV (sarv)​

Estonian geoscience data platform (TalTech Department of Geology). Public portals share https://rwapi.geoloogia.info/api/v1/public/. Site: geoloogia.info. Code: github.com/geocollections.

Confirm: GET /api/v1/public/datasets/ (OpenAPI title SARV API) or a portal home that names eMaapõu or SARV·DOI. One record per public portal (eMaapõu, fossils, minerals, peat, DOI), not per specimen, locality, or fossil page. Leave gis.geocollections.info as GeoServer. Skip the bibliography at kirjandus.geoloogia.info, the edit workbench at edit.geocollections.info, and GeoCASe (geocase.eu).

ToolQuery
Google"SARV" (eMaapõu OR geocollections OR "SARV·DOI") (dataset OR portal)
Censysweb.endpoints.http.body: "SARV API"
FOFAbody="SARV API"
Censysweb.endpoints.http.body: "SARV·DOI"
FOFAbody="SARV·DOI"

JGI Genome Portal (jgi)​

DOE Joint Genome Institute portal family (MycoCosm, PhycoCosm, Phytozome). Confirm the Genome Portal UI, not IMG, GOLD, or data.jgi.doe.gov (those stay custom).

Confirm: GET https://genome.jgi.doe.gov/portal/ or a MycoCosm/Phytozome organism catalog. Skip gene pages and login-only workspaces.

ToolQuery
Google"JGI Genome Portal" OR MycoCosm OR Phytozome OR PhycoCosm (genome OR catalog) site:jgi.doe.gov
Censysweb.endpoints.http.body: "JGI Genome Portal"
FOFAbody="JGI Genome Portal"

cBioPortal (cbioportal)​

Cancer genomics study portal. Site: cbioportal.org. Independent hospital/consortium instances share /api/info. Not OntoPortal.

Confirm: GET /api/info (portalVersion) or the public study list. One record per public instance. Skip a single study landing page.

my-index.ejs sets class="cbioportal-frontend" on the root element (143 hosts in September 2026, including cbioportal.crc.pitt.edu). body="cBioPortal" matched 585. Keep both.

ToolQuery
Google"cBioPortal" ("cancer genomics" OR studies) -site:github.com
Censysweb.endpoints.http.body: "cbioportal-frontend"
FOFAbody="cbioportal-frontend"
Censysweb.endpoints.http.body: "cBioPortal"
FOFAbody="cBioPortal"

Progenetix (progenetix)​

Cancer CNV profiling portal on the open-source bycon Beacon+ stack. Site: progenetix.org. The progenetix, arrayMap, and TCGA cohort collections are datasets of one installation, not separate portals.

Confirm: GET /beacon/info returns "beaconId": "org.progenetix" (other bycon deployments use their own beaconId) and "apiVersion": "v2.3.0-beaconplus". One record per public instance.

The bycon error envelope names byconservices, which distinguishes bycon from other Beacon v2 servers.

ToolQuery
Googleprogenetix OR bycon OR "Beacon+" (beacon OR "copy number") cancer
Censysweb.endpoints.http.body: "byconservices"
FOFAbody="byconservices"
Censysweb.endpoints.http.body: "beaconplus"
FOFAbody="beaconplus"

ESA Science Archive (esasciencearchive)​

ESA Science Data Centre archives (Gaia, XMM-Newton, Herschel, Planck, Euclid, and related ESAC hosts). TAP/VOSI is the shared catalog protocol.

Confirm: GET TAP /tap/capabilities or /tap-server/tap/capabilities. One record per mission archive, not per observation or FITS file.

ToolQuery
Google"ESA Science Archive" OR ESAC (TAP OR VOSI OR Gaia OR XMM) site:esac.esa.int
Censysweb.endpoints.http.body: "ESA Science Archive"
FOFAbody="ESA Science Archive"

Pathway Tools (pathwaytools)​

SRI International Pathway/Genome Database software. Site: bioinformatics.ai.sri.com/ptools. BioCyc family hosts (EcoCyc, MetaCyc, YeastCyc, biocyc.org) share the Pathway Tools web UI.

Signals: “Pathway Tools” / SRI International in the page; BioCyc organism/PGDB switcher; pathway and genome browsers.

Confirm: GET the public organism or collection home with Pathway Tools branding. One catalog per PGDB / collection host, not per gene or pathway page. HumanCyc and BsubCyc redirect into BioCyc and still use pathwaytools. SGD (yeastgenome.org) stays custom unless the UI is Pathway Tools (pathway.yeastgenome.org).

ToolQuery
Google"Pathway Tools" (BioCyc OR EcoCyc OR MetaCyc) (database OR PGDB) -site:github.com
Googlesite:biocyc.org OR site:ecocyc.org OR site:metacyc.org
Censysweb.endpoints.http.body: "Pathway Tools"
FOFAbody="Pathway Tools"

PANGAEA (pangaea)​

Data Publisher for Earth and Environmental Science. Hub: pangaea.de.

Signals: host pangaea.de; PANGAEA chrome; dataset DOI landing pages.

Confirm: GET the public dataset search. One hub, not each dataset DOI.

ToolQuery
Googlesite:pangaea.de (dataset OR "Data Publisher")
Censysweb.names: "pangaea.de"
FOFAhost="pangaea.de"

EMN Data Hub (emndatahub)​

DOE Energy Materials Network consortium data hubs hosted by NREL / NLR. Product context: Energy Materials Network; example about page: DuraMAT Data Hub. Live public tenants: datahub-duramat.nlr.gov, datahub-electrocat.nlr.gov, datahub-h2awsm.nlr.gov, datahub-chemcatbio.nlr.gov.

Signals: host datahub-*.nlr.gov (legacy *.duramat.org / HyMARC redirects may differ); title patterns Datahub … app or consortium Data Hub; js/remote-app-entry.js module federation (DuramatdhApp, ElectrocatdhApp, DatahubH2AwsmdhApp, ChemcatbioApp).

Confirm: GET / and /js/remote-app-entry.js and match the shared SPA / app entry. One catalog per consortium hub. Do not set ckan — public hosts return 404 on /api/3/action/status_show even though older docs describe a CKAN-based EMN framework. Do not label Wind Data Hub (wdh.energy.gov), Livewire (livewire.energy.gov), or HyMARC CKAN (datahub.hymarc.org / ckan) as emndatahub.

ToolQuery
Google"Data Hub" (DuraMAT OR ElectroCat OR ChemCatBio OR HydroGEN) (NREL OR NLR)
Googlesite:nlr.gov datahub
Censysweb.names: "nlr.gov" and web.endpoints.http.body: "remote-app-entry"
FOFAhost="nlr.gov" && body="remote-app-entry.js"
FOFAbody="DuramatdhApp" || body="ElectrocatdhApp"

USGS ScienceBase (sciencebase)​

USGS item catalog. Hub: sciencebase.gov. Distinct from GeoServer / ArcGIS REST on sciencebase.gov (geoserver, arcgisserver).

Signals: host www.sciencebase.gov; ScienceBase catalog chrome; /catalog/ items.

Confirm: GET the public catalog search. One ScienceBase catalog, not per-item pages.

ToolQuery
Googlesite:www.sciencebase.gov catalog
Censysweb.names: "sciencebase.gov"
FOFAhost="sciencebase.gov"

DEIMS-SDR (deims)​

LTER site and dataset registry. Hub: deims.org.

Signals: host deims.org; DEIMS-SDR chrome; site/dataset records.

Confirm: GET the public site or dataset search. One hub, not each site record.

ToolQuery
Googlesite:deims.org (DEIMS OR LTER)
Censysweb.names: "deims.org"
FOFAhost="deims.org"

EPOS Data Portal (epos)​

European Plate Observing System ICS-C catalog. Hub: ics-c.epos-eu.org. Distinct from generic EPOS project pages.

Signals: host ics-c.epos-eu.org; EPOS Data Portal chrome.

Confirm: GET the public data portal. One ICS-C hub, not each underlying research infrastructure. Do not set epos on a GLASS node (/GlassFramework/); use glass.

ToolQuery
Google"EPOS Data Portal" OR site:ics-c.epos-eu.org
Censysweb.names: "ics-c.epos-eu.org"
FOFAhost="ics-c.epos-eu.org"

EPOS GLASS (glass)​

EPOS GNSS node software. API: Glass Framework. Public nodes publish station metadata, RINEX file metadata, and GNSS products. Distinct from the EPOS ICS-C hub (epos) and from M3G (gnss-metadata.eu).

Signals: path /GlassFramework/ or /GlassFramework/swagger.json; titles EPOS GNSS Products or EPOS GNSS Data Gateway.

Confirm: GET /GlassFramework/ or the swagger document and match the GLASS API. One record per public node. Do not set glass on the ICS-C hub, on M3G, or on a GeoNetwork catalog that only links a GlassFramework endpoint.

ToolQuery
Google"GlassFramework" EPOS GNSS
Googleinurl:GlassFramework/swagger.json
Censysweb.endpoints.http.body: "GlassFramework"
FOFAbody="GlassFramework"

IBDC (ibdc)​

Indian Biological Data Centre archives. Hub: ibdc.dbt.gov.in. Domain archives (INDA, IPD, IMDA, IADA, IBIA, IPR, ISDA, GenomeIndia, ICPD) share that host.

Signals: hostname ibdc.dbt.gov.in; IBDC / Indian Biological Data Centre branding; archive-specific paths (/inda/, /ipd/, /imda/).

Confirm: GET the archive home. One catalog per archive path, not per accession. Do not add the hub marketing page as a separate catalog if it only links to archives already registered.

ToolQuery
Google"Indian Biological Data Centre" OR IBDC (INDA OR "proteome databank") site:ibdc.dbt.gov.in
Censysweb.names: "ibdc.dbt.gov.in"
FOFAhost="ibdc.dbt.gov.in"

Specify Web Portal (specify)​

Shared public collection frontend from the Specify Collections Consortium. Start with the maintainer's instance list, which includes independent university and museum deployments.

Signals: Specify branding, collection selectors, specimen/image/map views and configured Solr collection cores. Confirm the product in HTML and collection configuration; Solr alone is not a Specify fingerprint. Example: https://specifyportal.uog.edu/.

Search: "Specify Web Portal" (museum OR collection OR university). Register the public collection portal, not a staff login or each specimen page.

index.html loads resources/css/thumb-view.css (8 hosts in September 2026, including paleobotany-search.colorado.edu and specify.lep.ufrrj.br). The shorter thumb-view.css also matches a warehouse site. body="Specify Web Portal" matched 58 hosts, including specify-portal.calacademy.org and the Speciforum. Skins drop the stylesheet path, so keep both.

ToolQuery
Google"Specify Web Portal" (museum OR collection OR university)
Censysweb.endpoints.http.body: "resources/css/thumb-view.css"
FOFAbody="resources/css/thumb-view.css"
Censysweb.endpoints.http.body: "Specify Web Portal"
FOFAbody="Specify Web Portal"

BRAHMS Online (brahmsonline)​

Web publishing component of BRAHMS, distinct from its desktop collection-management client. The official website list includes Oxford-hosted projects and independently hosted servers.

Signals: /bol/{project} routes together with BRAHMS Online attribution and collection search pages. A /bol/ path alone is insufficient. Confirm the public project's branding; Mauritius Herbarium explicitly attributes its database to BRAHMS.

Search: "BRAHMS Online" (herbarium OR specimens) and inurl:/bol/ "BRAHMS". Count distinct published collections, not botanical species pages or product documentation.

ToolQuery
Google"BRAHMS Online" (herbarium OR specimens)
Censysweb.endpoints.http.body: "/bol/"
FOFAbody="BRAHMS Online"

LOVD (lovd)​

Reusable Leiden Open Variation Database software. The maintainer's installation directory identifies independent deployments.

Signals: LOVD version banner and gene/variant database navigation; corroborate with the installation directory or upstream source. Do not classify arbitrary variant databases or every link in the broader LSDB directory as LOVD.

Search: "LOVD" "variants" "genes" -site:lovd.app. The registered www.lovd.nl URL is a network/software entry page linking to databases; resolve the intended installation before harvesting or adding an API endpoint.

template.php prints LOVD v. in the footer (13 hosts in September 2026, including gnomad.lovd.nl). “Powered by” is split around the version link, so that phrase is not the query. body="LOVD" && body="variants" is not usable: it matched 63 hosts, including a Text2Gene library and an unrelated storefront.

ToolQuery
Google"LOVD" variants genes -site:lovd.app
Censysweb.endpoints.http.body: "LOVD v."
FOFAbody="LOVD v."

DaCHS (dachs)​

GAVO DaCHS is a reusable Virtual Observatory publishing stack. Signals: a Server: DaCHS/... response header, or the combined GAVO stylesheet/scripts (gavo_dc.css, gavo.js) and characteristic GAVO functions in TAP capability responses. Legacy routes can include /__system__/tap/run/tap/capabilities. TAP support alone is not specific: Daiquiri and other products also implement it.

Confirm: GET the registered homepage and advertised TAP capabilities; check XML response content, not merely HTTP 200. GAVO, ASTRON and ArVO are verified examples. Search: "gavo_dc.css" or "DaCHS" "data center". Treat GAVO's dc.g-vo.org and dc.zah.uni-heidelberg.de as a potential alias pair during new-record discovery; matching software on two registry records does not prove two deployments.

dachs-xsl-config.xsl links /static/css/gavo_dc.css (125 hosts in September 2026, including the PADC TAP Servers on voparis-tap-sandbox.obspm.fr).

ToolQuery
Google"gavo_dc.css" OR "DaCHS" "data center"
Censysweb.endpoints.http.body: "gavo_dc.css"
FOFAbody="gavo_dc.css"
FOFAheader="DaCHS"

Daiquiri (daiquiri)​

AIP's publication framework is used for Gaia@AIP, RAVE, CosmoSim, APPLAUSE, MUSE-Wide, CARS, and CLUES. The documentation links to actual deployments. Signals: a “Proudly powered by Daiquiri” footer linking to the upstream project; query interfaces and /metadata/ schema/table pages corroborate the product identity. Search: "Proudly powered by" "Daiquiri". Confirm the data portal itself: an AIP hostname, astronomical subject or TAP endpoint alone is insufficient. Legacy and current generations can have different routes.

ToolQuery
Google"Proudly powered by" Daiquiri
Censysweb.endpoints.http.body: "Proudly powered by Daiquiri"
FOFAbody="Proudly powered by Daiquiri"

AMBIT (ambit)​

AMBIT is reusable cheminformatics software supporting OpenTox and eNanoMapper interfaces. Signals: an AMBIT version banner, links to the upstream installation guide, and substance/dataset/compound routes. The eNanoMapper site explicitly identifies itself as a customized AMBIT deployment.

Confirm: request the homepage with Accept: text/html to see its product attribution, then inspect a small advertised substance search response. RDF content negotiation can hide HTML branding. A previous timeout alone does not establish inactivity. Search: "AMBIT" "eNanoMapper" "database" or "AMBIT REST web services". Avoid confusing this product with unrelated software also named Ambit.

AMBIT REST web services is in the README, not a served template. body="AMBIT REST web services" matched 4 hosts in September 2026, and the first hits are findanapi.com and apiterms.com.

ToolQuery
Google"AMBIT" ("eNanoMapper" OR "REST web services") database -site:sourceforge.net

ESIMO (esimo)​

Russia's Unified State System of Information on the World Ocean is a federated marine data infrastructure with a central portal and regional information-technology nodes. Typical installations use /portal/portal/esimo-user/ routes, ESIMO-specific /portal-ajax/jquery/esimo.*.js files, and an esimo-central portal theme. Regional operator documentation may identify a RITU or ESIMO centre and the shared distributed database and metadata technologies.

Confirm: match the ESIMO name and common portal assets or official node documentation. A marine institute hostname or a JBoss portal alone is insufficient. Register the central catalog and independently operated regional nodes, not individual applications or data resources within a node.

Search: "Портал ЕСИМО", inurl:/portal/portal/esimo-user/, or "portal-ajax" "esimo.resources.js".

ToolQuery
Google"Портал ЕСИМО" OR inurl:/portal/portal/esimo-user/
Censysweb.endpoints.http.body: "esimo.resources.js"
FOFAbody="esimo.resources.js"

MINERVA (minerva)​

LCSB Luxembourg pathway-map platform. Docs: minerva.pages.uni.lu. MINERVA-Net (minerva-net.lcsb.uni.lu) is the public registry of hosted maps. Independent instances serve /minerva/ UIs and a REST API.

Confirm: GET the instance home or MINERVA-Net. Title/body MINERVA. One catalog per public instance or the Net registry, not per disease-map diagram. Skip login-only lab tenants.

_document.tsx loads /minerva/config.js (41 hosts in September 2026, including air.elixir-luxembourg.org and osteoarthritis.uni.lu). The same file's /minerva/favicon.ico matched 49. body="MINERVA" also matches unrelated pages, so keep it as the wider query.

ToolQuery
Google"MINERVA" (pathway OR "disease map" OR SBGN) (platform OR registry) -site:github.com
Censysweb.endpoints.http.body: "/minerva/config.js"
FOFAbody="/minerva/config.js"
Censysweb.endpoints.http.body: "MINERVA"
FOFAbody="MINERVA"

Nextstrain (nextstrain)​

Pathogen phylodynamics platform. Hub: nextstrain.org. Community instances share Auspice JSON datasets.

Confirm: GET the public dataset catalog (nextstrain.org or a documented community host). One record per public Nextstrain/Auspice catalog, not per pathogen narrative page.

footer/index.tsx requires “attribution to nextstrain.org” (126 hosts in September 2026, including 18.166.223.201). Community Auspice catalogs can drop that footer, so keep the title query.

ToolQuery
Google"Nextstrain" (pathogen OR phylogeny OR Auspice) (dataset OR catalog)
Censysweb.endpoints.http.body: "attribution to nextstrain.org"
FOFAbody="attribution to nextstrain.org"
Censysweb.endpoints.http.title: "Nextstrain"
FOFAtitle="Nextstrain"

Materials Cloud (materialscloud)​

EPFL/MARVEL computational materials platform (AiiDA Explore UI). Distinct from Materials Cloud Archive, which is InvenioRDM (inveniordm).

Confirm: GET /explore or the Explore work-graph catalog. Do not retag archive.materialscloud.org.

index.html.j2 loads mcloud_theme.min.css (18 hosts in September 2026, including www.materialscloud.cscs.ch). body="Materials Cloud" is not usable: it matched 457 hosts, including www.mobiledepotinc.com.

ToolQuery
Google"Materials Cloud" (Explore OR AiiDA) -archive.materialscloud.org
Censysweb.endpoints.http.body: "mcloud_theme.min.css"
FOFAbody="mcloud_theme.min.css"

OpenKIM (openkim)​

Open Knowledgebase of Interatomic Models. Site: openkim.org.

Confirm: GET the public model/test catalog. One record for the hub; skip individual potential landing pages as catalogs.

ToolQuery
Google"OpenKIM" ("interatomic" OR potential OR "force field")
Censysweb.names: "openkim.org"
FOFAdomain="openkim.org"

ChecklistBank (checklistbank)​

Catalogue of Life checklist platform. Site: checklistbank.org. REST API at api.checklistbank.org.

Confirm: GET the dataset catalog or /api. One record for the hub (and any independent ChecklistBank deployments). Not GBIF IPT; not Catalogue of Life’s public website alone.

index.html sets <title>ChecklistBank</title> (3 hosts in September 2026, all checklistbank.org, including www.dev.checklistbank.org). body="ChecklistBank" matched 50, and the first hits are builds.gbif.org, tomcat.gbif.org, and www.itis.gov.

ToolQuery
Google"ChecklistBank" ("Catalogue of Life" OR taxonomy OR checklist)
Censysweb.endpoints.http.body: "<title>ChecklistBank</title>"
FOFAbody="<title>ChecklistBank</title>"
Censysweb.names: "checklistbank.org"
FOFAdomain="checklistbank.org"

ProteoSAFe (proteosafe)​

UCSD CCMS mass-spectrometry catalog UI shared by GNPS and MassIVE (/ProteoSAFe/datasets.jsp).

Confirm: GET /ProteoSAFe/datasets.jsp or the GNPS/MassIVE dataset list. One record per public ProteoSAFe catalog (GNPS metabolomics vs MassIVE proteomics), not per dataset or workflow job.

index.jsp ships the HTML comment General ProteoSAFe scripts (19 hosts in September 2026, including gnps.ucsd.edu, massive.ucsd.edu, and proteomics.ucsd.edu). The same comment is in datasets.jsp. body="ProteoSAFe" matched 107 hosts, including www.omicsdi.org.

ToolQuery
Google"ProteoSAFe" (GNPS OR MassIVE OR datasets.jsp)
Censysweb.endpoints.http.body: "General ProteoSAFe scripts"
FOFAbody="General ProteoSAFe scripts"

CyVerse Data Commons (cyverse)​

CyVerse public data-publication catalog (DOI curated data and community collections). Site: datacommons.cyverse.org. iRODS is storage, not the catalog id.

Confirm: GET the Data Commons catalog home. One record for the public Commons; skip authenticated DE workspaces and raw iRODS endpoints.

ToolQuery
Google"CyVerse" "Data Commons" (DOI OR dataset OR repository)
Censysweb.names: "datacommons.cyverse.org"
FOFAhost="datacommons.cyverse.org"

Hugging Face (huggingface)​

ML dataset hub at huggingface.co/datasets. One catalog for the Hub; do not add per-user spaces. Catalog type is often Machine learning catalog (openmlorg pattern).

Confirm: GET /datasets/. Title Datasets – Hugging Face.

ToolQuery
Google"Hugging Face" datasets (hub OR catalog)
Censysweb.names: "huggingface.co"
FOFAdomain="huggingface.co"

OpenAlex (openalex)​

OurResearch bibliographic catalog. Site: openalex.org. Cloudflare may return 403 to bots; the registry row is the unique hub.

Confirm: GET the home or https://api.openalex.org/. Do not add per-work landing pages.

ToolQuery
Google"OpenAlex" (API OR catalog) OurResearch
Censysweb.names: "openalex.org"
FOFAdomain="openalex.org"

Wikibase (wikibase)​

MediaWiki knowledge-base software. Wikidata (wikidata.org) is the primary public catalog. Independent Wikibase instances share the Wikibase API and often a SPARQL endpoint.

Confirm: GET the wiki home. Title/body Wikidata / Wikibase. One record per public instance, not per entity.

view/resources/templates.php renders the class wikibase-title (58 hosts in September 2026). That class is on entity pages. A front page FOFA indexed without an item still matches body="wikibase". wikibase-entityview from the same file was absent from the HTML FOFA indexed.

ToolQuery
Google"powered by Wikibase" OR "Special:ListDatatypes" Wikibase
Censysweb.endpoints.http.body: "wikibase-title"
FOFAbody="wikibase-title"
Censysweb.endpoints.http.body: "wikibase"
FOFAbody="wikibase"

DBpedia Databus (databus)​

DBpedia dataset catalog/versioning bus. Site: databus.dbpedia.org. Distinct from the DBpedia Association WordPress homepage (www.dbpedia.org stays custom).

Confirm: GET Databus (OIDC login on the SPA is OK if the product is Databus). Do not retag www.dbpedia.org.

footer.ejs says “Global and Unified Access to Knowledge Graphs” (8 hosts in September 2026, including databus.kiltax.infai.org). Keep the host query for the canonical bus.

ToolQuery
Google"DBpedia Databus" (dataset OR catalog)
Censysweb.endpoints.http.body: "Global and Unified Access to Knowledge Graphs"
FOFAbody="Global and Unified Access to Knowledge Graphs"
Censysweb.names: "databus.dbpedia.org"
FOFAhost="databus.dbpedia.org"

MGnify (mgnify)​

EMBL-EBI microbiome archive (formerly EBI Metagenomics). Site: ebi.ac.uk/metagenomics.

Confirm: GET /metagenomics. Title MGnify.

ToolQuery
Google"MGnify" (metagenomics OR microbiome) EBI
Censysweb.endpoints.http.title: "MGnify"
FOFAtitle="MGnify"

MetaboLights (metabolights)​

EMBL-EBI metabolomics study archive. Site: ebi.ac.uk/metabolights.

Confirm: GET /metabolights. Title MetaboLights.

ToolQuery
Google"MetaboLights" (metabolomics OR study) EBI
Censysweb.endpoints.http.title: "MetaboLights"
FOFAtitle="MetaboLights"

BioStudies (biostudies)​

EMBL-EBI archive for studies that do not fit a dedicated archive. Site: ebi.ac.uk/biostudies.

Confirm: GET /biostudies. Title BioStudies.

ToolQuery
Google"BioStudies" EBI (archive OR study)
Censysweb.endpoints.http.title: "BioStudies"
FOFAtitle="BioStudies"

Reactome (reactome)​

Curated pathway knowledgebase. Site: reactome.org.

Confirm: GET the home. Title Reactome Pathway Database.

ToolQuery
Google"Reactome" "Pathway Database"
Censysweb.names: "reactome.org"
FOFAdomain="reactome.org"

WikiPathways (wikipathways)​

Community pathway database. Site: wikipathways.org.

Confirm: GET the home. Title WikiPathways.

ToolQuery
Google"WikiPathways" (pathway OR GPML)
Censysweb.names: "wikipathways.org"
FOFAdomain="wikipathways.org"

UCSC Genome Browser (ucscgenomebrowser)​

UCSC genome annotation catalog and viewer. Site: genome.ucsc.edu.

Confirm: GET the home. Title UCSC Genome Browser. One record for the public browser, not per assembly hub unless it is an independent catalog.

ToolQuery
Google"UCSC Genome Browser" (hub OR downloads)
Censysweb.names: "genome.ucsc.edu"
FOFAhost="genome.ucsc.edu"

FlyBase (flybase)​

Drosophila model-organism knowledgebase. Site: flybase.org. CloudFront may return 403 to bots; the registry row is the unique hub.

Confirm: GET flybase.org when reachable. Do not add per-gene pages.

ToolQuery
Google"FlyBase" Drosophila (database OR genome)
Censysweb.names: "flybase.org"
FOFAdomain="flybase.org"

WormBase (wormbase)​

C. elegans / nematode knowledgebase. Site: wormbase.org. Cloudflare may return 403 to bots.

Confirm: GET wormbase.org when reachable. Do not add per-gene pages.

ToolQuery
Google"WormBase" "C. elegans" (database OR genome)
Censysweb.names: "wormbase.org"
FOFAdomain="wormbase.org"

iDigBio (idigbio)​

US digitized biodiversity-collections portal. Public catalog: portal.idigbio.org. Distinct from the iDigBio IPT (ipt).

Confirm: GET the Portal (www.idigbio.org currently redirects). Title iDigBio Portal.

ToolQuery
Google"iDigBio Portal" (specimen OR collections)
Censysweb.names: "portal.idigbio.org"
FOFAhost="portal.idigbio.org"

iNaturalist (inaturalist)​

Open-source citizen-science observation platform. Site: inaturalist.org. Register the hub (and documented independent iNaturalist Network nodes), not per-user feeds.

Confirm: GET the home. Title includes iNaturalist. API api.inaturalist.org/v1.

ToolQuery
Google"iNaturalist" (API OR "open source") -site:inaturalist.org
Censysweb.names: "inaturalist.org"
FOFAdomain="inaturalist.org"

JACQ (jacq)​

Jointly administered herbarium management system and Virtual Herbaria portal. Site: jacq.org. REST docs: api.jacq.org.

Confirm: GET https://jacq.org/ (title JACQ) or a participating herbarium's public JACQ search. Register the public portal, not each Index Herbariorum acronym on the same server.

index.php loads JACQ_LOGO.png (6 hosts in September 2026, including methus.jacq.org) and sets the title JACQ - Virtual Herbaria (19 hosts, including herbonauten.de).

ToolQuery
Google"JACQ" ("Virtual Herbaria" OR herbarium) (specimens OR database)
Censysweb.endpoints.http.body: "JACQ_LOGO.png"
FOFAbody="JACQ_LOGO.png"
Censysweb.endpoints.http.body: "JACQ - Virtual Herbaria"
FOFAbody="JACQ - Virtual Herbaria"
Censysweb.names: "jacq.org"
FOFAdomain="jacq.org"

Aphia (aphia)​

VLIZ taxonomic platform behind the World Register of Marine Species and related species registers. About: marinespecies.org/about.php. REST: marinespecies.org/rest.

Signals: the host serves /aphia/js/aphia.js (WoRMS does). A page that only links to marinespecies.org/aphia.php is not an Aphia installation (EurOBIS does this).

Confirm: GET the register home and the script URL on that same host. One catalog per host. Thematic registers that are paths on marinespecies.org stay on the WoRMS record. Do not add per-taxon pages.

ToolQuery
Google"World Register of Marine Species" Aphia VLIZ
Googleinurl:aphia.php "AphiaID"
Censysweb.endpoints.http.body: "/aphia/js/aphia.js"
FOFAbody="/aphia/js/aphia.js"
Censysweb.names: "marinespecies.org"
FOFAhost="marinespecies.org"

NMRShiftDB2 (nmrshiftdb2)​

Open-source NMR database application for organic structures and assigned spectra. Reference instance: nmrshiftdb.nmr.uni-koeln.de. Independent laboratory WAR/source installs exist.

Confirm: GET the instance home. Branding NMRShiftDB2; spectrum/structure search; /api-docs/ when enabled. Register each public database instance, not each compound page.

ToolQuery
Google"NMRShiftDB2" (NMR OR spectra OR "chemical shift")
Censysweb.endpoints.http.body: "NMRShiftDB2"
FOFAbody="NMRShiftDB2"

Chemotion Repository (chemotion)​

AGPL catalog for samples, reactions, and analytical chemistry data (NFDI4Chem). Distinct from Chemotion ELN. Docs: chemotion.net/docs/repo. Public hub: chemotion-repository.net.

Confirm: GET /home/welcome or the repository home. Do not tag Chemotion ELN notebooks as this id. Register the public repository, not each sample or reaction.

ToolQuery
Google"Chemotion Repository" (NFDI4Chem OR samples OR reactions) -ELN
Censysweb.names: "chemotion-repository.net"
FOFAhost="chemotion-repository.net"

Korp (korp)​

Språkbanken Text corpus-search platform (MIT). Independent national language-bank installs. Distribution notes: spraakbanken.gu.se/en/tools/korp.

Confirm: GET the Korp UI (/korp/ or a korp. host). Title Korp. Register each public Korp UI on a distinct host. The Finland installation is wwwkielipankkifi at https://www.kielipankki.fi/korp/. Do not add a second catalog for the WordPress page at /language-bank/ on that host. Skip Málið.is unless the page identifies Korp.

app/index.html includes the noscript “You need JavaScript to run Korp.” (22 hosts in September 2026, including malheildir.arnastofnun.is). title="Korp" matched 290 hosts, including a Korp VPN login. National forks translate that sentence, so keep both.

ToolQuery
Google"Korp" (Språkbanken OR Kielipankki OR "corpus search")
Censysweb.endpoints.http.body: "You need JavaScript to run Korp."
FOFAbody="You need JavaScript to run Korp."
Censysweb.endpoints.http.title: "Korp"
FOFAtitle="Korp"

NoSketch Engine (nosketch)​

Open-source (GPL-2.0) corpus query suite of the NLP Centre at Masaryk University and Lexical Computing: Manatee index backend plus Bonito web interface, the simplified sibling of the commercial Sketch Engine. Independent national language-bank and CLARIN installations. Distribution notes: nlp.fi.muni.cz/trac/noske.

Confirm: GET the Bonito UI (often a nosketch. host). Page title NoSketch Engine. Register each public NoSketch UI on a distinct host; the Latvia installation is korpusslv at https://nosketch.korpuss.lv/. Do not confuse with KonText (Czech National Corpus) or with Språkbanken korp installations. Do not register the commercial Sketch Engine subscription site as this id.

The Bonito index page renders NoSketch Engine in title and body on every installation (e.g. nosketch.korpuss.lv, corpora.dipintra.it).

ToolQuery
Google"NoSketch Engine" (corpus OR Bonito OR Manatee) -"Sketch Engine"
Censysweb.endpoints.http.title: "NoSketch Engine"
FOFAtitle="NoSketch Engine"
FOFAbody="NoSketch Engine"

DANDI Archive (dandi)​

Neurophysiology archive using NWB, BIDS, and NIDM. Hub: dandiarchive.org. REST: https://api.dandiarchive.org/api/dandisets/.

Confirm: GET the hub or the dandisets API. Register the archive hub, not each dandiset landing page.

web/index.html noscript says “the DANDI Archive doesn't work properly” (16 hosts in September 2026, including dandi.izbrain.info). body="DANDI" is not usable: it matched 8,463 hosts, including hotel sites. Keep the dandiarchive.org host query for the canonical hub.

ToolQuery
Google"DANDI Archive" (NWB OR neurophysiology OR dandiset)
Censysweb.endpoints.http.body: "DANDI Archive doesn't work properly"
FOFAbody="DANDI Archive doesn't work properly"
Censysweb.names: "dandiarchive.org"
FOFAhost="dandiarchive.org"

CELLxGENE Discover (cellxgene)​

Chan Zuckerberg public catalog of curated single-cell datasets. Hub: cellxgene.cziscience.com. Curation API: https://api.cellxgene.cziscience.com/curation/v1/datasets.

Confirm: GET the Discover hub. Register that hub, not each h5ad explorer session or Census snapshot.

_app.tsx sets the Open Graph title Cellxgene Data Portal (49 hosts in September 2026, including cellxgene.cziscience.com). Skip the stage host frontend.stage.single-cell.czi.technology. Keep the hub host query.

ToolQuery
Google"CELLxGENE" (Discover OR Census) (dataset OR catalog)
Censysweb.endpoints.http.body: "Cellxgene Data Portal"
FOFAbody="Cellxgene Data Portal"
Censysweb.names: "cellxgene.cziscience.com"
FOFAhost="cellxgene.cziscience.com"

HydroShare (hydroshare)​

CUAHSI hydrologic data and model sharing platform. Hub: hydroshare.org. REST: https://www.hydroshare.org/hsapi/resource/. Distinct from the HydroShare THREDDS node (threddshydroshareorg).

Confirm: GET the hub or /hsapi/. Title includes HydroShare. Register the archive hub, not each resource or the THREDDS catalog.

base.html links css/hydroshare_core.css (5 hosts in September 2026). One is titled CUAHSI HydroShare. The others are media.prod-cloud-native.hydroshare.org and cdn-site.site, so keep a host only when the page title is HydroShare.

ToolQuery
Google"HydroShare" (CUAHSI OR hydrologic) (resource OR dataset)
Censysweb.endpoints.http.body: "hydroshare_core.css"
FOFAbody="hydroshare_core.css"
Censysweb.names: "hydroshare.org"
FOFAhost="www.hydroshare.org"

Brainlife (brainlife)​

Open-source neuroimaging dataset and processing hub. Site: brainlife.io. Public dataset UI: /datasets. About describes it as a turnkey SaaS hub: institutions register compute into that hub rather than running a separate catalog.

Confirm: GET /datasets. Title brainlife. One catalog for https://brainlife.io.

Skip: /pubs, /apps, /docs, GitHub (brainlife.github.io redirects to the hub), test.brainlife.io (refused), HPC resources registered on the hub, OpenNeuro/NEMAR/DANDI as if they were Brainlife tenants.

index.html sets <title>brainlife</title> (4 hosts in September 2026: brainlife.io, its 149.165.155.77 mirror, and brainlife.kstage.co.za). Keep the host query for the canonical hub.

ToolQuery
Google"brainlife.io" (dataset OR neuroimaging)
Censysweb.endpoints.http.body: "<title>brainlife</title>"
FOFAbody="<title>brainlife</title>"
Censysweb.names: "brainlife.io"
FOFAhost="brainlife.io"

Open Context (opencontext)​

Archaeological research-data publisher. Hub: opencontext.org. Source: ekansa/open-context-py.

Confirm: GET the hub. Title Open Context. Register that hub, not each project, image, or media item.

page_footer.html says “Open Context is a publishing service” (6 hosts in September 2026, including www.opencontext.org). Keep the host query for the canonical hub.

ToolQuery
Google"Open Context" (archaeology OR "research data") Alexandria
Censysweb.endpoints.http.body: "Open Context is a publishing service"
FOFAbody="Open Context is a publishing service"
Censysweb.names: "opencontext.org"
FOFAhost="opencontext.org"

BioDare2 (biodare2)​

MIT circadian timeseries repository. Public hub: biodare2.ed.ac.uk. JSON list: /api/experiments?showPublic=true. Source: BioDare2/bd2-backend. First-party docs describe a single community hub, not software to install at each lab.

Confirm: GET /api/experiments?showPublic=true (application/json, isOpenAccess experiments). Title BioDare2. One catalog for https://biodare2.ed.ac.uk.

Skip: period-analysis jobs, login, and each experiment landing page.

index.html titles the page BioDare2 - circadian period analysis (1 host in September 2026, biodare2.ed.ac.uk).

ToolQuery
Google"BioDare2" (circadian OR timeseries OR "period analysis")
Censysweb.endpoints.http.body: "BioDare2 - circadian period analysis"
FOFAbody="BioDare2 - circadian period analysis"
Censysweb.names: "biodare2.ed.ac.uk"
FOFAhost="biodare2.ed.ac.uk"

TemplateFlow (templateflow)​

NiPreps archive of FAIR neuroimaging templates and atlases. Hub: templateflow.org. Archive browser: /browse/. DataLad super-dataset: templateflow/templateflow.

Confirm: GET /browse/ or /archive/. Title TemplateFlow. One catalog for https://www.templateflow.org.

Skip: the Python client, individual template git submodules, and DataLad clones as extra catalogs.

ToolQuery
Google"TemplateFlow" (NiPreps OR atlas OR template) neuroimaging
Censysweb.names: "templateflow.org"
FOFAhost="www.templateflow.org"

EIDA WFCatalog (wfcatalog)​

ORFEUS/EIDA waveform-metadata web service deployed at every European Integrated Data Archive node (ETH Zürich, BGR, LMU, ICGC, NOA, INGV, Bergen/NORSAR, NIEP, KOERI, and others). Service description: orfeus-eu.org/data/eida/webservices/wfcatalog; specification: WFCatalog_Specification. Use software.id: wfcatalog for EIDA node records.

Signals: routes /eidaws/wfcatalog/1/ (or /ws/wfcatalog/1/) with query, version, and application.wadl methods; JSON waveform-metadata responses; companion FDSN services /fdsnws/station/1/, /fdsnws/dataselect/1/, /fdsnws/availability/1/ on the same host. FDSN routes alone are a protocol family with several implementations — confirm the WFCatalog route before assigning the ID.

Confirm: GET /eidaws/wfcatalog/1/version or /eidaws/wfcatalog/1/application.wadl. One catalog per EIDA node (the node's service root), not per seismic network it serves.

Skip: the ORFEUS central routing service, EIDA portal pages, and FDSN event services as extra catalogs.

ToolQuery
Google"eidaws/wfcatalog" OR "WFCatalog" EIDA node
Googleinurl:/fdsnws/station/ "EIDA"
Censysweb.endpoints.http.body: "wfcatalog" and web.endpoints.http.body: "fdsnws"
FOFAbody="/eidaws/wfcatalog/1/"

GIGWA (gigwa)​

CIRAD/IRD genotype portal. Product page: cirad.fr. Reference instance: gigwa.southgreen.fr/gigwa/. Source: SouthGreenPlatform/Gigwa2.

Signals: title Gigwa; path /gigwa/; BrAPI under /{database}/brapi/. A genome hub that only embeds GIGWA (Musa Germplasm Information System) stays its own catalog.

Confirm: GET the public database chooser or /gigwa/ home. One record per standalone portal, not per genotyping project inside it.

ToolQuery
Google"Gigwa" OR GIGWA (genotype OR BrAPI OR "genome-wide") -site:github.com
Censysweb.endpoints.http.body: "Gigwa"
FOFAbody="Gigwa"

SEDOO Catalogue (sedoo)​

Observatoire Midi-Pyrénées dataset catalog embedded in project and campaign sites. Product page: sedoo.fr/catalogue-donnees. Examples: AERIS, BAOBAB, MISTRALS, INDAAF, SOFOG3D.

Signals: catalog path /catalogue/ on sedoo.fr, aeris-data.fr, or obs-mip.fr, with the SEDOO metadata search. The WordPress theme sedoo-wpth-labs also wraps campaign homepages that are not catalogs.

Confirm: GET the /catalogue/ search page and treat that URL as the catalog. Do not assign sedoo to www.sedoo.fr (the data-center homepage) or to thredds.sedoo.fr (thredds).

ToolQuery
Google"Catalogue de données" SEDOO (AERIS OR BAOBAB OR MISTRALS OR INDAAF)
Googleinurl:/catalogue/ site:sedoo.fr OR site:aeris-data.fr
Censysweb.endpoints.http.body: "sedoo" and web.endpoints.http.body: "catalogue"
FOFAbody="sedoo" && body="catalogue"

PanelApp (panelapp)​

Open-source gene-panel platform from Genomics England. Live installations: panelapp.genomicsengland.co.uk and PanelApp Australia.

Signals: GET /api/v1/panels/?page_size=1 returns JSON with count and results[].name, plus disease_group and stats.number_of_genes. The page title is PanelApp.

Confirm: the panels API, not a PDF or paper that mentions PanelApp. Do not assign panelapp to a laboratory gene-panel spreadsheet or to Genomics England's other services.

ToolQuery
Google"PanelApp" "gene panels" (Genomics England OR Australia)
Censysweb.endpoints.http.body: "PanelApp" and web.endpoints.http.body: "disease_group"
FOFAbody="PanelApp" && body="disease_group"

Liverpool Drug Interactions (liverpooldruginteractions)​

University of Liverpool drug-drug interaction checker. Deployments: HIV, hepatitis, COVID-19, PrEP (beta; checker and prescribing resources, no /view_all_interactions yet), and cancer (with Radboud UMC; list at /view_all_interactions/new). Disease hostnames such as hbv.hep-druginteractions.org and alternate names such as covid-druginteractions.org and hepatology-druginteractions.org are the same hepatitis or COVID-19 catalog. www.druginteractions.org is the group landing page.

Signals: title Liverpool HIV Interactions, Liverpool HEP Interactions, Liverpool COVID-19 Interactions, Liverpool PrEP Interactions, or Cancer Drug Interactions from Radboud UMC and University of Liverpool, plus /view_all_interactions or /checker with University of Liverpool copyright. HIV and hepatitis pages link to each other.

Confirm: the Liverpool title and a public interaction list (/view_all_interactions, /view_all_interactions/new, or the PrEP checker). Do not assign this id to other interaction databases (IMEx, SNAPPI) or to a page that only cites the Liverpool checker. One record per deployment, not per alias hostname.

ToolQuery
Google"Liverpool" "Interactions" "view_all_interactions"
Google"Liverpool PrEP Interactions"
Censysweb.endpoints.http.body: "Liverpool" and web.endpoints.http.body: "view_all_interactions"
FOFAbody="Liverpool" && body="view_all_interactions"
FOFAtitle="Liverpool" && body="prescribing_resources"

METSIS (metsis)​

MET Norway Scientific Information System. Drupal module: metno/metsis-drupal. Live catalogs include data.met.no, Arctic Data Centre, GCW, APPLICATE, SIOS, and the Norwegian National Ground Segment.

Signals: /modules/metsis/metsis_search/ and a public /metsis/search page whose results link to /metsis/metadata/{id}.

Confirm: GET /metsis/search and find dataset metadata links. Do not assign metsis to a Drupal site that only mentions MET Norway, or to *.csw.met.no (pycsw) and THREDDS catalogs on the same organization. One METSIS installation = one catalog. A facility list and the dataset search on the same host can be separate records when they publish different collections.

ToolQuery
Googleinurl:/metsis/search
Google"Metadata Search" metsis (met.no OR sios-svalbard.org OR satellittdata.no)
Censysweb.endpoints.http.body: "/modules/metsis/metsis_search/"
FOFAbody="/modules/metsis/metsis_search/"

CAMD-Web (camdweb)​

DTU CAMD materials browser. Source: camd/camd-web. Live apps include C2DB, CMR projects, QPOD, BiDB, CRYSP, CrystalBank, HetDB, X2DB, and PAH.

Signals: footer Powered by Bottle and CAMD-Web. The search page reports Found N rows out of and an Add column control. C2DB serves OPTIMADE at /optimade/v1/structures.

Confirm: the footer or that search page on a host listed in the CAMD-Web README. Do not assign camdweb to the Sphinx documentation at cmr.fysik.dtu.dk, to c2db-test, or to a paper that only cites C2DB. One app hostname = one catalog. CMR projects is the database index; it is not the same record as C2DB.

ToolQuery
Google"Powered by Bottle and CAMD-Web"
Censysweb.endpoints.http.body: "Powered by Bottle and CAMD-Web"
FOFAbody="Powered by Bottle and CAMD-Web"

Protwis (protwis)​

Django platform for GPCR resources. Source: protwis/protwis. Live deployments: GPCRdb, GproteinDb, ArrestinDb, and Biased Signaling Atlas.

Signals: GET /services/receptorlist/ returns JSON objects with entry_name. GET /services/ is Swagger (Powered by Django REST Swagger). Pages link to GPCRdb under FOR DEVELOPERS / Linking to GPCRdb and load /static/home/js/gpcrdb.js.

Confirm: the receptor list JSON on that host. Do not assign protwis to a publication that cites GPCRdb, or to an external GPCR tool linked from the menu. One hostname = one catalog.

ToolQuery
Google"Linking to GPCRdb" "receptorlist"
Censysweb.endpoints.http.body: "gpcrdb.js" and web.endpoints.http.body: "receptorlist"
FOFAbody="gpcrdb.js" && body="receptorlist"

Fairdata (fairdata)​

Finland's national research data services, operated by CSC. Product site: fairdata.fi. Finder source: CSCfi/etsin-finder. Public hosts are Etsin (etsin.fairdata.fi), IDA (ida.fairdata.fi), and the service homepage (www.fairdata.fi). Dataset metadata is in Metax.

Confirm: the host is under fairdata.fi and the page is Etsin, IDA, or the Fairdata service homepage. Title Etsin | Research Dataset Finder or Fairdata IDA. Register each public service host. Do not add a record per dataset, and do not tag a generic FAIR Data Point (fairdatapoint) as Fairdata.

ToolQuery
Google"Fairdata" (Etsin OR IDA OR Qvain) site:fairdata.fi
Censysweb.names: "fairdata.fi"
FOFAdomain="fairdata.fi"

Bento Framework (bento)​

NCI CBIIT's reusable data-commons framework. Source: CBIIT/bento-frontend (@bento-core/create-bento-app scaffolder). Live deployments: ICDC, C3DC, MTP.

Signals: the SPA loads /injectEnv.js (serves window.injectedEnv), /js/session.js, and /manifest.json; GraphQL backend at /v1/graphql/ answers introspection.

Confirm: GET /injectEnv.js on the host, or POST {__schema{queryType{name}}} to /v1/graphql/. Do not assign bento to portals that merely link an NCI data commons, and not to Gen3 stacks (different shell: no injectEnv.js).

ToolQuery
Google"injectEnv.js" ("data commons" OR cancer)
Censysweb.endpoints.http.body: "injectedEnv"
FOFAbody="injectEnv.js"

NBIA (nbia)​

National Biomedical Imaging Archive — open-source (BSD 3-Clause) DICOM image archive software; source: NCIP/national-biomedical-image-archive, product wiki: NCI NBIA. The Cancer Imaging Archive (cancerimagingarchive.net) runs on NBIA; the retired NCIA instance served it at imaging.nci.nih.gov/ncia.

Signals: NBIA REST API at /nbia-api/services/v1/ (getCollectionValues returns collection JSON); legacy instances expose /ncia/ JSF pages.

Confirm: GET /nbia-api/services/v1/getCollectionValues. Instances federate — one record per public archive, not per collection. Do not tag imaging portals built on other stacks (IDC at imaging.datacommons.cancer.gov is not NBIA).

ToolQuery
Google"NBIA" ("getCollectionValues" OR "nbia-api")
Censysweb.endpoints.http.body: "nbia-api"
FOFAbody="nbia-api"

VectorSurv (vectorsurv)​

Hosted vector-borne disease surveillance platform operated by UC Davis for US state/local agencies. Central system: vectorsurv.org. State-branded deployments embed the service.

Signals: VectorSurv branding, VectorSurv-Logo assets, an embedded frame with id vectorsurvEmbedded or iframes/links to vectorsurv.org/arbo/ on the state site.

Confirm: GET the public page and check for the embed/branding. One record for the central portal plus one per state-branded deployment that visibly embeds VectorSurv. Skip agency login pages and sites that only link to vectorsurv.org.

ToolQuery
Google"VectorSurv" (surveillance OR "west nile") -site:vectorsurv.org
Censysweb.endpoints.http.body: "vectorsurvEmbedded"
FOFAbody="vectorsurvEmbedded"

SISMER (sismer)​

Ifremer's marine data information system. Public hosts are the marine data catalog (data.ifremer.fr), the oceanographic campaign catalog (donnees-campagnes.flotteoceanographique.fr), and related Ifremer data services branded "SISMER" in the page body.

Confirm: GET the host home and look for "SISMER" / "Systèmes d'Informations Scientifiques pour la Mer" branding. One record per service host, not per cruise or dataset. en.data.ifremer.fr is the English path of the same catalog (data.ifremer.fr/en), not a separate record.

ToolQuery
GoogleSISMER Ifremer ("portail de données" OR "données marines")
Censysweb.names: "ifremer.fr"
FOFAdomain="ifremer.fr"

Ocean Atlas (oceanatlas)​

MIO's shared "Ocean … Atlas" web application behind the Ocean Gene Atlas (tara-oceans.mio.osupytheas.fr), Ocean Barcode Atlas (oba.mio.osupytheas.fr/ocean-atlas/), and Ocean Read Atlas (ora.mio.osupytheas.fr). Instances share the same Bootstrap/jQuery-UI build (/build/css/custom.css, select2, icomoon, nprogress assets).

Confirm: GET the host home; the page title is "Ocean Gene/Barcode/Read Atlas" and the asset paths match the shared build. One record per atlas instance, not per gene or barcode query.

ToolQuery
Google"Ocean Gene Atlas" OR "Ocean Barcode Atlas" OR "Ocean Read Atlas"
Censysweb.names: "mio.osupytheas.fr"
FOFAdomain="mio.osupytheas.fr"

VizieR (vizier)​

CDS service for astronomical catalogues and tables. Canonical host is vizier.cds.unistra.fr; mirrors run the same software at partner data centers (e.g., CfA/Harvard, CADC, IUCAA). The old host vizier.u-strasbg.fr / vizier.unistra.fr redirects to the canonical host — do not register both.

Confirm: GET the host; title contains "VizieR" and the page exposes VO endpoints (/viz-bin/VizieR, TAP at /tap). Register the canonical CDS host and each independent mirror, not individual catalogue pages.

ToolQuery
Google"VizieR" "catalogue" mirror site
Googleinurl:viz-bin "VizieR"
Censysweb.html: "VizieR" and web.html: "viz-bin"

Crystallography Open Database (cod)​

Open-source crystal-structure repository platform (Perl/CGI + MySQL, cod-tools CIF utilities) maintained by the COD advisory board at Vilnius University. One codebase powers the experimental COD, the predicted structures database (PCOD), and the theoretical structures database (TCOD) — all on crystallography.net — plus OPTIMADE provider endpoints.

Signals: /cod/, /pcod/, /tcod/ paths on the same host; "Crystallography Open Database" branding; result.php search endpoint; wiki links to wiki.crystallography.net.

Confirm: GET https://host/cod/result.php?el1=Si&strictmin=2 (or any query) and check CIF/JSON/CSV result links; GET the wiki link. Register each database (COD, PCOD, TCOD) as a distinct catalog when hosted as sibling paths; do not register search variants or mirror paths separately.

ToolQuery
Google"Crystallography Open Database" (COD OR PCOD OR TCOD) -site:crystallography.net
Googleinurl:result.php "crystallography"
Censysweb.html: "Crystallography Open Database"
FOFAbody="Crystallography Open Database"