💾 Archived View for kennedy.gemi.dev › docs › crawling.gmi captured on 2023-04-26 at 12:55:41. Gemini links have been rewritten to link to archived content

View Raw

More Information

⬅️ Previous capture (2022-03-01)

➡️ Next capture (2023-06-14)

-=-=-=-=-=-=-

🔭 Notes on Crawling and Indexing

Home

Kennedy creates its search index by crawling content only within Geminispace. It will not crawl or index other content with other protocols like Gopher or HTTP.

Crawler details

Kennedy crawls Geminispace using the following IP addresses:

Crawler speed

Kennedy throttles itself and waits 1.5 seconds between making requests to the same IP address. This increases the amount of time it takes to crawl multiple capsules hosted from the same IP address, such as Flounder.online.

Robots.txt Support

Kennedy will respect sites that are using the simplified robots.txt protocol defined for Gemini.

Robots.txt subset for Gemini

Specifically, Kennedy will follow the Deny rules defined for the follow user-agents:

Note: There are a number of robots.txt files in Geminispace which use rules outside of the simplified standard above. These include:

Kennedy does not currently respect these rules.

Crawler Limits

Kennedy has the following limits: