Skip to content

robots.txt: declare the sitemap (and restore access to it) #326

Description

@jroundshp

robots.txt in this repo currently declares no sitemap, and https://www.gutenberg.org/sitemap.xml returns 404. For a catalog of 77,000+ books, a declared sitemap helps search engines find new releases quickly and confirms the full catalog is indexed.

autocat3 already implements a sitemap index (Sitemap.py, routed at /ebooks/sitemaps/ and /ebooks/sitemaps/{page}), so most of the work is done. Two things are needed:

  1. Add one line to robots.txt in this repo:

    Sitemap: https://www.gutenberg.org/ebooks/sitemaps/
    
  2. Server side (per Eric this may be an Apache config that got lost): https://www.gutenberg.org/ebooks/sitemaps/ currently 301-redirects to http://www.gutenberg.org/ebooks, so crawlers cannot reach the sitemap at all right now. That redirect needs to be removed or corrected before the robots.txt line will do any good.

Related: the sitemap <loc> URLs are generated with an http:// scheme (same root cause as the canonical-tag issue filed in autocat3).

Happy to submit a PR for the robots.txt part.

(Filed at Eric Hellman's request, following a site review I sent him in July 2026.)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions