# Biblion — https://schema.org/WebSite # # NOTE ON WHAT THESE COMMENTS MAY SAY: this file is served to the public exactly # as written, and it is the first thing almost every crawler and every scanner # asks for. So it names no component, no source file, no internal name and no # figure out of the logs — the standing rule that the site does not tell anyone # what it is built on, and a comment ships just as surely as a rule does. # Rewritten on 2026-08-23 for precisely that reason: the comments that stood # here named things the site says nowhere else, and narrated one measured # incident in enough detail to be a working recipe for repeating it against us. # Before adding a line, ask whether you would be content to see it quoted back. # # What is blocked here is only what renders no page a crawler could read: the # data and redirect endpoints. Everything else that should stay out of an index # says so with a robots META TAG instead, which is deliberate and matters — # Disallow stops a crawler FETCHING a page, and a page that is never fetched is # a page whose noindex is never read. Blocking the private pages here would # leave them eligible to appear as bare URLs on the strength of inbound links # alone. The list of those pages is held once, with the pages, and not here. # # Nothing sensitive is named in this file on purpose: robots.txt is public, and # listing a path here advertises it to exactly the people who should not have it. # NO BLANK LINE BETWEEN A User-agent AND ITS RULES. One stood here until # 2026-08-14 and made this group EMPTY under any parser that follows the letter # of the format: a blank line ends a group, so the eighteen Disallows below # belonged to nothing. The largest crawlers evidently ignore blank lines — they # reached zero disallowed endpoints throughout — but that is their leniency and # not this file being right. A comment line is safe here; an empty one is not. User-agent: * # Endpoints — they answer data or a redirect, so there is no page and no meta # tag on which to carry a directive. Since 2026-08-01 each of them also answers # at its extensionless address, so BOTH forms are listed and both must stay: a # crawler can reach either, and a file that blocks one form only blocks nothing. # The suffixed form says more about the site than a public file otherwise would; # that is a trade made knowingly, because an unblocked endpoint is the worse of # the two outcomes by a distance. Do not "tidy" those lines away. Disallow: /favorite.php Disallow: /favorite Disallow: /save_position.php Disallow: /save_position Disallow: /mark_read.php Disallow: /mark_read Disallow: /mark_studied.php Disallow: /mark_studied Disallow: /article_suggest.php Disallow: /article_suggest Disallow: /article_time.php Disallow: /article_time Disallow: /mood.php Disallow: /mood Disallow: /export.php Disallow: /export Disallow: /logout.php Disallow: /logout # Text and data mining is RESERVED under Article 4(3) of Directive (EU) # 2019/790, which permits commercial mining of lawfully accessible content # unless the rightsholder reserves it in machine-readable form. Saying nothing # here would therefore be a grant, not a neutral position. This is one of FOUR # places the reservation is made — the others are a meta tag carried by every # page, a clause in the terms of use, and llms.txt beside this file. THEY MUST # KEEP SAYING THE SAME THING: a reservation that contradicts itself in one # place invites the argument that it was never properly made at all. # # SINCE 2026-08-10 THE RESERVATION CARRIES NAMED EXCEPTIONS (Daniel's call, # shape B of a considered pair): the general rule stays reserved, and the # operators expressly allowed below are permitted to include this site in # model training. The Allow lines ARE the machine-readable permission the # terms refer to. Why these and not everyone: these are first-party model # operators answerable for what their models do; a blanket grant would hand # the commentary to every scraper with a GPU. # # CCBot IS DELIBERATELY NOT AMONG THEM. Common Crawl is a public corpus that # anybody may mine, so permitting it would be a grant to all comers — the # blanket grant again, wearing a single crawler's name. # # What is reserved is our own work: the commentary, the maps, the devotional # retellings, the design and the site's own text. The biblical texts are public # domain and are not reserved by any of this. # ⚠ EVERY GROUP BELOW REPEATS THE ENDPOINT BLOCKS ABOVE, AND MUST. # # A ROBOTS GROUP DOES NOT INHERIT. A crawler obeys only the most specific group # whose name matches it, and nothing from `User-agent: *` reaches it — so when # the named exceptions were added on 2026-08-10, each `Allow: /` silently # canceled the whole endpoint list for that operator. Nothing failed and # nothing logged; the file simply meant something other than it looked like, # and went on meaning it until 2026-08-14, when the cost showed up in the logs # and was not small. # # The proof it was the groups and not the rules: a crawler with no group of its # own falls back to `User-agent: *`, and those reached a disallowed endpoint # not once over the same days. # # THE ALLOW LINES STAY. They are the machine-readable training permission the # terms refer to, and that policy is unchanged — a longer, more specific # Disallow beats `Allow: /` for that path alone, which is exactly the shape # wanted: train on the writing, do not walk the endpoints. # # So the list appears six times. This file format has no include, and the only # alternative to repeating it is a file that lies. IF AN ENDPOINT IS ADDED TO # THE `*` GROUP IT MUST BE ADDED TO ALL FIVE GROUPS BELOW, or that endpoint is # open to precisely the crawlers that reach for it hardest. User-agent: GPTBot Disallow: /favorite.php Disallow: /favorite Disallow: /save_position.php Disallow: /save_position Disallow: /mark_read.php Disallow: /mark_read Disallow: /mark_studied.php Disallow: /mark_studied Disallow: /article_suggest.php Disallow: /article_suggest Disallow: /article_time.php Disallow: /article_time Disallow: /mood.php Disallow: /mood Disallow: /export.php Disallow: /export Disallow: /logout.php Disallow: /logout Allow: / User-agent: ClaudeBot Disallow: /favorite.php Disallow: /favorite Disallow: /save_position.php Disallow: /save_position Disallow: /mark_read.php Disallow: /mark_read Disallow: /mark_studied.php Disallow: /mark_studied Disallow: /article_suggest.php Disallow: /article_suggest Disallow: /article_time.php Disallow: /article_time Disallow: /mood.php Disallow: /mood Disallow: /export.php Disallow: /export Disallow: /logout.php Disallow: /logout Allow: / User-agent: Google-Extended Disallow: /favorite.php Disallow: /favorite Disallow: /save_position.php Disallow: /save_position Disallow: /mark_read.php Disallow: /mark_read Disallow: /mark_studied.php Disallow: /mark_studied Disallow: /article_suggest.php Disallow: /article_suggest Disallow: /article_time.php Disallow: /article_time Disallow: /mood.php Disallow: /mood Disallow: /export.php Disallow: /export Disallow: /logout.php Disallow: /logout Allow: / User-agent: Applebot-Extended Disallow: /favorite.php Disallow: /favorite Disallow: /save_position.php Disallow: /save_position Disallow: /mark_read.php Disallow: /mark_read Disallow: /mark_studied.php Disallow: /mark_studied Disallow: /article_suggest.php Disallow: /article_suggest Disallow: /article_time.php Disallow: /article_time Disallow: /mood.php Disallow: /mood Disallow: /export.php Disallow: /export Disallow: /logout.php Disallow: /logout Allow: / User-agent: meta-externalagent Disallow: /favorite.php Disallow: /favorite Disallow: /save_position.php Disallow: /save_position Disallow: /mark_read.php Disallow: /mark_read Disallow: /mark_studied.php Disallow: /mark_studied Disallow: /article_suggest.php Disallow: /article_suggest Disallow: /article_time.php Disallow: /article_time Disallow: /mood.php Disallow: /mood Disallow: /export.php Disallow: /export Disallow: /logout.php Disallow: /logout Allow: / User-agent: CCBot Disallow: / # Fully qualified per the sitemap protocol (Bing requires it; Google merely # tolerates a relative path) and using the same canonical host the site uses to # name itself everywhere else, or the two would disagree about what the site is # called. Extensionless since 2026-08-01, matching every entry inside the # sitemap itself. Sitemap: https://biblionapp.com/sitemap