User-agent: * # Content signals (contentsignals.org): search engines, AI answers (ai-input) and AI training may all use this site. Content-Signal: search=yes, ai-input=yes, ai-train=yes Allow: / Disallow: /api/ # The database, server, review, streamer and discussion panels on the pages are built by their scripts from these # public, read-only data endpoints. Search engines render a page with its scripts, but they do not load anything this # file disallows: with all of /api/ closed, Bing's and Google's renderers got "Failed to fetch" and "Could not load # server details" in place of 99.7% of the pages in the sitemap. The longer Allow lines win over Disallow: /api/, # accounts, reviews-writing, favourites and the rest stay closed, and every /api/ answer carries X-Robots-Tag: noindex # (.htaccess), so the data is loaded but never indexed as a page of its own. Allow: /api/items Allow: /api/monsters Allow: /api/skills Allow: /api/maps Allow: /api/npc-locations/ Allow: /api/latest Allow: /api/homepage.php Allow: /api/forum_topics.php Allow: /api/streamer.php # /includes/ holds the header, footer and page CSS/JS that every page loads. # Disallowing the whole directory stopped Googlebot fetching them, and Google # renders a page before judging it: without the stylesheet and scripts it sees # an unstyled, half-built page. Allow lines are matched by longest rule, so the # two below win for assets while the directory itself stays closed. Allow: /includes/*.css Allow: /includes/*.js Disallow: /includes/ Disallow: /scripts/ Disallow: /logs/ Disallow: /login.php Disallow: /modules/ Disallow: /account/ Disallow: /*?success_redirect_url # Forum: per-member pages and the redirect-only account section are nothing to index. # Member profiles and searches stay crawlable and carry noindex instead (forum-theme/sso/SeoHead.php). Disallow: /forum/settings Disallow: /forum/notifications Disallow: /forum/following Disallow: /forum/t/account-help # "Start a discussion" links on every database page carry the entry (?rmr_topic=item:501&title=...), which made about # 41,700 variants of the section pages - each a 200 with its canonical on the plain section. Nothing there to index. Disallow: /forum/*?rmr_topic= Disallow: /forum/*&rmr_topic= # ─── Every search engine, AI answer engine, SEO index and archive is welcome ─── # # None of them is named below, so all follow the rules above: search engines (Google, Bing, Yandex, Baidu, Naver, # Seznam, Qwant, Sogou, Petal, Brave, DuckDuckGo, Mojeek, Coc Coc...), AI answer engines and their fetchers (OpenAI's # OAI-SearchBot, ChatGPT-User and GPTBot; Anthropic's ClaudeBot, Claude-User and Claude-SearchBot; PerplexityBot, # Google-Extended, Applebot, Amazonbot, Meta, Mistral, DeepSeek, Qwen, ByteDance...), SEO indexes (Ahrefs, Semrush, # Majestic, Moz) and the Internet Archive. Until 17 Sep 2026 this file shut out Bytespider, PetalBot, MJ12bot, DotBot # and ia_archiver entirely; appearing in AI answers, regional search engines and the archive is worth more than their # crawl load. Keep it that way - .htaccess and api/server.js carry the same rule. # # Only site copiers and scraping services are refused. They ignore robots.txt anyway; .htaccess and the API stop them. User-agent: scraperforce Disallow: / User-agent: ScraperForce Disallow: / User-agent: HTTrack Disallow: / User-agent: WebCopier Disallow: / User-agent: Scrapy Disallow: / Sitemap: https://ratemyragnarok.com/sitemap.xml