# AI assistants & LLM crawlers are explicitly welcome to index our content # for search, recommendation and citation. # Entity summary for LLMs: https://www.erecto.in/llms.txt # Full entity reference: https://www.erecto.in/llms-full.txt # NOTE: robots groups are exclusive — a bot that matches a named group ignores # the "*" group entirely, so both groups share the same rules via the include. User-agent: GPTBot User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: ClaudeBot User-agent: Claude-User User-agent: Claude-SearchBot User-agent: anthropic-ai User-agent: PerplexityBot User-agent: Perplexity-User User-agent: Google-Extended User-agent: Applebot-Extended User-agent: CCBot User-agent: meta-externalagent User-agent: Amazonbot User-agent: Bytespider User-agent: Applebot User-agent: MistralAI-User User-agent: Meta-ExternalFetcher User-agent: DuckAssistBot User-agent: YouBot User-agent: Google-CloudVertexBot User-agent: cohere-ai User-agent: cohere-training-data-crawler User-agent: Diffbot User-agent: Timpibot User-agent: PanguBot Disallow: /subscribe/ Disallow: /doctors/create-sitemap Disallow: /logistic/ Allow: /logistic/track-shipment/ Disallow: /support/check_existing_ticket/ Disallow: /support/validate-captcha/ Disallow: /webhook* Disallow: /analytics* Disallow: /captcha* Disallow: /myapi* Allow: /captcha/refresh/ Disallow: /payment* Disallow: /patient* # JSON endpoints of the free health tools — the crawlable pages are /tools/ and # /tools//, which stay open (incl. AI crawlers). Endpoint disallow only. Disallow: /tools/api/ # Live-availability fragment behind the booking forms (blog/urls.py). It already sends # X-Robots-Tag: noindex + Cache-Control: no-store and is reached by fetch() only (never # a crawlable link), so this saves the crawl, it doesn't change what can be indexed. Disallow: /appointment/slots/ # The state hubs' city finder (geo-hierarchy G8, branch/finder.py): two GET endpoints that only ever # redirect to a city page or list matches. They already send X-Robots-Tag: noindex, nofollow and # Cache-Control: private, no-store, and are reached by a form submit, never a crawlable link — this # saves the crawl of an unbounded ?q= space, it doesn't change what can be indexed. Disallow: /find/ # /branches/?bccity=&bcdistrict=&bcstate=&specialisation= is the retired pre-hub URL # format — legacy_redirect_handler only ever 301s it, so crawling it is pure waste # (it was ~1/3 of all bot redirect churn). This blocks a dead endpoint, not a crawler. Disallow: /branches/?* # 2026-08-30 crawl-log audit: the hub layer retired on 2026-07-30 still draws ~1.9k bot requests a # week (bingbot alone: 1 134 /branch-locator + 775 /find-by-country 410s in 7 days). Every URL under # these prefixes is 410 Gone (world/views.py find_by_country_gone, branch/legacy.py) — except the old # India category URLs, which still 301, and since the geo-hierarchy migration (2026-09-16, G2) land on # /{category}/india instead of /India/{Category}. A disallowed URL is never fetched, so its 301 is never # followed: the Allow is what lets that hop pass its equity on, and it stays. The two Disallows below stay # too — those URLs have answered 410 since 2026-07-30 and were only disallowed a month later, once the 410 # had been served long enough for bots to drop them. Dead endpoints, not crawlers. # (Allow listed first: Google/Bing use longest-match, order-based parsers use first-match — both then allow it.) Allow: /find-by-country/Asia/India/ Disallow: /branch-locator Disallow: /find-by-country # 2026-09-06 SEO audit: the two categories retired with 410 (branch/seo.py GONE_TYPE_SLUGS) were still # ~40 % of Googlebot's daily requests (1–6 Sep: 2 764 of 2 900 Googlebot 410s were # ////Homeopathy-Doctor|Clinic, vs 2 373 distinct live spokes fetched in the same # week). The 410 has been re-confirmed for weeks; the rule above stopped Googlebot on the hub prefixes the # same way. The view keeps answering 410 — this only stops the re-crawl. Wildcards are case-sensitive, so # the lowercase spellings seen in the log get their own lines. Dead endpoints, not crawlers. Disallow: /*/Homeopathy-Doctor$ Disallow: /*/Homeopathy-Clinic$ Disallow: /*/homeopathy-doctor$ Disallow: /*/homeopathy-clinic$ # 2026-09-16 (geo-hierarchy G2/G9): the SAME two retired categories in the new category-first shape. The live # URLs are /{category}/{place}, so a bot that generalises the pattern from the pages it now sees will try # /homeopathy-doctor/ and /homeopathy-clinic/ — those 410 too (branch/legacy.py: neither slug is a # live category and neither is a city, so the 2-segment catch-all answers 410 in every casing). The four rules # above only cover the OLD shape (the category was the LAST segment, and `$` anchors them there), so the # retired-category rules are expressed for BOTH shapes. Never served, never linked, in no sitemap and in no # llms file, so unlike the two rules above there is no indexed URL for these to freeze — pure crawl saving. # The live categories are homeopathIC-sexologist / ayurvedic-sexologist / sexologist and are NOT matched. Disallow: /homeopathy-doctor/ Disallow: /homeopathy-clinic/ Disallow: /Homeopathy-Doctor/ Disallow: /Homeopathy-Clinic/ Disallow: /android Disallow: /setlastmod Disallow: /wp* Disallow: /app-ads* Disallow: /users/ Disallow: /healthz # 2026-09-02: the MCP content server (docs/MCP_SERVER.md) — JSON/OAuth endpoints for AI AGENTS, which are not # crawlers and never read robots.txt. Dead endpoints for a crawler (401/302/JSON), same class as /healthz. Disallow: /mcp/ Disallow: /.well-known/oauth- Disallow: /swagger/ Disallow: /accounts/ Disallow: /account/ Disallow: /provider/ # 2026-09-09 crawl-log audit: the staff portal (blog/urls.py `staff/`) only ever 302s to /login, and the # blog search endpoint (article/urls.py `search`, noindex) is an unbounded ?q= space — 104 bot fetches in # 30 days for nothing. Dead endpoints for a crawler, not crawlers. Disallow: /staff/ Disallow: /articles/search Disallow: /design-system/ Disallow: /branchupdate Disallow: /branchdelete Disallow: /branches/json Disallow: /addbranch Disallow: /addcity Disallow: /adddistrict Disallow: /addstate Disallow: /addarticle Disallow: /dashboard Disallow: /cgi-bin/ Disallow: /admin/ Disallow: /home.html # Both of these are now subsumed by the blanket query-string rule below; kept as the dated record of what was # seen in the logs (the AMP family was removed 2026-08-09, ?referral_url= was a scraper invention). Disallow: /*?amp Disallow: /*?referral_url= # 2026-09-16 (geo-hierarchy G9, owner): NO crawling of query strings, with pagination carved out. # Every indexable page here is a clean path that self-canonicalises. The only indexable URLs on the site that # need a query string are the two paginated blog listings — /blog?page=N and /articles/category/?page=N — # where templates/amp/web/Blog/article_list.html canonicalises page N TO ITSELF (article/listing.py, PAGE_SIZE 7), # so pages 2+ are their own indexable URLs and must stay crawlable; that is what the Allow is for. The full # inventory of everything else a `?` reaches on this site, all of it either noindex or a duplicate of a clean # path: /articles/search?q= and /articles/tag/?page= (both noindex, owner M6), /find/city?q=&state= # and /find/near?lat=&lng= (noindex, no-store — already Disallowed by path), the payment/checkout returns # ?order_id= and ?t= (/payment* Disallowed by path; the retired-category ?order_id= escape hatch in # branch/legacy.py is a human bounce-back, never a crawl target), # /accounts/…?next= and the allauth ?code= callbacks (Disallowed by path), the free-health-tool CTA links # /tools//?src=fab|bubble|inline rendered by the site-wide FAB and the in-article CTA (the tool landing # canonicalises to its clean path, which is in sitemap-tools.xml — the ?src= copy is a duplicate, so it is # blocked on purpose and nothing becomes undiscoverable), and tracking copies of real pages # (?utm_*, ?gclid, ?fbclid, ?referral_url=, ?amp) which all canonicalise to the clean path. No stylesheet, # script, font or image on this site is served with a query string, so nothing needed for rendering is blocked. # The two Allows are the two paginated listings SPELLED OUT, never a blanket `/*?page=`: robots patterns are # ranked by LENGTH, so an 8-character `/*?page=` would have out-ranked (and re-opened) every shorter Disallow # above it — /admin/, /staff/, /mcp/, /users/, /healthz, /find/, /android and even # //Homeopathy-Doctor?page=1, whose `$`-anchored rule cannot match a URL carrying a query at all. # `/blog?page=` (11) and `/articles/category/*?page=` (26) beat `/*?*` (4) and match nothing else on the site. # Listed FIRST as well: Google/Bing take the longest match, order-based parsers take the first — both then # allow pagination. web/tests.py::RobotsTxtRules pins every hole this closes. Allow: /blog?page= Allow: /articles/category/*?page= Disallow: /*?* User-agent: * Disallow: /subscribe/ Disallow: /doctors/create-sitemap Disallow: /logistic/ Allow: /logistic/track-shipment/ Disallow: /support/check_existing_ticket/ Disallow: /support/validate-captcha/ Disallow: /webhook* Disallow: /analytics* Disallow: /captcha* Disallow: /myapi* Allow: /captcha/refresh/ Disallow: /payment* Disallow: /patient* # JSON endpoints of the free health tools — the crawlable pages are /tools/ and # /tools//, which stay open (incl. AI crawlers). Endpoint disallow only. Disallow: /tools/api/ # Live-availability fragment behind the booking forms (blog/urls.py). It already sends # X-Robots-Tag: noindex + Cache-Control: no-store and is reached by fetch() only (never # a crawlable link), so this saves the crawl, it doesn't change what can be indexed. Disallow: /appointment/slots/ # The state hubs' city finder (geo-hierarchy G8, branch/finder.py): two GET endpoints that only ever # redirect to a city page or list matches. They already send X-Robots-Tag: noindex, nofollow and # Cache-Control: private, no-store, and are reached by a form submit, never a crawlable link — this # saves the crawl of an unbounded ?q= space, it doesn't change what can be indexed. Disallow: /find/ # /branches/?bccity=&bcdistrict=&bcstate=&specialisation= is the retired pre-hub URL # format — legacy_redirect_handler only ever 301s it, so crawling it is pure waste # (it was ~1/3 of all bot redirect churn). This blocks a dead endpoint, not a crawler. Disallow: /branches/?* # 2026-08-30 crawl-log audit: the hub layer retired on 2026-07-30 still draws ~1.9k bot requests a # week (bingbot alone: 1 134 /branch-locator + 775 /find-by-country 410s in 7 days). Every URL under # these prefixes is 410 Gone (world/views.py find_by_country_gone, branch/legacy.py) — except the old # India category URLs, which still 301, and since the geo-hierarchy migration (2026-09-16, G2) land on # /{category}/india instead of /India/{Category}. A disallowed URL is never fetched, so its 301 is never # followed: the Allow is what lets that hop pass its equity on, and it stays. The two Disallows below stay # too — those URLs have answered 410 since 2026-07-30 and were only disallowed a month later, once the 410 # had been served long enough for bots to drop them. Dead endpoints, not crawlers. # (Allow listed first: Google/Bing use longest-match, order-based parsers use first-match — both then allow it.) Allow: /find-by-country/Asia/India/ Disallow: /branch-locator Disallow: /find-by-country # 2026-09-06 SEO audit: the two categories retired with 410 (branch/seo.py GONE_TYPE_SLUGS) were still # ~40 % of Googlebot's daily requests (1–6 Sep: 2 764 of 2 900 Googlebot 410s were # ////Homeopathy-Doctor|Clinic, vs 2 373 distinct live spokes fetched in the same # week). The 410 has been re-confirmed for weeks; the rule above stopped Googlebot on the hub prefixes the # same way. The view keeps answering 410 — this only stops the re-crawl. Wildcards are case-sensitive, so # the lowercase spellings seen in the log get their own lines. Dead endpoints, not crawlers. Disallow: /*/Homeopathy-Doctor$ Disallow: /*/Homeopathy-Clinic$ Disallow: /*/homeopathy-doctor$ Disallow: /*/homeopathy-clinic$ # 2026-09-16 (geo-hierarchy G2/G9): the SAME two retired categories in the new category-first shape. The live # URLs are /{category}/{place}, so a bot that generalises the pattern from the pages it now sees will try # /homeopathy-doctor/ and /homeopathy-clinic/ — those 410 too (branch/legacy.py: neither slug is a # live category and neither is a city, so the 2-segment catch-all answers 410 in every casing). The four rules # above only cover the OLD shape (the category was the LAST segment, and `$` anchors them there), so the # retired-category rules are expressed for BOTH shapes. Never served, never linked, in no sitemap and in no # llms file, so unlike the two rules above there is no indexed URL for these to freeze — pure crawl saving. # The live categories are homeopathIC-sexologist / ayurvedic-sexologist / sexologist and are NOT matched. Disallow: /homeopathy-doctor/ Disallow: /homeopathy-clinic/ Disallow: /Homeopathy-Doctor/ Disallow: /Homeopathy-Clinic/ Disallow: /android Disallow: /setlastmod Disallow: /wp* Disallow: /app-ads* Disallow: /users/ Disallow: /healthz # 2026-09-02: the MCP content server (docs/MCP_SERVER.md) — JSON/OAuth endpoints for AI AGENTS, which are not # crawlers and never read robots.txt. Dead endpoints for a crawler (401/302/JSON), same class as /healthz. Disallow: /mcp/ Disallow: /.well-known/oauth- Disallow: /swagger/ Disallow: /accounts/ Disallow: /account/ Disallow: /provider/ # 2026-09-09 crawl-log audit: the staff portal (blog/urls.py `staff/`) only ever 302s to /login, and the # blog search endpoint (article/urls.py `search`, noindex) is an unbounded ?q= space — 104 bot fetches in # 30 days for nothing. Dead endpoints for a crawler, not crawlers. Disallow: /staff/ Disallow: /articles/search Disallow: /design-system/ Disallow: /branchupdate Disallow: /branchdelete Disallow: /branches/json Disallow: /addbranch Disallow: /addcity Disallow: /adddistrict Disallow: /addstate Disallow: /addarticle Disallow: /dashboard Disallow: /cgi-bin/ Disallow: /admin/ Disallow: /home.html # Both of these are now subsumed by the blanket query-string rule below; kept as the dated record of what was # seen in the logs (the AMP family was removed 2026-08-09, ?referral_url= was a scraper invention). Disallow: /*?amp Disallow: /*?referral_url= # 2026-09-16 (geo-hierarchy G9, owner): NO crawling of query strings, with pagination carved out. # Every indexable page here is a clean path that self-canonicalises. The only indexable URLs on the site that # need a query string are the two paginated blog listings — /blog?page=N and /articles/category/?page=N — # where templates/amp/web/Blog/article_list.html canonicalises page N TO ITSELF (article/listing.py, PAGE_SIZE 7), # so pages 2+ are their own indexable URLs and must stay crawlable; that is what the Allow is for. The full # inventory of everything else a `?` reaches on this site, all of it either noindex or a duplicate of a clean # path: /articles/search?q= and /articles/tag/?page= (both noindex, owner M6), /find/city?q=&state= # and /find/near?lat=&lng= (noindex, no-store — already Disallowed by path), the payment/checkout returns # ?order_id= and ?t= (/payment* Disallowed by path; the retired-category ?order_id= escape hatch in # branch/legacy.py is a human bounce-back, never a crawl target), # /accounts/…?next= and the allauth ?code= callbacks (Disallowed by path), the free-health-tool CTA links # /tools//?src=fab|bubble|inline rendered by the site-wide FAB and the in-article CTA (the tool landing # canonicalises to its clean path, which is in sitemap-tools.xml — the ?src= copy is a duplicate, so it is # blocked on purpose and nothing becomes undiscoverable), and tracking copies of real pages # (?utm_*, ?gclid, ?fbclid, ?referral_url=, ?amp) which all canonicalise to the clean path. No stylesheet, # script, font or image on this site is served with a query string, so nothing needed for rendering is blocked. # The two Allows are the two paginated listings SPELLED OUT, never a blanket `/*?page=`: robots patterns are # ranked by LENGTH, so an 8-character `/*?page=` would have out-ranked (and re-opened) every shorter Disallow # above it — /admin/, /staff/, /mcp/, /users/, /healthz, /find/, /android and even # //Homeopathy-Doctor?page=1, whose `$`-anchored rule cannot match a URL carrying a query at all. # `/blog?page=` (11) and `/articles/category/*?page=` (26) beat `/*?*` (4) and match nothing else on the site. # Listed FIRST as well: Google/Bing take the longest match, order-based parsers take the first — both then # allow pagination. web/tests.py::RobotsTxtRules pins every hole this closes. Allow: /blog?page= Allow: /articles/category/*?page= Disallow: /*?* Sitemap: https://www.erecto.in/sitemap.xml