Softechinfra
Technology

The 'llms.txt' File: Should Your Site Have One in 2026?

Only ~10% of websites use llms.txt and no major AI crawler officially requests it. Here is when it is still worth shipping for an Indian services business — with a copy-paste template.

Softechinfra TeamSoftechinfra Team
April 11, 202611 min read
The 'llms.txt' File: Should Your Site Have One in 2026?

Anthropic, Stripe, Cursor, Cloudflare, Vercel, Mintlify, Supabase, and LangGraph all ship an llms.txt file at the root of their docs domains. Around 10% of websites overall now have one. But the inconvenient truth is that no major AI crawler — GPTBot, ClaudeBot, PerplexityBot, Google-Extended — officially requests llms.txt at any meaningful volume. Server-log audits show GPTBot occasionally fetches it; the rest essentially do not. So: should you ship one? Short answer: yes, with caveats and a clear understanding of what it does and does not do. This post is the honest take, the copy-paste template we use for Softechinfra-tier services businesses, and the trade-offs.

~10%
Of websites ship llms.txt in 2026
0
Major AI labs that have publicly committed to read it (as of Q1 2026)
5 KB
Typical file size — costs nothing to host
2 hr
Time to write a good one for a services site

TL;DR — should you ship it

Yes if you are a docs-heavy domain, a SaaS, or a services firm with 30+ valuable URLs. No if your site is 8 pages and a contact form. The cost is two hours and 5 KB. The upside is being on the right side of a future standard if Anthropic or OpenAI ever decide to honor it; the downside is essentially zero. Treat llms.txt as an entity-signal investment and a documentation discipline exercise — not as a magic citation lever. The current 2026 reality is what Google's John Mueller said on Reddit: AI engines do not need a separate file to read your site, they can read it directly.

What is llms.txt, in 30 seconds

llms.txt was proposed by Jeremy Howard (Answer.AI, fast.ai) in September 2024 as a Markdown-formatted file at the root of a website — /llms.txt — that tells large language models which content on the site is most important to read and in what order. The spec is intentionally simple: a Markdown document with an H1 title, an optional short description, and a list of grouped Markdown links to your most useful pages. A common companion file, llms-full.txt, contains the actual Markdown bodies of those pages so an LLM can ingest the entire useful corpus in one fetch without spidering the HTML.

Why this matters now (April 2026 trigger)

Three things happened in Q1 2026 that put llms.txt back on every SEO meeting agenda. First, the Discoverability Co schema study showed pages with comprehensive structured data lift AI citations by 44% — and llms.txt is structured data's lazy cousin. Second, the Mintlify "real llms.txt examples" roundup went viral on r/programming. Third, Google's John Mueller publicly likened llms.txt to the old "keywords" meta tag — which the SEO community read as "Google does not care, but it costs nothing to ship." The split between "ship it for safety" and "do not waste time on it" defines the 2026 debate.

Which AI labs and tools actually read it (the honest list)

AN
Anthropic (Claude)
Ships their own at docs.anthropic.com/llms.txt and docs.claude.com/llms-full.txt. ClaudeBot does not request it at scale; Anthropic's own AI tooling (Claude Code, Claude Skills) does.
CR
Cursor, Cline, Continue
AI coding tools fetch llms.txt for documentation context. If your product is dev-tool-adjacent, this is the highest-impact reader you have today.
OP
OpenAI / GPTBot
Occasionally fetches it. OpenAI's documented recommendation is to use robots.txt for crawler control, not llms.txt. Treat as opportunistic.
PP
Perplexity, Google, Meta
Effectively do not fetch it. PerplexityBot and Google-Extended scrape your site directly. Meta does not engage with the spec.

The honest pattern: if your llms.txt is going to be read at all today, it will be by AI coding assistants and Claude-family tooling on docs-heavy domains. For an Indian B2B services firm, the immediate utility is small. The medium-term bet is whether a standard emerges.

When to ship llms.txt (the decision matrix)

Site typeShip llms.txt?Why
Docs site / API referenceYes, todayAI coding tools fetch it. Direct user value.
Services firm with 50+ blog postsYes, this quarterEntity-graph signal, free, sets discipline
8-page brochure siteNoNot enough content to need a curated index
Edtech / consumer SaaSMaybeUseful if your product is dev-adjacent or has public docs
E-commerceProbably notllms.txt is text-content focused; not designed for catalog
Affiliate / SEO content farmNoThe signal you broadcast is "we want LLMs to scrape us harder," which is the opposite of what you actually want

A copy-paste llms.txt for a Softechinfra-tier services firm

This is what we ship for our own site and for the 14-person Pune logistics-tech client we wrote up in another post. Adapt for your domain.

markdown
# Softechinfra
  
  > Softechinfra is an Indian IT services firm building custom software, AI automation, CRM, web, and mobile applications for SMBs and growth-stage companies. We also ship two in-house product TalkDrill (English fluency for Indian adults).
  
  ## Services
  
  - [AI Automation](https://www.softechinfra.com/services/ai-automation): n8n workflows, custom AI agents, RAG systems, voice and chat bots
  - [Custom Software Development](https://www.softechinfra.com/services/custom-software-development): Next.js, FastAPI, Python, MongoDB, PostgreSQL
  - [CRM Development](https://www.softechinfra.com/services/crm-development): SuiteCRM, Zoho, custom builds
  - [Mobile App Development](https://www.softechinfra.com/services/mobile-app-development): React Native, Flutter, native iOS/Android
  - [SEO and GEO](https://www.softechinfra.com/services/seo): Classical SEO + Generative Engine Optimization audits
  - [Cloud and DevOps](https://www.softechinfra.com/services/cloud-devops): AWS, Hetzner, GCP, CI/CD
  
  ## Founder
  
  - [Vivek Kumar](https://viveksinra.com): Founder and CEO. Personal blog, founder's notes on Indian SMB tech.
  
  ## In-house products
  
  - a client platform: AI-powered creative writing & exam prep for students 11+ in India
  - [TalkDrill](https://talkdrill.com): English speaking & fluency app for Indian adults, 5,000+ active users
  
  ## Blog (priority pages)
  
  - [GEO in 2026: Zero-Click Search Playbook](https://www.softechinfra.com/blog/generative-engine-optimization-2026-zero-click-search)
  - [How to Get Cited by Perplexity](https://www.softechinfra.com/blog/how-to-get-cited-by-perplexity-7-step-audit)
  - [14-Point Schema Markup Checklist](https://www.softechinfra.com/blog/14-point-schema-checklist-ai-overviews)
  - [Brand Mentions Without Backlinks](https://www.softechinfra.com/blog/brand-mentions-without-backlinks-ai-visibility)
  
  ## Case studies
  
  - [Softechinfra Perplexity Citation in 11 Weeks](https://www.softechinfra.com/blog/softechinfra-perplexity-citation-case-study)
  
  ## Contact
  
  - Email: contact@softechinfra.com
  - Phone: +91 [number]
  - Location: India

Three things to notice. First, the description block (the > line) is the most-cited line — keep it tight, name your entity, name what you do, name where you operate. Second, the link list uses Markdown anchor format with a short description per link — this is the spec, follow it. Third, group your links by purpose (Services / Founder / Products / Blog / Contact) so any LLM ingesting it has clear sections to reason over.

llms-full.txt — when to ship the longer variant

The companion file llms-full.txt contains the full Markdown body of the pages listed in llms.txt. Anthropic, Vercel, and LangGraph all ship both. The use case is: an AI coding tool wants to ingest your entire useful documentation corpus in one fetch instead of crawling 30 HTML pages.

For a services firm, llms-full.txt makes sense only if your blog is genuinely useful as a knowledge corpus to an LLM. If your blog is mostly news-recap and "10 reasons to use Cloud," skip it. If your blog has technical depth — actual code, real numbers, named clients — ship llms-full.txt with your top 20–30 posts concatenated as Markdown.

Privacy caveat: llms-full.txt is a literal full-text dump. Do not put anything in it that you do not want crawled, embedded, or potentially reproduced inside an LLM's answer. Strip client names, NDA-covered details, and any internal-only documentation before publishing.

The DIY ship plan (2 hours, including validation)

Quick pre-flight checklist:

  • You can write to /public/llms.txt or your hosting equivalent
  • You have a list of your 20 to 30 most valuable URLs ready
  • You have a tight 1-line description of who you are, what you do, who you serve
  • You have decided whether to ship llms-full.txt as well (yes if your blog is technical and useful)
  • You have stripped any client names or NDA-covered details from the corpus you plan to publish
1
Step 1 — Draft your description block (30 min)
One H1 (your brand name) and one blockquote line that names: who you are, what you do, who you serve, where you operate. Cut every adjective. The line gets quoted into AI answers when it works.
2
Step 2 — Curate the link list (45 min)
Pick your 20–30 most valuable URLs. Group them under H2 headers (Services / Products / Blog / Contact). Each link gets a 10–20 word description after the colon. Skip anything you would not point a serious buyer to.
3
Step 3 — Save as /public/llms.txt or /llms.txt (15 min)
For Next.js: place in /public/llms.txt and confirm GET /llms.txt returns 200 with Content-Type text/plain or text/markdown. For Cloudflare Pages: same. For WordPress: drop in wp-content with a redirect rule.
4
Step 4 — Validate and submit (30 min)
Validate at llmstxt.org. Add a self-reference in your sitemap.xml. Optionally submit to llms.txt hub for discoverability. Done.

The r/SEO debate — the honest read

If you spend an hour on r/SEO and r/bigseo threads from Q1 2026, the consensus is mixed. The skeptical camp (loudest) says: AI labs do not request it, John Mueller compared it to the keywords meta tag, it is a waste of time. The pragmatic camp says: it costs 2 hours, the file is 5 KB, and there is no downside. The bullish camp says: the standard could be formalized in 2027 and you want to be early.

Our take after shipping it for ~40 client sites since mid-2025: the direct citation impact is unmeasurable. The indirect impact — forcing a team to write a clear, structured description of what they do and what their best content is — is significant. It is a documentation discipline exercise dressed up as an SEO tactic. If you have docs-heavy infrastructure, ship it now. If you do not, ship it when your blog hits 30+ valuable posts.

Common mistakes

Treating llms.txt as a sitemap.xml replacement. They are different. sitemap.xml lists every URL for indexing crawlers. llms.txt curates the most useful URLs for LLM context. Ship both.

Stuffing every URL into llms.txt. The spec is about curation. A site that lists 800 URLs in llms.txt is signaling "I do not know what is important on my own site," which is the opposite of useful.

Forgetting to update it. llms.txt is a living document. When you publish a major new post, add it. When a service page deprecates, remove it. We rebuild ours quarterly.

Mismatched URLs. Every URL in llms.txt must return 200. Broken links signal a low-quality site to any LLM that does fetch the file.

No description block. The > quote line is the part most likely to be lifted into AI answers. Skip it and you waste the most-cited line in the file.

A real example

We shipped llms.txt + llms-full.txt for a 6-person Mumbai fintech client in February 2026. They had 22 blog posts and an active developer audience. After 8 weeks, server logs showed Claude Code and Cursor fetching their llms.txt regularly when their users asked AI tools about Indian KYC and UPI integration. Did this translate to leads? Indirectly — three engineering leads who later became clients said they "knew the brand from being recommended by Claude." Not direct attribution, but consistent enough to repeat the exercise for every dev-tool client.

For the founder's perspective on why standards like llms.txt matter even when adoption is slow, our founder Vivek Kumar writes about early-standard bets on his blog.

FAQ

Does Google read llms.txt?

No, not in any meaningful volume as of May 2026. Google-Extended does not fetch it. Google's John Mueller publicly compared the file to the old "keywords" meta tag, suggesting Google has no plans to use it.

Does Anthropic's ClaudeBot read llms.txt?

Server-log audits suggest no, not at scale. But Anthropic's own user-facing AI tooling (Claude Code, Claude Skills) does fetch llms.txt for context on docs domains. So Anthropic the company benefits from your file even if ClaudeBot the crawler does not log many requests.

Can I put llms.txt in a subdirectory like /docs/llms.txt?

You can, but the convention is the root domain. If you have a docs subdomain (docs.yoursite.com), put it at docs.yoursite.com/llms.txt. Most readers expect root placement.

Does llms.txt replace robots.txt?

No. They solve different problems. robots.txt tells crawlers what they can and cannot fetch. llms.txt tells LLMs what is most important on your site. Ship both, and use robots.txt to allow AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended) if you want to be cited.

Should I block AI crawlers in robots.txt and skip llms.txt entirely?

Only if your business model genuinely depends on traffic that is dying — affiliate sites where the user must click through to monetize. For every other model, allowing AI crawlers and shipping llms.txt is the dominant strategy. Blocking removes you from citation pools, which is the new top of funnel.

What is the file size limit?

There is no hard limit, but Anthropic's llms-full.txt is ~2 MB and tools generally handle that fine. We keep llms.txt under 10 KB and llms-full.txt under 5 MB.

Will llms.txt still matter in 2027?

Honest answer: nobody knows. Either it becomes a standard with Anthropic or OpenAI signing on publicly, or it joins the meta-keywords graveyard. The 2-hour investment hedges both outcomes.

Want llms.txt + a Full GEO Setup Done End-to-End?

We ship llms.txt + llms-full.txt + the 14 priority schema types + a content-extractability rewrite of your top 10 pages in one engagement. Fixed scope, 5 working days for a 50-page site. Comes with a 60-day re-check after the first AI citations land.

Book a GEO Setup Call
Tags:
llms.txtAI SEOGEOAI CrawlersDocumentationAnthropicOpenAI
Share this post:
Softechinfra Team

Softechinfra Team

Insights and updates from the Softechinfra team.