Piceci Services
โ† All articles
GEOTechnical SEOSchema

llms.txt and schema markup for AI crawlers: a technical setup guide

The technical layer of AI visibility: robots directives, llms.txt, structured data and server rendering, with checks.

Piceci Services/August 26, 2026/5 min read
โ„– 04Piceci ยท Journal
GEOEssay

llms.txt and schema markup for AI crawlers: a technical setup guide

Content strategy fails when the technical layer blocks it. This is the setup we ship on client sites: explicit crawler policy, a llms.txt that describes the business, entity-level structured data, and rendering that does not require JavaScript.

Who asks this question

  • Technical teams responsible for AI and search visibility
  • Companies whose pages are never cited despite good content
  • Developers auditing an existing marketing site
  • Businesses operating in several languages or regions

How to implement it

  1. Declare AI crawler policy explicitly in robots.txt, allowing the ones you want.
  2. Publish llms.txt with a short business description, locations, services and key URLs.
  3. Add Organization schema with legal name, address, certifications and sameAs profiles.
  4. Add Service or Product schema on commercial pages and Article on posts.
  5. Server-render content so answers exist in the initial HTML response.
  6. Maintain a sitemap that lists every language variant with correct hreflang.
  7. Validate structured data and re-run the check after each release.
  8. Monitor server logs for AI crawler user agents to confirm access.

What to measure

  • AI crawler requests per week in server logs
  • Structured data validation errors
  • Percentage of pages fully rendered server-side
  • Sitemap and hreflang coverage of published URLs
  • Pages cited by assistants after the technical fix

Common mistakes

  • Default-blocking every AI user agent in robots.txt
  • Structured data with a different company name than the footer
  • Client-side rendering for the main content of key pages
  • A llms.txt that is a link dump with no description of the business
  • Broken hreflang between language variants

Frequently asked questions

What is llms.txt used for?

It is a plain-text file at the site root that describes the business and points to the pages worth reading. It is not an official standard and not a ranking factor, but it is a low-cost way to make your positioning and key URLs machine-readable.

Should I block AI crawlers?

Only if you do not want to be cited. If discovery inside assistants matters commercially, allow the crawlers you want and keep the sensitive areas of the site disallowed as usual.

Does schema markup help AI visibility?

It helps disambiguate your entity: legal name, locations, services and certifications. That improves attribution accuracy and makes selection more likely, though it does not guarantee a citation.

Piceci Services is a HubSpot Solutions Partner and ISO 27001 certified consultancy operating from Milan and Dubai. If you want this reviewed against your current setup, book a call and we will walk through it with your data.

Work with us

Want a second pair of eyes on your CRM?

We run free 15-minute reviews โ€” no slides, no pitch, just a look at your setup.

Book a review โ†’