There is a file on your site, a few hundred bytes of plain text, that decides whether you exist to the systems your buyers now ask. Most teams have never opened it. Fewer still know what their CDN quietly bolted on top of it.
That file is robots.txt, and it is a fair place to start, because the next stretch of B2B growth belongs to a specific kind of company: the machine-readable one. Buyers ask engines. Engines read the web. What they can read, they can cite. What they cannot read does not exist.
We have been watching crawlers grow up for 26 years, since the first AdWords auctions. This is the fifth migration we have worked through, from search to SEO as a service to social to mobile to AI answers, and it rewards the same discipline the first one did: make it effortless for a machine to establish who you are and why you are credible.
Here is what that means in practice, with honest labels on what is proven and what is still speculative.
What engines read, and what they skip
Start with three hard truths.
Gated content is unread. Your best case study, your licensed analyst report, your benchmark data: if it sits behind a form in a PDF, no answer engine will ever quote it. A form is a wall. Walls do not get cited.
Script-built pages are half-read. Most of the crawlers behind AI answers do little or no JavaScript rendering today. If your proof points, pricing, or product descriptions render client-side, assume a meaningful share of the machines your buyers consult never see them. Server-render whatever you would be upset to lose.
Paid placements are never cited as evidence. Sponsored slots buy attention from humans. Engines assemble answers from organic, corroborated sources. You cannot buy your way into an answer, which is exactly why answers are worth earning.
Schema that earns its keep
Structured data will not rescue weak content, and we will not claim a straight line from markup to model output. What schema provably does is remove ambiguity, and ambiguity is the enemy of citation. Four types matter most in B2B.
Organization. Your legal name, your brands, your official profiles tied together with sameAs, one logo, one address, one description. This is entity disambiguation, and it is foundational.
Product. What it is and who it is for, in plain language, marked up so the offer cannot be confused with the company or the category.
FAQPage. Real questions your buyers ask, answered in liftable form. Not filler. The questions your sales team hears on calls.
Article. Authors, dates, provenance. Machines increasingly weigh who said something and when.
Honest label: schema's effect on classic search features is well established. Its direct effect on AI answers is partly inferred from how these systems retrieve and ground. It is also cheap, standard, and low-regret. Do it anyway.
One company, one story, everywhere
Entity consistency is the unglamorous core of this work. Same name, same claims, same numbers, on your site, your LinkedIn page, the directories, the review platforms, the partner listings, the press. A model that finds three versions of your category, your customer count, or your founding year has an easy option available: leave you out and cite a competitor whose facts agree with each other.
Audit the top twenty places your company is described. Reconcile every conflict, then assign an owner so it stays reconciled. This is boring. It is also one of the highest-leverage projects on this list.
Let the right bots in
Now back to that file. Robots.txt, plus whatever bot management your CDN or security layer enforces, is your access policy for the machines that answer your buyers. Many security products block AI crawlers by default, which means plenty of companies have opted out of AI answers without anyone ever deciding to.
Decide deliberately. There are crawlers that gather training data, crawlers that fetch pages for live answers, and agents that browse on a buyer's behalf. GPTBot, ClaudeBot, PerplexityBot, and their peers each publish documentation, and the lists change, so review your policy quarterly. Whatever your position on training access, blocking the retrieval bots that assemble live answers is blocking your buyer mid-question.
Proof that can be quoted
If a claim matters in a sales call, it should exist on an open page. Ungated case studies in HTML. Pricing structure explained, even where exact numbers vary. Comparison pages that state your side of the argument. Security and compliance posture in text, not only inside a portal. This is the substance layer of AI search optimization: engines build answers from evidence, and the evidence has to live somewhere a machine can reach.
The llms.txt question
Every month someone asks about llms.txt, the proposed standard for handing models a curated index of your site. Our honest read: worth watching, cheap to try, unproven. Adoption across the major engines is mixed at best, and we have seen no convincing public evidence that it changes answers today. If it costs an afternoon, fine. If it displaces any of the fundamentals above, you have traded the proven for the fashionable.
The readable company wins
None of this is exotic, and that is the point. The last migration rewarded companies that wrote for readers and tuned for crawlers. This one rewards something plainer: say true things, in open formats, consistently, everywhere, and let the machines carry them to buyers you will never meet.
So run the quiet test. If a machine read everything public about your company tonight, what could it say with confidence?
If you are not sure, that is exactly what our AI Search Diagnostic answers. Talk to Danton and we will show you what the engines can see, and what they are missing.
