← Back to the library

GUIDESEO & AI search

Get found by ChatGPT and Perplexity: what to check on your website

A 45-minute check with no installs: ask the AI tools your customers' questions, check robots.txt for the right crawlers, score your pages with PLUTO, fix one thing and log it.

WHAT YOU’LL GET

A list of the questions you are and are not quoted for, a checked robots.txt, a baseline PLUTO score, and one fix published and logged.

WHO IT’S FOR

Business owners with a website, and whoever edits it, who want to know why AI search does not mention them.

DIFFICULTY

Beginner

TIME

About 45 minutes

WORKS WITH

ChatGPT Perplexity Any AI chat

Get the full file

The whole resource as one Markdown file for your notes or your AI workspace.

FREEOpen the tool

What this checks and who it is for

More customers now ask ChatGPT or Perplexity "who does aircond servicing near Subang?" instead of scrolling Google. Those tools search the web and quote a few sources. If your site blocks their search crawler, or never says plainly what you do and where, you make it much harder to be one of those sources.

This guide is a check you can do in about 45 minutes, without installing anything: ask the tools, check the door (your robots.txt), check the page itself, fix one thing, write it down. It is for business owners with a website and whoever edits it. Nothing here guarantees you will be quoted; it removes the reasons you cannot be.

What you need first

  • Your website address.
  • ChatGPT with search, and Perplexity (free accounts are enough to look).
  • A spreadsheet or notes page for the results.
  • Someone who can edit the site, for the fix.

Step 1: Ask the questions your customers ask

Write down five to ten real questions, in the words customers use. Mix these kinds:

  • the service plus your area ("kedai repair iPhone Bangsar"),
  • a comparison ("which is better, X or Y, for a small office"),
  • a price question ("how much is ... in KL"),
  • your business name.

Ask each one in ChatGPT (search on) and Perplexity. For each answer, note: were you mentioned, were you linked as a source, and who was linked instead. The businesses being quoted show you what a quotable page looks like in your trade.

Step 2: Check the door (robots.txt)

Open https://yourbusiness.com/robots.txt in a browser. It tells crawlers what they may read. Look for these names, checked against each company's own documentation on 30 Sep 2026:

  • OAI-SearchBot: the crawler for ChatGPT's search features. OpenAI says sites that block it will not appear in ChatGPT search answers.
  • GPTBot: OpenAI's training crawler. OpenAI treats it separately, so you can allow search and still refuse training.
  • PerplexityBot: the crawler that surfaces and links sites in Perplexity search. Perplexity says it is not used to train AI models.
  • Google-Extended: controls whether Google may use your content for Gemini. Google says it does not affect whether you appear or rank in Google Search.

What the rules mean in practice, per the same documentation:

  • Disallow: / under OAI-SearchBot: OpenAI says your pages will not be shown in ChatGPT search answers, though they can still appear as plain navigational links.
  • Disallow: / under PerplexityBot: Perplexity's search crawler stays out. Perplexity also runs a separate fetcher, Perplexity-User, which visits a page when a user asks about it, and says that fetcher generally ignores robots.txt. So blocking PerplexityBot is not a complete block.
  • No line for a bot is not a block. A crawler that is not named follows the User-agent: * section, if there is one. If nothing there disallows it, it is allowed.

If you want to allow the two search crawlers explicitly, the lines look like this:

TEXT
User-agent: OAI-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

Careful when you edit: a crawler with its own section follows only that section and ignores the User-agent: * rules. If your file already keeps bots out of an admin area or a private folder, copy those Disallow lines into the new sections too, and do not delete existing rules you do not understand.

Blocking training crawlers is a fair business choice; blocking the search ones by accident is the common mistake. OpenAI notes a robots.txt change can take about a day to show up in its search.

Step 3: Get a baseline score

Run your address through PLUTO, the free checker I built. It reads up to 12 of your public pages and scores six things: whether AI crawlers can reach you, whether pages are written as answers, whether they say specifically what you do, business details machines can read, technical health, and reasons to trust and contact you. The score shows on screen before any contact form. It is a diagnosis, not a ranking promise. The library entry See what people and search engines find on your website with PLUTO walks through each score.

Step 4: Check the page a customer lands on

Open your home page or main service page and check, honestly:

  • The first paragraph says what you do, for whom, and where. "We provide solutions" fails; "Aircond servicing and repair for homes and offices in Subang Jaya and Shah Alam" passes.
  • Opening hours, phone or WhatsApp, and address are written as text, not only inside an image.
  • Common questions are answered on the page in short, plain answers. Real questions from Step 1 are the best source.
  • Prices or price ranges are there, if you are willing to publish them.
  • Business details are marked up as structured data. Check in Google's Rich Results Test, which runs the page like a browser does.

Step 5: Fix one thing, then write it down

Pick the single biggest gap: usually a blocked crawler or a vague first paragraph. Draft the fix with AI from your real facts only, have a person read it, publish it, and log the date and the change. Ask the same Step 1 questions again in a few weeks and compare. Search tools change their answers for many reasons, so the log is what lets you tell a change from noise.

Going deeper with an agent

If you use Claude Code or Codex, the ai-seo skill in marketingskills turns this into a fuller audit: citation patterns, page structure and crawler access, with a prioritised list. Ask it to stay read-only on the first run.

Watch-outs

  • Do not fake it. Invented reviews, made-up FAQs or copied competitor text can hurt trust with people and machines. Use only true facts.
  • No guarantees. Allowing crawlers and writing clear pages makes you quotable. It does not make you quoted.
  • Changes go live through a person. Robots.txt mistakes can hide a whole site; have whoever manages the site check the edit.

Sources

OpenAI crawler documentation (developers.openai.com/api/docs/bots), Perplexity crawler documentation (docs.perplexity.ai), Google's common crawlers page (developers.google.com), all read on 30 Sep 2026. The check order draws on the ai-seo skill by Corey Haines, in my own words.