URL Health: Find and Fix Broken Links in AI Answers
AI assistants sometimes link to pages on your site that do not exist. URL Health checks every link to your own site that your tracked AI answers exposed, confirms broken ones with two checks, suggests a possible fix and emails you a digest.
Key Takeaways
- AI assistants sometimes link to pages on your site that do not exist. URL Health checks every link to your own site that your tracked AI answers exposed and tells you which ones are broken.
- A link counts as broken only after two separate checks found it missing (404, 410 or soft 404). Timeouts, server errors and firewall blocks are shown as Unreachable and are never reported as broken.
- For a broken link, the Possible fix column suggests a working page on your site ("Did you mean ...?") and a redirect rule to copy, one wildcard rule when several links share a pattern. Your team verifies it; Ayzeo never changes your site.
- An email digest lists the broken links shown in answers, weekly by default or after every full scan. URL Health is included in the Pro and Enterprise plans and needs no setup.
Why Broken Links in AI Answers Matter
When an AI assistant recommends your brand, it often adds a link to a page on your site. Sometimes that page does not exist. The address looks right, sits on your domain and even carries the title of a real page, but it returns an error. The reader asked a buying question, got your brand as the answer and landed on a 404.
Nobody on your team sees these links. They are not linked from your site, so your own crawlers and SEO tools never find them, and the assistant keeps showing the same invented address answer after answer. URL Health finds them in the answers Ayzeo already tracks for your prompts and checks each one on your live site.
What Gets Checked
URL Health looks at the answers stored for your tracked prompts and collects every link that points at your own site: the project's domain and its subdomains, plus any other domains Ayzeo knows belong to the same brand. Links to other websites are never checked. Each link is filed under one or more of three layers:
- Answer text: links written into the answer itself, as a link, a full address, or your domain followed by a path. These are the links a reader can click, so they matter most.
- Visited: pages of your site the model opened while answering.
- Consulted: pages of your site the model looked at as sources.
One page is one row. Tracking parameters such as utm_source=chatgpt.com are removed, http and https count as the same address, and a trailing slash makes no difference. When the assistants wrote one link in more than one form, for example with and without tracking parameters, the row's detail view lists each form.
Only answers to live prompts count. If you delete a prompt, its answers leave every number on this page, and a link that appeared only in those answers is no longer shown or checked. Checking a link never runs a prompt and uses no AI queries.
When Links Are Checked
A check runs after each completed scan of your prompts, so it follows your project's scan schedule. Not every link is requested every time:
- New links are checked on the next check.
- Links that looked missing, and broken links, are checked on every check: the second check confirms or clears a missing link, and a fixed link shows up as soon as the next check finishes.
- Working links and redirects are rechecked at most once a week, and only while answers keep showing them.
- Links that could not be reached are rechecked after one, three and then seven days.
The checks come from Ayzeo's link checker, which identifies itself as AyzeoLinkCheck, paces its requests to your site and follows your robots.txt. The details for your web team are on ayzeo.com/bot.
What the Statuses Mean
| Status | What it means | Broken? |
|---|---|---|
| 200 | The page opened. | No |
| 200, content unverified | The page answered 200, but the check could not confirm that your site answers a made-up address with a 404, so it cannot prove this page is real. It is not counted as broken. | No |
| 301 → /path | The link redirects to a working page, and the status shows where it lands. A redirect you set up yourself shows here. | No |
| 404 or 410 | The page does not exist, confirmed by two separate checks at least 15 minutes apart. If a later check finds the page loading but cannot confirm it is a real page (for example it now lands where made-up addresses on your site land, and that could not be checked this time), it stays broken, with the note "last check: page loaded, could not confirm the fix". | Yes |
| Soft 404 | The page answers 200 but behaves like a missing page: it ends up where a made-up address on your site ends up (often the home page) while your working pages in that section open normally, or its title says the page was not found. Confirmed by two checks. A later check that could not look at that evidence again keeps it a soft 404, with the note "last check: could not recheck the soft 404". | Yes |
| 404, rechecking | One check found the page missing. The next check confirms it or clears it. Not counted yet, so a page that was down for a minute during a deploy never reaches you as broken. | Not yet |
| Unreachable (...) | The check got no reliable answer: a timeout, a server error such as HTTP 503, a firewall that blocked it, or a connection that failed. A link that redirects to another website whose answer could not be confirmed shows "redirects off your site": the checker never reads pages outside your site. A redirect to another of your own hosts that the check did not get to follow this time shows "redirect not followed". A new link that redirects to the page your site sends unknown addresses to shows "could not rule out a soft 404" until a check can compare it with a working page in the same section. The reason is in brackets. A status measured earlier is kept, with a note such as "last check: timeout". | Never |
| Unreachable (robots.txt) | Your robots.txt asks the link checker not to request that path, so it was not checked. A link checked before the rule keeps the status measured then, with the note "last check: robots.txt". | Never |
| Not checked yet | A new link, waiting for the next check. | No |
Broken always means 404, 410 or soft 404, confirmed twice. Every count of broken links, on this page, in the email digest, in reports and through MCP, uses that one rule.
Using the URL Health Page
Open URL Health from the AI Visibility group in the project sidebar. The header shows when the last check finished and has an Export CSV button.
- Date range, models and tags: the same filters as the analytics pages. They decide which answers are counted; tags appear when your project uses them.
- Shown in answers / All URLs: the default lens, Shown in answers, lists links that appeared in the answer text. All URLs adds the pages the models visited or consulted.
- Status chips: All, Broken, Redirected, Unreachable, Working.
- Search: find a link by any part of its address.
The table is sorted by In answers, the links readers saw most often first, and shows every status by default. Choose the Broken chip to list only the broken links, with the one readers saw most often at the top. Columns: URL with its layers (Answer text, Visited, Consulted), Status, In answers, Models, First / last seen, Possible fix and Last checked. A New marker flags a link that broke since your last digest when you arrive from the email.
Select a row to open its detail: the prompts whose answers showed the link (each with a link to Prompt Insights), the models, the first and last time it was seen, each form the assistants wrote it in when there was more than one and, for a redirect, every step of the chain.
Export CSV downloads every row that matches the current filters and search, not only the rows on screen.
Possible Fix: "Did You Mean ...?"
For a broken link, URL Health looks for the page the assistant most likely meant. It tries a few variations on your own site: the same address with one path segment removed, the same page in the URL structure your other markets use, a page with the same last path segment, or a page whose address matches the broken link's title. A candidate is shown only if it opened correctly on the same host when it was checked and its last path segment is the broken link's or matches its title. When nothing qualifies, the column says No pattern match.
Every possible fix can be copied with Copy rule as a redirect in the plain from -> to form redirect managers and spreadsheets take, for example /uk/clearance-uk/tents -> /uk/tents. When three or more broken links share the same pattern, for example the same extra path segment, the copied rule is one wildcard redirect such as /uk/clearance-uk/* -> /uk/* that covers all of them. No wildcard is offered when it would also catch a page on your site that works; each link then keeps its own rule.
A possible fix is a pattern match, not knowledge of what the assistant intended. Check it before you redirect: make sure the target is the page a reader of that answer would want. Ayzeo never changes your site.
How to Fix a Broken Link
You cannot correct the assistant's answer, but you can make the link work:
- Pick the target. Start from the possible fix, or choose the right page yourself.
- Add a permanent redirect (301) from the broken address to that page, or one wildcard rule when several links share a pattern. Your web team, CMS or CDN can set it up.
- Wait for the next check. Broken links are rechecked on every check, so after your next scan the row changes to 301 → /path. If you receive the digest with fix notices switched on, it lists the link as fixed.
The redirect keeps working when the assistants repeat the invented address, which they usually do: every answer that shows it now sends a reader to a real page.
The Broken-Link Email Digest
The digest is a summary of the links shown in AI answers over the last 30 days that are currently broken, grouped by model. Each link carries its status, how many answers and prompts showed it, a few of those prompts, when it was first and last seen, and the possible fix. Links that broke since your previous digest are marked New. A short section counts broken pages the models only visited or consulted, which no reader saw. The Open the broken links button opens URL Health filtered to exactly those links.
When your project is connected to Google Analytics, a broken link that received visits from AI assistants while it was broken also shows how many: in the last 30 days, counted from the day the link was confirmed broken when that falls inside them. A link without such visits shows no line, and a link listed under several models shows it once, under the first.
Set it up in Project Settings → Email Notifications (owners and admins):
- Broken links in AI answers: turns the digest on or off. It is on by default.
- Weekly summary (Mondays), the default, or After every full scan. A full scan runs every prompt on every model, so on a project that scans daily, "after every full scan" means a daily email.
- Also tell me when a broken link is fixed: adds the links that work again since the previous digest.
The digest goes to the project's owners and admins, except anyone who has switched off report emails in their profile. When nothing is broken and nothing was fixed, no email is sent.
Where Else Link Status Shows Up
- Badges: links to your own site in AI Sources & Destinations, in the answer view and in the Sources & pages tab of Prompt Insights carry a small badge when they are broken, rechecking or redirected. Select the badge to open that link in URL Health. Working links carry no badge.
- Monthly report: the monthly report email and the Enhanced PDF report include the number of broken links shown in AI answers.
- MCP: the
get_url_healthtool gives your AI assistant the same list, with totals per status and filters for layer, status, model and date range. See Connect AI Clients to Ayzeo (MCP Server).
When Your Site Blocks the Checks
Some firewalls and bot-protection services answer automated requests with a challenge or an HTTP 403. The affected links then show as Unreachable (blocked by the site), which is never counted as broken but also tells you nothing. When a host keeps refusing, URL Health stops checking it for that round and shows a banner such as "shop.example.com refuses our link checks (HTTP 403)". A banner that says a host does not answer in time, answers with server errors or has a robots.txt that could not be read points to the site itself: there is nothing to allowlist, and later checks try that host again. A banner that says the robots.txt of a host asks crawlers to wait too long between requests means its Crawl-delay is longer than three minutes, too slow for a check to finish: a Crawl-delay of three minutes or less in an AyzeoLinkCheck group of that file lets the checks through.
Ask your web team to allow the AyzeoLinkCheck user agent in the firewall. The user agent string, the request rate and how the checker follows robots.txt are documented at ayzeo.com/bot. Once they have, links not checked yet get their status on the next check. Links already shown as Unreachable wait for the first check one, three or seven days after their last one, depending on how many checks in a row failed.
Plans and Setup
URL Health is included in the Pro and Enterprise plans. There is nothing to set up: the first check starts within about an hour of your next scan, or the next morning (Berlin time) if the scan ends between 20:00 and 08:00, and until then the page says so. Every project member can open the page; the email settings are for owners and admins.
Frequently Asked Questions
Q: Does URL Health change anything on my website?
A: No. It only requests pages to see whether they open. Possible fixes are suggestions for your team; any redirect is set up by you.
Q: Why does a link I know is broken say "404, rechecking"?
A: A link counts as broken only after two separate checks at least 15 minutes apart found it missing. The first check marks it as rechecking, the next one confirms it. This keeps a page that was down for a moment out of your reports.
Q: Why is a link "Unreachable" instead of broken?
A: Because the check got no reliable answer: a timeout, a server error or a firewall block says something about the connection, not about whether the page exists. Unreachable links are rechecked later and are never counted as broken.
Q: Does checking links use AI queries or run my prompts?
A: No. URL Health reads the answers your scans already stored and requests the linked pages on your site. No prompt is run and no AI query is spent.
Q: I added a redirect. When will the link show as fixed?
A: After the next check, which runs after your next scan. Broken links are rechecked on every check, so the row changes to a redirect as soon as that check finishes, and the digest lists the fix if you switched fix notices on.
Q: Why is a link from an old answer missing?
A: URL Health counts answers to live prompts only. If the prompt behind an answer was deleted, its links are no longer shown or checked. The date range also applies: a link last seen before the selected period is not listed.
Q: What does the Google Analytics line in the digest count?
A: Visits (sessions) that came from AI assistants and landed on the link while it was broken, as recorded by your connected Google Analytics property: over the last 30 days, and only from the day the link was confirmed broken when that falls inside them, so visits to the page while it still worked are not counted. It appears only when Google Analytics is connected and recorded at least one such visit, once per link. A server that answers a missing page without your analytics tag records no visits there.
Q: Can I get the digest after every scan instead of weekly?
A: Yes. In Project Settings, Email Notifications, choose After every full scan. The digest then follows each scan of every prompt on every model; a scan of a single prompt does not send one.