Joost de Valk, who built Yoast SEO, recently wrote that almost nobody gets XML sitemaps right. He pointed his Sitemap Inspector at Adobe, Anthropic, Manchester City, GOV.UK, X and Semrush, and found broken sitemaps nearly everywhere.
I wanted to know if Ireland does any better. So I ran the same kind of checks against 30 of the best-known websites in the country: state bodies, banks, airlines, utilities, retailers, universities and media. The answer is no, not really.
What "right" means
A sitemap has one job. Google's documentation puts it simply: list the URLs you want to see in search results. They should be canonical URLs, and they should load. Every entry is a request for Google to spend a crawl on that page.
That gives two kinds of problems. Some things are allowed but not ideal: missing lastmod dates, dates without a time, or thousands of pages sharing one date. Other things are simply wrong: URLs that redirect, return errors, carry a noindex, or declare a different page as canonical. Each of those hands search engines a URL, then tells them not to use it.
For every site I read robots.txt, parsed the sitemap or sitemap index, and then fetched the first 100 listed URLs to check status codes, redirects, noindex and canonicals. I checked all of them on 27 September 2026, from Dublin, with a normal browser user agent. Four sites (An Post, Tesco, the Irish Independent and Dublin Bus) blocked automated requests, so they are left out. Where results came from geo-redirects or bot walls rather than the sitemap, I've ignored them.
Transport for Ireland: English pages, Irish canonicals
Transport for Ireland's sitemap index lists 2,744 URLs. Of the first 100 I checked, 53 are English news articles whose canonical tag points to the Irish-language version. The page at /news/public-transport-fares-determinations-for-2016/ declares /ga/news/public-transport-fares-determinations-for-2016/ as its canonical.
The same page says lang="en-US" and carries hreflang tags for en, ga and x-default. So the sitemap says "index this English page", the hreflang says "this is the English version", and the canonical says "no, index the Irish one". Three signals, two answers.
This is a common multilingual plugin mistake, and it's exactly the kind of thing nobody notices by looking at the website.
Brown Thomas: sale pages that point somewhere else
Brown Thomas lists over 48,000 unique URLs across 20 sitemap files. Of the first 100, 81 declare a different page as their canonical. Almost all of them are "exclusive preview" sale URLs such as /exclusive-preview/platinum-sale-15-off/the-hand-treatment-100ml/577555.html, and each one canonicalises to the regular product page under /beauty/… or /brands/….
The canonical is doing its job: the sale URL is a duplicate of the product. The problem is that the sitemap lists the duplicate, and puts all 78 of those sale URLs at the very top, which is the first thing a crawler reads. It's a small share of a big file, but it shows the sitemap and the canonical logic aren't talking to each other.
eir: most of the list at once
eir has two sitemap indexes, /sitemap.xml and /sitemap_index.xml, and robots.txt mentions neither. Between them they list about 10,000 URLs. One of the entries in /sitemap.xml is /business/sitemap.xml, which is itself a sitemap index. Google's rules are clear on that: a sitemap index can list sitemaps, not other indexes.
Of the first 100 URLs I checked:
- 28 redirected to 19 different destinations. Several land on
/opencms/export/sites/default/…paths, which look like leftovers from a previous CMS. - 5 returned errors: four 404s, including
/broadband/about/and/email/, and a 500 on/mobile/phones/prepay/. - 1 had a
noindex. - 23 declared a different page as canonical. The terms and conditions page points to the homepage. The cookies list points to the mobile bill-pay page.
Only 36 of the 100 were URLs the site actually wants indexed. And 7,012 of the URLs share the same date-only lastmod of 15 May 2024.
Discover Ireland: "undefined"
Fáilte Ireland's Discover Ireland sitemap lists 3,111 URLs. For 2,632 of them, the lastmod value is the literal word undefined. That is a JavaScript variable that never got a value, written straight into the XML. It isn't a valid date, so crawlers simply throw it away.
Then there are the pages that don't exist. Thirteen of the first 100 URLs returned a 404, and every one of them is a Dublin place: Balbriggan, Baldoyle, Ballinteer, Ballsbridge, Blackrock, Blanchardstown, Booterstown, Cabinteely and more. The national tourism site asks Google to index pages about Dublin that aren't there.
Dublin Airport and Bord Gáis: pages that have gone
Dublin Airport lists 1,103 URLs. Eight of the first 100 return a 404, including /corporate/airport-development/north-runway/about and the north runway news page. More than half of all entries, 582, share one date: 12 March 2026.
Bord Gáis Energy lists 418 URLs with no lastmod on any of them. Of the first 100, eight redirect and five return a 404. The dead ones include old product pages for Hive sensors and the Hive Active Plug, and a business customer stories page.
Irish Rail: a sitemap that isn't there
The simplest failure is having no sitemap at all. Irish Rail's robots.txt doesn't mention one. Ask for /sitemap.xml and you get a 200 response containing a small HTML page with a meta refresh to /en-ie. A crawler asking for a sitemap is told "here it is", then handed a redirect page.
UCD: HTML pretending to be XML
UCD has a sitemap index of 98 department sitemaps, none with a lastmod. Three of them, for Planning, Landscape and Electrical Engineering, return an HTML web page titled "Google Sitemap", served with a text/xml content type. It's a human-readable page listing links, which is useful to people and unreadable to crawlers.
RTÉ: one character
RTÉ's sitemap is otherwise in good shape. The index lists 125 child sitemaps, the dates are real timestamps, and the first 100 URLs I checked were all clean.
But the index declares its namespace as https://www.sitemaps.org/schemas/sitemap/0.9. The protocol requires http://, and an XML namespace is an exact string, not a web address. Strict parsers see an unknown document with zero sitemaps in it. Mine did, until I made it more forgiving. The child sitemaps use the correct namespace, so it's a one-character fix in one file.
Allowed, but not ideal
Two patterns came up again and again that aren't errors, but make lastmod meaningless:
- gov.ie lists 119,301 URLs. Of those, 90,237 carry one of two dates, 11 or 12 April 2025. That looks like a site migration stamping every page at once. The first 100 URLs were clean.
- Dunnes Stores gives every one of its 5,614 entries the same
lastmod: today. The date is generated when the sitemap is requested, so it says nothing about when a page changed. Also, 323 URLs appear more than once.
Google has said it uses lastmod only when it's consistently accurate. If every page changed today, none of them did.
Who got it right
Revenue lists 15,032 URLs with real timestamps down to the second, no date shared by more than four pages, and the first 100 URLs I checked were all clean. Citizens Information and the Guinness Storehouse also came back with 100 clean URLs out of 100. So it can be done, even on large government sites.
Why this keeps happening
None of this is hard. Joost's explanation fits what I found: the sitemap usually comes from one system, while redirects, canonicals and noindex live somewhere else, and nobody checks whether the two agree. The file validates, so nobody opens it again.
You can see it in the results. Transport for Ireland's multilingual setup decides canonicals, and the sitemap generator doesn't ask it. Brown Thomas's commerce platform knows sale pages are duplicates, and its sitemap doesn't. Discover Ireland's front end fills in the dates, and nothing checks that a date came out.
On WordPress this is easier to get right, because a good SEO plugin builds the sitemap from the same data it uses for canonicals and noindex. That's the point: one system should own all of those signals.
What to do about your own site
If you run a small business website, the stakes are smaller but the logic is the same. Open yoursite.ie/robots.txt and check that it names a sitemap. Open the sitemap and look at a handful of URLs. Do they load? Do they redirect? Are old pages you've deleted still listed? Then check Google Search Console for "Submitted URL" warnings, which is Google telling you exactly this.
If big Irish brands with web teams get this wrong, there's a good chance your sitemap has a problem or two as well. It's one of the first things I check.