A sitemap’s job is narrower than people assume — it’s not a list of every URL on your site, it’s a list of URLs you actually want indexed and consider worth a crawler’s time. Padding it with low-value pages doesn’t help those pages rank; it just dilutes crawl budget on pages that do matter.
What tends to not belong in a sitemap
- Tag archive pages, if your site has thin tag pages with only one or two posts each
- Author archive pages on a single-author blog (redundant with the homepage)
- Paginated comment pages, cart/checkout/account pages on WooCommerce sites
- Any page set to
noindex— including it in the sitemap while also noindexing it sends a mixed signal
Excluding post types from WordPress’s default sitemap
add_filter('wp_sitemaps_post_types', function($post_types) {
unset($post_types['attachment']);
return $post_types;
});
add_filter('wp_sitemaps_taxonomies', function($taxonomies) {
unset($taxonomies['post_tag']);
return $taxonomies;
});
robots.txt: fewer surprises than people expect
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /cart/
Disallow: /checkout/
Disallow: /my-account/
Sitemap: https://example.com/wp-sitemap.xml
The Allow: /wp-admin/admin-ajax.php line matters more than it looks — a lot of frontend functionality (AJAX add-to-cart, live search, infinite scroll) routes through admin-ajax.php, and blocking the whole /wp-admin/ path without this exception can quietly break how Google renders and evaluates the page.
A robots.txt mistake worth checking for specifically
Sites migrated from staging sometimes go live with Disallow: / still in place — a rule meant to keep the staging environment out of search results that never got removed after launch. This single line silently blocks the entire site from being crawled, with no error or warning anywhere, and it’s a genuinely common cause of “why isn’t my new site showing up in Google at all” weeks after launch.