Skip to main content
← Architecture
Shipped and running

Your rebuild will drop URLs. Google will tell you in six weeks.

A rebuild changes how pages are made. The old addresses stop existing without anyone deciding they should, and Google notices before you do. Here is the check that proves no URL moved, the four things it cannot see, and the failure that cost more than any of them - from a 10,723-page site that went through it.

6 September 2026

Check it yourself. curl the sitemap before and after, then diff the URL sets

A terminal showing the sitemap diff: 10,722 URLs before, 10,723 after, nothing dropped

Nobody rebuilds a website intending to lose its pages. It happens anyway, and it happens in a specific way: an address gets tidied because the new one reads better, a template gets a cleaner path, a folder is dropped because the new system does not need it. Every one of those decisions is sensible on its own. Together they are why traffic falls off a cliff six weeks after a launch that nobody thought was risky.

One word before the example. A page’s URL is its address - the part after the domain name, like /word/serendipity. Google’s memory of a page is tied to that exact address. Change it, and as far as Google is concerned the old page has died and a new, unknown page has appeared. The years of ranking the old address earned do not move across by themselves.

A dictionary site rebuilds

Take an online dictionary. It has forty thousand pages, one per word, and none of them was written by hand. A template builds each page from a database row: the word, its meaning, its pronunciation. The address of each page is built the same way - a small function that takes the word and turns it into /word/serendipity.

The site is rebuilt on a new framework. The developer, tidying as she goes, notices that the old addresses have an untidy trailing part that the new system does not need, and simplifies the function. /word/serendipity-1 becomes /word/serendipity. Cleaner. Better. And forty thousand addresses have just changed at once, because they all come from one line of code.

Nothing warns her. The build passes. The new site looks right, because it is right on its own terms. The old addresses simply stop existing, and the only party who notices is a search engine that will not say anything for weeks.

Why nobody notices for six weeks

A dropped address does not throw an error. No test goes red. The people who built the site are looking at the new site, where every link works. The people who would hit the old addresses - visitors arriving from Google, from old bookmarks, from links on other sites - are strangers, and strangers do not file bug reports. They see “page not found” and leave.

Google does notice, slowly. It re-visits the old addresses, finds nothing, and after a few weeks removes them. The traffic graph bends downwards. By then the team has shipped a fortnight of other changes, and the question “which of these did it?” has no cheap answer. The fix has to happen before the launch, or it costs ten times as much afterwards.

Why “keep the URLs” is not a plan

Everyone says it at the start of a rebuild. On a twenty-page site it even works, because twenty is a number a person can hold in their head and check by eye before launch.

On a site with thousands of pages built from templates, it is exactly where things go wrong. The templates are the danger: change one line in the function that builds an address and every page in that family moves at once. The blast radius of a one-character edit is the whole site. So “keep the URLs” has to become something a machine checks, not something a person promises.

There are two halves to that, and they are different jobs. The first is a test inside the build that fails when an address, a title or a heading changes - that protects you while you work. The second is proving, from outside, that the site which actually went live kept them. This piece is about the second, because it is the half you can run against any site, including one you did not build.

The check: two lists, compared

It is deliberately simple. Every site that cares about search already publishes a complete list of its own addresses. It is called a sitemap, it lives at /sitemap.xml, and search engines read it to learn what the site contains. Because it is a full list of addresses, it doubles as a record of what the site looked like on a given day. Save it before the launch, and again after:

two snapshots
1# Before the deploy, and again after it. Nothing clever  - 
2# the sitemap is the list of URLs you already publish.
3curl -s https://example.com/sitemap.xml > before.xml
4
5# ...deploy...
6
7curl -s https://example.com/sitemap.xml > after.xml

Then compare the two lists. Not the totals - the lists. A total that matches can still hide one page dropped and another added, which is the case that hurts most, because the number looks reassuring.

diff.py - twelve lines
1import re, sys
2
3def urls(path):
4    xml = open(path, encoding='utf-8', errors='ignore').read()
5    return set(re.findall(r'<loc>([^<]+)</loc>', xml))
6
7before, after = urls(sys.argv[1]), urls(sys.argv[2])
8
9print('before:', len(before))
10print('after: ', len(after))
11print('gone:  ', len(before - after))
12
13# The count is reassurance. The list is the answer.
14for u in sorted(before - after):
15    print('  DROPPED', u)

Three numbers and a list. The list is the part that matters: it names every address that existed before the launch and does not exist after it, which is the question everyone asks two months too late.

FamilyShare of the siteBeforeAfter
/convert/
9,0459,045
/time/
1,0001,000
/dst/
248248
/country/
247247
/abbreviation/
136136
/utc/
3838
static
89
Totalone page added, none removed10,72210,723
Before and after, per template family, from a real migration. One page appeared - a new static route - and nothing was removed.

What the check cannot see

A passing diff should not make anyone relax, and this is the part usually left out. An address is one of at least five things a rebuild can change, and the sitemap only knows about one of them.

  • The URL/convert/EST_to_FJTa diff catches this
  • The page titlewhat the result shows in searcha diff does not
  • The H1the heading Google matches against the querya diff does not
  • The canonicalwhich version of the page countsa diff does not
  • Internal linkshow authority moves between pagesa diff does not
  • The redirect mapwhat happens to anything that did movea diff does not
A sitemap lists addresses. Everything else that decides whether a page keeps its rankings is invisible to it.

A page can keep its address and lose its title, its main heading, its “canonical” tag - the line in the page that tells Google which address is the official one when several show the same content - and every internal link that pointed at it. The diff returns zero and the rankings still go. A build test can cover titles and headings too. The canonical tag and the links between pages are usually guarded by nothing but attention, which is not a guarantee and should not be described as one.

Where this breaks

The sitemap must come from the same place as the pages. If it is written by hand, or built by a separate job, it can describe a site that no longer exists and the diff will happily confirm a fiction. It only works as evidence when the file is generated from the same data that builds the pages.

An address in the sitemap can still be broken. Present in the file, missing on the server: the diff passes. So take a sample of the survivors and check what they actually return:

sample forty, expect 200
1# A URL can survive the diff and still be broken: present in the
2# sitemap, 404 at the server. Sample the survivors.
3shuf -n 40 <(grep -o '<loc>[^<]*' after.xml | sed 's/<loc>//') |
4  xargs -P 8 -I{} sh -c 'printf "%s %s\n" "$(curl -s -o /dev/null -w %{http_code} "{}")" "{}"' |
5  grep -v '^200' || echo "all sampled URLs returned 200"

Pages outside the sitemap are outside the check. Old campaign pages, PDFs, anything only ever linked from an email - none of it is in the file, so none of it is protected. For those, the equivalent snapshot is a Search Console export of every page that had impressions, taken before the work starts.

And none of it matters if the site is down. The site these numbers come from was unreachable for about a week shortly before the migration - not a proper “we are down, come back later” answer, which tells a search engine to wait, but a dead connection, which tells it nothing. Every address was preserved perfectly throughout, and preserved addresses on an unreachable server are worth exactly nothing. The fix afterwards was ordinary: the web server now shows a maintenance page with a 503 status and a “retry after” header, so an outage looks like an outage rather than an absence. That was the more expensive lesson of the two.

The site these numbers come from

The dictionary is made up. The migration is not. timezones.in is a site with 10,723 pages generated from six templates - 9,045 of them are time-conversion pages built from one address function, which is exactly the one-line blast radius described above. It moved to Next.js 16 in one week in September 2026. The first sitemap snapshot was taken on the morning of 6 September and the second after the deploy the same day: 10,722 addresses before, 10,723 after, none dropped. The build test stopped the addresses moving; the diff proved they had not; and the second of those is checkable by anyone, which is the next section. Whether the rankings held is a separate claim. The site was only verified in Search Console that month, so there is no “before” to compare against - an honest gap, and one worth writing down rather than leaving out.

Run it on your own rebuild

The whole method is two downloads and twelve lines of Python, and it works on any site with a sitemap - including one you are being asked to take over from another agency. Take the first snapshot before anyone touches anything. It is worthless afterwards, and afterwards is when you will want it.

You can run the same check against the site above. Its sitemap is public and split by template family: timezones.in/sitemap.xml. The fuller story of that site - what it is, what it cost, what else went wrong - is on its project page.

The habit to take away is smaller than the tooling. Before any change that touches how pages are made, save the list of addresses the site publishes today. After the change, compare. If the list is the same, you have proved one thing and only one thing - that the addresses survived - and that is still the one thing most teams find out the hard way, from a graph, six weeks after it was cheap to check.

Related on this site: when a redirect rule swallows a page that still works, web application development, legacy modernisation and hiring Next.js developers.

Questions about this

Usually because page addresses changed. A new framework or template tidies a URL, the old address stops existing, and Google drops the page from its index. Nothing in the build fails, so the first sign is a traffic graph a few weeks later.

Not answered here?

Ready to Build Something
That Actually Works?

Stop patching legacy code. Let's engineer a platform that scales with your ambition.