Disclaimer: This article is for education and for lawful, authorised professional research. Use these methods only where you have a legitimate purpose and a lawful basis, and follow the laws and platform terms that apply to you, including data protection law such as the UK GDPR and EU GDPR. See our Terms of Service.
TL;DR
- The problem: The current version of a company website, domain or public page rarely tells the whole story. Earlier records often say more.
- The idea: Treat historical data sets (archived pages, registration history, old public documents, past infrastructure) as anchors that connect an organisation's past to its present.
- The workflow: Move from slow, one-site-at-a-time manual checks to structured research across third-party data sources with UserSearch and OneScan.
- The outcome: You will learn to connect domains, trading names and infrastructure to the organisation behind them, and to write findings that stand up to review.
2.1 How Historical Data Sets Reach a Researcher
To use historical records well, you need to understand how they come into existence and how they end up in a searchable form. A record does not appear in a research tool the moment something changes on the web. It passes through a fairly predictable chain.
- Publication: An organisation publishes something: a website, a domain registration, a filing on an official register, a press release, a PDF brochure, a job advert.
- Change: Over time that material is edited, moved, renamed or taken down. Domains lapse, companies rebrand, pages are rewritten.
- Archiving: Web archives, registration-history services and data providers keep copies of earlier versions, often with a date attached.
- Aggregation: Third-party data sources index these copies and standardise them into a searchable format. This is the stage where most analysts work.
As an analyst, you are typically interacting with stage 4. Even so, you will meet very different kinds of historical material:
- Bulk lists: Large, flat lists of domains, hostnames or addresses with little context. Useful for quick yes/no checks, but they offer limited evidential value on their own.
- Structured records: Registration history, archived DNS, company register snapshots. These carry fields you can pivot on: registrant organisation, name servers, creation dates, hosting IPs, registered office changes.
- Unstructured documents: Archived web pages, PDFs, public text posts and old brochures. Messier to work with, but often the only place a former trading name, an old contact domain or a parent company is mentioned.
2.2 Why Old Records Still Matter
Why care about a web page from 2014? Because organisations are creatures of habit, just like people. Teams rarely rebuild everything from scratch; they carry forward what already works. We call this the "reuse pattern".
Consider a fictional online retailer, Kestrel Lane Trading, whose 2016 website footer listed [email protected] and a support phone number. That single archived page tells you:
- Naming habits: The organisation uses a
role@domainemail pattern, which helps you recognise its later domains. - Date anchor: The archive date gives you a point in time to compare with register filings and domain creation dates.
- Pivot point: If the same support number or the same analytics tag appears on a newer site under a different brand name, you have a strong lead that both sites are run by the same organisation.
This technique, the reuse pivot, is one of the most reliable ways analysts attribute a website or a network of domains to the business that operates it. Infrastructure, templates, tag IDs and contact details are expensive to replace, so they tend to survive rebrands.
2.3 Manual Historical Research (The Slow Way)
Before using a platform, it is worth understanding how to do this manually. The slow way teaches you what the data looks like, but it has real practical drawbacks.
Search Operators for Archived Public Documents
Organisations often leave older documents on their own servers long after the links to them have gone. Search operators can bring these back into view.
site:example.com filetype:pdf "annual report"
site:example.com "formerly known as"
"example.com" -site:example.com "registered office"What each line does: site: limits results to one domain, filetype: limits results to one document type, quotation marks force an exact phrase, and the leading minus sign excludes the organisation's own site so you see third-party mentions instead.
Web Archive Queries
The Internet Archive's Wayback Machine keeps dated snapshots of many public websites. Its CDX API lets you list those snapshots from the command line rather than clicking through a calendar.
curl "https://web.archive.org/cdx/search/cdx?url=example.com/*&from=2015&to=2019&fl=timestamp,original&collapse=urlkey&output=json"url=example.com/*asks for every archived path under the domain.fromandtorestrict the years returned.flchooses the fields to return (here the snapshot timestamp and original URL).collapse=urlkeyreturns one row per unique URL, which keeps the output manageable.
The limitation: Archives are patchy. Smaller sites may have only a handful of snapshots, and pages that relied on scripts often archive badly. Treat an empty result as "not archived", not as "never existed".
Local Bulk File Analysis
Some analysts download large historical data files to their own machine and search them with command-line tools such as grep or ripgrep.
rg -i "example.com" historical_domains_2019.csvrg is ripgrep, -i makes the search case-insensitive, and the final argument is the file to search.
The drawback: Holding large bulk files on a work laptop is a liability. If those files contain personal data, you have a data protection problem (storage, retention, security, lawful basis) that is out of proportion to the one answer you needed. It is usually better to query a source and keep only the relevant result.
2.4 Moving the Work into UserSearch
This is where UserSearch changes the workflow. Instead of visiting many archives one by one or keeping bulk files locally, you query third-party data sources through one account and keep only what is relevant to the case.
UserSearch 2.0 gives you access to 100+ third-party data sources across 18 Search types, including Domain Intelligence (ownership, history, favicon), Email Intelligence, IP Intelligence, Corporate Intelligence and Cyber Intelligence (Shodan Modules). You do not need a separate account with each provider.
The OneScan Workflow
- Choose the input: Enter a domain, a company email address, an IP address or a username.
- Select sources: OneScan runs one input across several selected data sources. Pick the sources that fit your question.
- Check the cost: The Credit cost is the sum of the selected sources and is shown before you run the search.
- Review: OneScan merges results with source attribution, so for every item you can see:
- Source: Which provider returned it.
- Context: Whether the record is a registration entry, an archived page or an infrastructure record.
- Timeline: Whether it reflects 2015 or last month.
If you have several inputs of the same kind, Bulk search lets you run up to five through one Module in a batch.
2.4a Structured and Unstructured Sources Side by Side
Why query more than one kind of source? Because they cover different parts of the public record, and knowing the difference makes the pivot work.
Unstructured sources index archived pages, public documents and text posts. They are good at finding a stray mention, such as a brochure PDF that names a parent company, or a public repository that lists a company's old domain in a configuration file.
Structured sources index registration and infrastructure records. They normalise fields such as registrant organisation, name servers, creation date and hosting IP, which lets you compare twenty domains on the same fields in one table.
Running both together through OneScan gives you the narrative from one and the joinable fields from the other. Remember that coverage and freshness depend on each third-party source, so results are leads to corroborate, not facts.
2.5 Fully Worked Scenarios
Theory is useful, but execution is what matters. Here are three research patterns using fictional organisations and placeholder identifiers.
Scenario A: The Supplier With a Short History (Due Diligence)
Context: A law firm acting for a client is reviewing a proposed supplier, Harrowby Components Ltd, as part of legal due diligence. The supplier's website, harrowby-components.example.com, looks polished, but the company claims fifteen years of trading.
The research:
- Initial search: You run the domain through Domain Intelligence. The registration history shows the domain was created eight months ago.
- The question: A new domain is not a problem in itself; companies rebrand. But it means the fifteen-year claim needs support from somewhere else.
- The pivot: An archived "About us" page on the new site mentions a former name, "Harrowby Engineering". A search operator query finds an older domain,
harrowby-eng.example.org. - The connection: Domain Intelligence on the older domain shows it was registered in 2011 and shares the same favicon and the same support phone number,
+44 7700 900123. - Corroboration: Corporate Intelligence and the official company register show a name change consistent with the rebrand.
- Outcome: The trading history is supported. The file records each finding with its source and date, and the client gets a clear, proportionate answer.
Scenario B: The Lookalike Domain (Brand Protection)
Context: The security team at a fictional travel company, Fenwick Coast Holidays, receives customer reports about a website offering "Fenwick" discounts at fenwick-coast-deals.example.net.
The research:
- The query: You enter the domain into Domain Intelligence and IP Intelligence through OneScan.
- The result: The domain was registered three weeks ago and resolves to
203.0.113.40. - The pivot: A Cyber Intelligence (Shodan Modules) search on that IP shows four other recently registered domains using travel-brand names on the same server.
- The analysis: The archived versions of those pages share the same template, the same analytics tag and the same contact email domain.
- Outcome: The team captures the pages with Forensic Capture (full-page capture, SHA-256 fingerprints, independent timestamps) and sends a documented takedown request to the hosting provider and registrar.
Scenario C: The Network of Online Shops (Journalism)
Context: A newsroom is researching a group of online shops selling the same unbranded supplements under different names. The editor wants to know whether one business runs them all.
The research:
- Domain search: You run the six shop domains through Domain Intelligence using Bulk search (five inputs in the first batch, one in the second).
- The hit: Four of the six share a registrant organisation name in older registration records, before privacy services were applied.
- Email research: Archived contact pages list
[email protected]. You run it in Email Intelligence and find it listed on a business directory page for a company. - Corroboration: The company appears on the official company register with a registered office matching the returns address printed on the archived shop pages.
- Outcome: The journalist has a documented chain from the shops to one registered business, with sources and dates for each link, ready for the right-of-reply stage.
2.6 Advanced Pivot Strategies
Beyond simple domain lookups, there are more advanced ways to use historical records to map how an organisation's online estate fits together.
The Shared Tag Pivot
Many websites carry an analytics or advertising tag ID in their page source. Web teams often reuse the same ID across every site they manage. If you find a distinctive tag ID in an archived page, search for it. You may find sister sites run by the same organisation, including ones that no longer link to each other.
The Hosting History Map
Historical DNS and hosting records often preserve which IP address a domain pointed to, and when. If you are researching five domains and two of them resolved to 198.51.100.25 in the same month, you have a correlation worth noting. Use IP Intelligence to check whether that address is shared hosting (weak link, many unrelated sites) or a dedicated server (stronger link).
Favicon Hashes: Anatomy of a Match
A site's favicon is a small image file, and many teams reuse the same one across projects. Researchers compute a hash of the favicon and search infrastructure data for other servers serving the same file.
pip install mmh3 requests
python3 -c "import mmh3, requests, codecs; r = requests.get('https://example.com/favicon.ico'); print(mmh3.hash(codecs.encode(r.content, 'base64')))"The first line installs the mmh3 hashing library and requests. The second fetches the favicon, base64-encodes it and prints the hash, which you can then use in a Shodan http.favicon.hash: query (a Shodan API key is needed for scripted searches). UserSearch's Domain Intelligence includes favicon data, so you can often skip this step.
Understanding these patterns lets you judge the strength of a link. A shared-hosting IP in common is weak; a shared tag ID, favicon and contact number together are strong.
2.7 Legal and Ethical Guardrails
Working with historical records still requires care and a clear purpose.
- Purpose first: Write down the lawful business purpose (due diligence, brand protection, security research, journalism) before you start, and keep your searches to what that purpose needs.
- Data minimisation: Avoid downloading bulk files. Use the UserSearch results view to keep only the records relevant to the case, and record the source and date of each one.
- Private mode or Forensic Mode: Choose deliberately. Forensic Mode stores search history and bookmarks in a Case so your work is reviewable; Private mode does not store history. Turn on two-step verification for sign-in.
- Corroborate: Historical records are leads. Confirm them with an official register or a second independent source before they go into a report.
2.8 Turning Old Records Into Findings You Can Defend
Historical data sets are the memory of the public web. They remember the old trading name, the lapsed domain and the server that five sites once shared. For the analyst, they are one of the most useful tools for attributing websites and domains to the organisations that run them.
The volume, though, is overwhelming, and manual checking is slow and incomplete. To turn scattered archives into findings, you need structure: one account, many sources, OneScan with source attribution, Cases to keep the work reviewable, and SargeBot to help draft a PDF report you then verify. Teams in research roles use the same workflow, and our guide to domain and infrastructure research goes further on registration data.
Stop guessing. Start researching with UserSearch at https://usersearch.com/.