Skip to main content

Grading OSINT Sources with the Admiralty Code: A Practical Guide

How the Admiralty (NATO) system grades source reliability A to F and information credibility 1 to 6, why most open-source items start at F6, how corroboration moves grades, and how to grade results from aggregated data sources.

· By UserSearch Team · 14 min read

Disclaimer: This article is for education and for lawful, authorised professional research. Use these methods only where you have a legitimate purpose and a lawful basis, and follow the laws and platform terms that apply to you, including data protection law such as the UK GDPR and EU GDPR. See our Terms of Service.

TL;DR

  • The Admiralty Code for OSINT source grading (also called the NATO system) grades two things separately: how reliable the source is (A to F) and how credible the specific piece of information is (1 to 6).
  • Most open-source material starts at F6: an unknown source making a claim nobody has confirmed yet. That is a starting grade, not a verdict.
  • Independent corroboration raises information credibility quickly. Source reliability rises slowly, over a history of reporting, and a single confirmation should not move it.
  • When results come from an aggregator, grade the underlying data source, not the platform you used to reach it. UserSearch's OneScan keeps source attribution on every result so you can.
  • We include a worked grading table, code for recording grades as MISP tags and STIX confidence values, and two scenarios.

The Problem with "It Looks Legit"

The Admiralty Code is a two-part grading system for OSINT and other intelligence: it rates how reliable a source is (A to F) and, separately, how credible a specific piece of information is (1 to 6). Ask five analysts why they believe a finding and you will usually hear some version of "it looks legit". The account is old. The website is professional. The channel has 40,000 subscribers. None of those are reasons. They are impressions, and impressions do not survive a hand-over to another team or a challenge from a sceptical manager.

OSINT makes this worse because volume is cheap. A single username search can return hundreds of hits across platforms, aggregators and archives, and every one of them arrives with the same visual weight on the screen. Without a grading scheme, the loudest source wins. With one, you can say precisely why a company register entry outranks a forum post, and precisely what would need to happen for that forum post to count.

The oldest and most widely used scheme for this is the Admiralty system. This guide explains how it works, why so much open-source material starts at the bottom of the scale, how corroboration moves grades up, and how to grade results that reach you through an aggregated platform.

How Does the Admiralty Code Grade OSINT Sources and Information?

The Admiralty system, formalised in NATO intelligence doctrine (AJP-2.1, under STANAG 2511) and used across the NATO and Five Eyes intelligence communities, rates every item on two independent axes. A letter describes the source. A number describes the information. You combine them into a two-character grade such as B2 or F6.

The wording below follows the open MISP admiralty-scale taxonomy, which many threat intelligence teams use as their machine-readable reference.

Source reliabilityMeaningInformation credibilityMeaning
ACompletely reliable1Confirmed by other sources
BUsually reliable2Probably true
CFairly reliable3Possibly true
DNot usually reliable4Doubtful
EUnreliable5Improbable
FReliability cannot be judged6Truth cannot be judged

Two rules make the system work, and both are routinely broken.

The axes are independent. A reliable source can pass on a doubtful claim (B4), and an unknown source can report something you have confirmed elsewhere (F1). If you find yourself only ever writing A1, B2 and C3, you are grading one thing twice.

F and 6 are not bad grades. They mean "no basis to judge yet". E and 5 are the negative grades. Treating F6 as "unreliable, false" throws away leads; treating it as "probably fine" invites mistakes. It means: do the work.

The UK's keystone intelligence doctrine, Joint Doctrine Publication 2-00 (fourth edition, 2023), treats evaluation, an appraisal of source reliability and information credibility, as a step in processing that comes before analysis. That order matters for commercial teams too. Grade on collection, then analyse. If you grade after you have formed a view, the view will seep into the grades.

Why Grading Matters for OSINT and CTI Teams Now

The system is decades old, but interest in it has spiked. In September 2024, Freddy Murstad and Sean O'Connor (co-author of the SANS FOR589 course) published a series on applying it to cyber threat intelligence, starting with Enhance Your Cyber Threat Intelligence with the Admiralty System. Their line that B2 and above is the kind of intelligence you want to bet on, but harder to get than you would think, is the most honest summary of OSINT grading we have read. Murstad returned to the theme at the SANS CTI Summit in January 2026 with From Your Gut to a Gold Standard, which covers the practical friction: rating subjectivity, data volume and AI-assisted analysis.

The research base has also moved. A 2025 study in Judgment and Decision Making by Kelly, Budescu, Dhami and Mandel, on how source reliability and information credibility shape judgments of information quality, found that people rating trustworthiness put substantially more weight on the source than on the information. In other words, analysts naturally over-trust good sources and under-use good information from unknown ones. A written grade for each axis is the practical counter to that bias.

For security teams, the payoff is prioritisation: you act on the B2 indicators first and park the F6 chatter for corroboration. Our cyber security teams page shows how that fits alongside Cyber Intelligence and Threat Intelligence searches in UserSearch.

Why Does So Much Open-Source Material Start at F6?

Classic intelligence sources are recruited, vetted and reassessed over years. Open sources are not. Most of what you collect in an OSINT case comes from an account, channel, website or dataset you have never used before. By definition, you have no history of its reporting, so the honest reliability grade is F. And most claims arrive alone, without independent confirmation, so the honest credibility grade is 6 or, at best, 3.

That is uncomfortable for teams who want quick answers, so they cut corners in predictable ways:

  • Grading the platform instead of the author. Telegram is not a source. A specific channel is. Reddit is not a source. A specific account with a specific posting history is.
  • Grading presentation. A polished website with a company logo tells you someone paid for a template.
  • Grading popularity. Follower counts can be bought, and repetition across many accounts often means one origin copied many times.

UK policing has taken a slightly different route. The College of Policing's intelligence report guidance replaced the older 5x5x5 model with 3x5x2: three source evaluations (reliable, untested, not reliable), five intelligence assessments and two handling codes. "Untested" is essentially the police version of F, and the guidance explicitly allows non-human sources, such as CCTV, to be graded. If you work with police partners, map your grades to theirs in the report so nobody has to guess.

Grading by Hand: A Repeatable Workflow

Here is the manual method we teach. It works in a spreadsheet, and it works better when the grades are stored as structured data your tools can read.

Step One: Split the Item into Source and Claim

Write down, in one line each, who the source is (the specific author, account, register or dataset) and what the specific claim is. "Channel @example_handle states that the domain support-example.example.com belongs to Halsworth Bank" is two things: an unknown channel, and a claim about domain ownership. Grade the channel on its history. Grade the claim on what else supports it.

Step Two: Grade the Source on History, Not Feel

Ask: have we used this source before, how often was it right, and does it have first-hand access to what it reports? An official register has direct access to what was filed with it; a news aggregator does not. A first-time source is F, whatever it looks like.

Step Three: Grade the Claim on Corroboration and Plausibility

Ask: is this confirmed by a source independent of the first, is it consistent with what we already know, and is it plausible? Consistent but unconfirmed is usually 3. Consistent and supported by one independent source is usually 2. Confirmed by independent sources with different access is 1.

Step Four: Store the Grade as a Machine Tag

If your team uses MISP, the taxonomy already exists. The PyTaxonomies library lets you load it in Python, check your tags are valid and export them. No API key is needed.

pip install pytaxonomies

python3
>>> from pytaxonomies import Taxonomies
>>> taxonomies = Taxonomies()
>>> adm = taxonomies.get('admiralty-scale')
>>> adm.description
>>> adm.machinetags_expanded()

Taxonomies() loads the published taxonomy set, get('admiralty-scale') selects this one, and machinetags_expanded() lists every valid tag with its human-readable meaning. The tags themselves look like this, and you can paste them straight into an event or a report appendix:

admiralty-scale:source-reliability="f"
admiralty-scale:information-credibility="3"

One friction point: the tag values use the short codes, while the descriptions carry the wording. Agree internally which one appears in written reports so a reader never sees "f" and "Reliability cannot be judged" used for the same thing in different places.

Step Five: Convert to STIX Confidence When You Share

If you share indicators in STIX 2.1, there is no separate Admiralty field. The STIX 2.1 specification instead maps the Admiralty credibility scale onto its 0 to 100 confidence property in an appendix. The OASIS stix2 Python library implements that mapping:

pip install stix2

python3
>>> from stix2.confidence.scales import admiralty_credibility_to_value, value_to_admiralty_credibility
>>> admiralty_credibility_to_value("2 - Probably True")
70
>>> value_to_admiralty_credibility(85)
'1 - Confirmed by other sources'

Note what the mapping loses. Only information credibility travels; source reliability has no slot in confidence, and "6 - Truth cannot be judged" maps to no value at all, so the property is simply left out. If reliability matters to the recipient, and it usually does, put it in a label or a note. Platforms such as OpenCTI solve this by storing reliability on the author entity; its documentation on reliability and confidence shows the default Admiralty vocabulary and how confidence levels restrict which users can overwrite data.

How Does Corroboration Move a Grade?

Corroboration is the engine of the whole system, and independence is what makes it count. Two sources that both copied the same original post are one source. Before you upgrade anything, ask whether the second source had its own access to the facts: a different dataset, a different collection method, a different moment in time.

Keep three rules in mind:

  • Independent confirmation raises the number, not the letter. An F source whose claim you confirm becomes F1 for that item. The source stays F until it builds a history of reporting with you.
  • Letters move slowly and can move down. After several confirmed reports, an F source can reasonably become C. One badly wrong report should pull it back.
  • Circular reporting is the classic failure. A claim appears on a forum, gets repeated by three news aggregators and a dozen accounts, and looks confirmed. Follow every copy back to its first appearance before counting it.

Here is a worked grading table from a supplier check, showing initial grades and grades after corroboration:

ItemSourceInitialCorroboration foundFinal
Company incorporated in 2019, company number 00000000National company registerA2Filing history and accounts consistentA1
Director describes herself as "logistics consultant"National company registerA3None; self-declared occupationA3
Domain carrowmere.example.com created March 2026Domain registration historyB2Earliest web archive copy April 2026B1
Forum post says the company "went bust last year"Unknown forum accountF6Register shows active status and recent filingF5
Telegram channel claims to be the firm's official supportNew channel, 3 weeks oldF6Company website links a different support channelF5
Same username on the firm's GitHub and a freelancer marketplaceTwo platformsF3Same avatar, matching bio link to the domainF2

Two points stand out. Corroboration can push credibility down as well as up: the forum and channel claims became 5 because independent evidence contradicted them. And the username match stayed at F2 rather than F1, because shared handles and avatars are strong evidence, not proof.

Grading Results from Aggregated Data Sources

Aggregated platforms create a grading trap. If a platform returns a result, is the platform the source? No. The platform is the route. The source is the provider that held the data, and your grade should reflect that provider's access and history, plus how fresh the record is.

This is why we built OneScan the way we did. It runs one input across several data sources you select, merges the results and keeps source attribution on every item, so you can see which provider returned what. The Credit cost is shown before you run it. Through one UserSearch account you reach 100+ third-party data sources, and each of them deserves its own reliability grade in your team's reference sheet. A company register returned through Corporate Intelligence might be A. A username existence check through Username Intelligence tells you an account with that handle exists (often A for that narrow fact), but says nothing on its own about who runs it, so the claim "this is the same operator" starts at 6 or 3.

Three habits help:

  • Keep a source reliability register. One row per provider, with a default letter, the reason and the date it was last reviewed. Revisit it quarterly.
  • Grade freshness in the comment. Coverage and freshness depend on the third-party source. A record last refreshed two years ago can be accurate history and still wrong today.
  • Store the route. Run the work inside a Case in Forensic Mode so the search history and bookmarks sit with the grades. A reviewer can then see which search produced which graded item.

Our OSINT fundamentals knowledge base has a short primer on corroboration if you are building a grading sheet for a new team.

Worked Scenario: Grading a Channel That Claims to Be Official Support

Context. A bank's brand protection team is told that a Telegram channel, @example_handle_support, is directing customers to a payments page on secure-pay.example.org. The team needs to know whether the channel is connected to the bank (it is not listed anywhere official) and how confident they can be before requesting takedown.

Actions. The analyst runs the handle through the Chat Messaging (Telegram) search type and the domain through Domain Intelligence, inside a Case. The channel is new to the team: F. Its claim to be official support starts at 6. Domain Intelligence shows the domain was registered nine days earlier through a registrar the bank does not use, with a favicon matching the bank's own site. The bank's official website lists a different support route. Our Telegram channel OSINT guide covers the channel-level checks in more depth.

Outcome. The claim "this channel is official" moves from F6 to F5 (improbable), supported by two independent, higher-graded sources: the bank's own site (A1 for its own support channels) and domain registration history (B2). The team's report states "highly likely the channel is not operated by the bank" and requests takedown with the grades attached. Nobody has to argue about impressions.

Context. An AML analyst at a payments firm reviews an incoming transfer from a wallet the customer says belongs to their trading company, Example Trading Ltd. A crypto forum post from an unknown account claims the same wallet belongs to an unrelated group.

Actions. The analyst runs the address through the Cryptocurrency search type and the company through Corporate Intelligence. The forum post is F6. The customer's declaration is from a known customer with two years of clean history with the firm: C2. The company's public website publishes the wallet as its payment address, and archive copies show it listed there for 14 months. A second, independent data source attributes the address cluster to a regulated exchange account rather than the unrelated group.

Outcome. The customer's claim is upgraded to C1 (confirmed by the company's own long-standing publication and an independent attribution). The forum claim drops to F5. The analyst records both, with dates, and closes the review with a note in the Case. Our anti-money laundering page shows how teams use these searches in reviews.

Advanced Grading Practices for Mature Teams

Calibrate the Team Every Quarter

Give five analysts the same ten items and compare grades. Where they disagree, write down the rule that settles it. The SANS authors put it plainly: Bob from the SOC might rate things differently than Alice from CTI. Calibration turns that from an argument into a house rule.

Never Average Grades Across Items

A finding supported by one A1 item and three F6 items is not "roughly C3". Report the strongest independent support and list the weak items as context. Averaging blurs where the weight really sits.

Keep Grades Separate from Judgments

Grades describe inputs. Judgments describe what you conclude from them. For the judgment, use a likelihood scale such as the UK probability yardstick, published in the government's guidance on explaining uncertainty in intelligence assessment, plus a confidence rating. "Highly likely, moderate confidence, based on items graded A1, B2 and F2" tells the reader everything.

Re-grade When Time Passes

Information ages. A B1 record of a company's directors from 2023 is good history and possibly stale fact. Add a "valid as of" date to every graded item and re-check anything older than your case's decision window.

Grades attach weight to statements about real people and organisations, so treat them with the same care as the data itself. Grade claims, not character: "claim that X controls the wallet: F6" is appropriate, while labelling a person as unreliable in a shared system is a statement about them that may be disclosable and must be accurate under data protection law.

Collect only what the question needs, record the lawful basis for the research, and keep graded data no longer than the case requires. Respect platform terms when you gather material. Where an item grades low, say so plainly and do not let it drive a decision about an individual on its own. And treat every automated or AI-assisted grade as a suggestion that an analyst confirms and signs.

Frequently Asked Questions About the Admiralty Code in OSINT

What is the Admiralty Code in OSINT?

The Admiralty Code, also called the Admiralty or NATO system, grades OSINT material on two independent axes: a letter from A to F for how reliable the source is, and a number from 1 to 6 for how credible the specific information is. The two combine into a grade such as B2 or F6, which records why a finding is believed.

What does F6 mean in the Admiralty Code?

In the Admiralty Code, F6 means the reliability of the source cannot be judged and the truth of the information cannot be judged. F6 is a starting grade, not a verdict. Most open-source material begins at F6 because analysts have no history with the account or website and the claim has not yet been confirmed independently.

What is the difference between source reliability and information credibility?

Source reliability describes the source itself, based on its history of reporting and its access to the facts, and it changes slowly. Information credibility describes one specific claim, based on independent corroboration, consistency and plausibility, and it can change quickly. A reliable source can pass on a doubtful claim, and an unknown source can report something confirmed elsewhere.

How does corroboration change an Admiralty Code grade?

Independent corroboration raises the information credibility number, not the source reliability letter. An F source whose claim is confirmed becomes F1 for that item, while the letter only improves after several confirmed reports. Corroboration can also lower a grade when independent evidence contradicts the claim, and two sources copying the same original post count as one.

How do you grade OSINT results from an aggregated platform?

When OSINT results arrive through an aggregated platform, grade the underlying data source, not the platform used to reach it. The platform is the route; the provider that held the data is the source. Keep a source reliability register with one row per provider, a default letter, the reason and a review date, and note how fresh each record is.

How do you convert an Admiralty grade to STIX confidence?

STIX 2.1 has no separate Admiralty field. Its specification maps the Admiralty credibility scale onto the 0 to 100 confidence property, and the OASIS stix2 Python library implements that mapping, so 2 - Probably True becomes 70. Source reliability has no slot in STIX confidence, so record the reliability letter in a label or a note.

From Gut Feel to Graded Evidence

The Admiralty system is not complicated. Two axes, six levels each, and a discipline of keeping them apart. What makes it powerful is that it forces every finding to carry its own reason for belief, and it makes the path from F6 to something you can act on visible to the whole team.

Stop guessing. Start researching with UserSearch. One account reaches 100+ third-party data sources, OneScan runs your input across the sources you choose and keeps attribution on every result so you can grade each provider properly, Cases in Forensic Mode keep the search history behind every grade, shared team Cases make calibration easy, and SargeBot can draft a PDF report that you verify, grade and sign.

About the author

UserSearch Team
Updated on Sep 30, 2026