Steel handcuffs on a black background

Return of the HMAC

Linking Patient Data Without Sharing Identities with HMAC, KMAC and Tokenisation.

Keyed hashing lets a GP system, an acute trust and a community provider join their records on the same patient without any of them, or the analysts downstream, ever seeing each other’s NHS numbers. That is the quiet workhorse behind most pseudonymised linkage in UK health data, and it rests on two small cryptographic functions: HMAC and KMAC.

This post explains what those functions are, why a plain hash of an NHS number is not good enough, and how a sound linkage pipeline is put together. It also covers where the approach breaks and the governance that has to sit around it because GDPR treats pseudonymised data as personal data.

Deterministic Tokenisation Functions

Both are keyed hash functions: put in a secret key and some data, and get back a fixed string of bits. Their original job is message authentication, proving a message came from someone holding the key and was not altered. Three properties make them useful for tokenisation:

  • Deterministic. The same key and the same input always give the same output, so the same patient always gets the same token.
  • Unpredictable without the key. Knowing the algorithm and the input is not enough to produce the output.
  • One-way. You cannot work back from the output to the input.

They are also symmetric: one secret key both creates and checks the output. There is no public/private key pair, so anyone who can verify a token can also generate one. That makes the key the whole of the security.

HMAC (RFC 2104, FIPS 198-1) wraps a conventional hash such as SHA-256. Those hashes suffer from length extension: knowing H(key + message) lets an attacker compute H(key + message + extra) without the key. HMAC defeats this by hashing twice, with the key mixed in at each stage:

HMAC(K,m)=H((K⊕︎opad)‖H((K⊕︎ipad)‖m))\mathrm{HMAC}(K, m) = H\big((K \oplus opad) \,\|\, H((K \oplus ipad) \,\|\, m)\big)

KMAC (NIST SP 800-185, 2016) is built on Keccak, the algorithm inside SHA-3. Keccak’s sponge design is not vulnerable to length extension, so KMAC simply absorbs the padded key and then the message in one pass. It adds two features that matter for health data:

  • Chosen output length. Ask for 128 or 256 bits directly; the length is bound into the result.
  • A customisation string. KMAC(K, x, “project-123”) and KMAC(K, x, “project-456”) give unrelated outputs from the same key, which is exactly what per-project pseudonyms need.
HMACKMAC
Built onAny hash, usually SHA-256Keccak / SHA-3 only
Length-extension defenceNested double hashBuilt into the sponge
Output lengthFixed by the hashChosen by the caller
Domain separationDo it yourselfBuilt-in customisation string
Platform supportEverywhereThinner, growing

Why a plain hash is not enough

A SHA-256 of an NHS number can be reversed in minutes, because there are only about 10 billion possible NHS numbers. An attacker hashes every one on a single GPU, builds a lookup table, and reads the identifiers straight back out of your “anonymised” dataset.

The same applies to most identifiers we care about. Date of birth plus postcode, or a hospital MRN, all come from small, guessable spaces. A salt stored alongside the data does not help either: whoever obtains the data usually obtains the salt.

A keyed hash closes this gap. Without the key, the attacker cannot build the lookup table at all, however much computing power they have. The problem shifts from protecting the data to protecting one key, which is a far more tractable job.

How keyed-hash tokenisation works

Tokenisation replaces each identifier with a keyed hash of it, so the analytics platform only ever sees tokens. If your platform follows a medallion architecture, tokenise before the bronze layer so identifiers never land in it. Done well, it takes three steps.

1. Normalise first. Every provider must clean identifiers the same way before hashing, or linkage silently fails. “943 476 5919” and “9434765919” are different inputs and give unrelated tokens. Agree the rules up front: NHS number as 10 digits with a valid check digit, dates in ISO format, postcodes upper-case with no spaces.

2. Tokenise each identifier, labelled by field. A prefix stops tokens from different fields colliding:

T_nhs = HMAC(K, "nhs|" + nhs_number)
T_dpp = HMAC(K, "dob_pc_sex|" + dob + postcode + sex)
T_name = HMAC(K, "name|" + soundex(surname) + first_initial + dob)

3. Re-tokenise per project. Researchers never receive the linkage-layer identifier. Each project gets its own pseudonym derived from it, so extracts from two projects cannot be joined by anyone outside the trusted environment:

# KMAC: the customisation string separates projects
token = KMAC(K, person_id, "project-123")
# HMAC: derive a per-project key first
K_project = HMAC(K_master, "project-123")
token = HMAC(K_project, person_id)

The resulting tokens are irreversible. If re-identification is genuinely needed, for direct care or recontacting a patient, the trusted party keeps a separate, tightly controlled lookup table. If a system needs to recover the original value routinely, keyed hashing is the wrong tool; use vault tokenisation or format-preserving encryption (FF1) instead.

Linking datasets end to end

A sound linkage pipeline leaves identifiers with the organisations that already hold them and moves only tokens.

Keyed-hash linkage pipeline: key service, providers tokenising locally, linkage environment, per-project re-tokenisation, trusted research environment
Identifiers stay with providers; only tokens reach the linkage layer

Each provider tokenises locally using a key it can call but never read. The linkage environment sees tokens and clinical data, and researchers see only project-specific pseudonyms. That clinical data is mostly coded terminology such as SNOMED CT, which graph databases handle well once records are linked.

There are two ways to run the key:

  • Tokenise at source. Every provider runs the same keyed hash inside its own environment, calling a shared key service. Identifiers never leave the building.
  • Trusted third party. Providers send identifiers over a secure channel to a separate linkage service that tokenises them. The key stays in one place, but that service does see identifiers.

Either way, apply separation of functions: identifiers and clinical data travel by different routes, so no single hop holds both.

The linkage environment then matches in tiers, most reliable first:

  1. Exact match on the NHS number token. This covers the large majority of records.
  2. For records that are left over, match on the date of birth, postcode and sex token.
  3. Then on weaker keys such as phonetic name plus date of birth, flagged as lower confidence.

Each matched cluster gets one person identifier, and match rates are recorded for quality reporting.

The weak spot: exact matches only

A keyed hash changes completely if a single character changes. “Smith” and “Smyth”, a transposed day and month, or last year’s postcode all produce unrelated tokens. Linkage quality therefore depends heavily on NHS number completeness and on how clean each source is.

There are three ways to claw back the missed matches:

  • More keys and phonetic encoding. Tokenise several identifier combinations, with Soundex or similar on names, and match through them in order of reliability.
  • Bloom-filter linkage. Names are split into character pairs, each pair is keyed-hashed into a bit array, and arrays are compared for similarity. This allows fuzzy matching without revealing names. It is more exposed to frequency attacks, so it needs hardening.
  • Link in the clear inside a trusted party. Match on real identifiers within a tightly governed linkage service, then release tokens only. This is often the most accurate route, and it is how national linkage services commonly work.

Governance matters as much as the cryptography

Tokenised data is pseudonymised, not anonymous. Under UK GDPR it remains personal data, because the key holder can re-identify it. As with any sensitive share, no two airlocks are the same: the controls have to fit the data and the people using it. In practice that means:

  • A lawful basis is still needed. For research uses that means consent or Section 251 support, plus a DPIA and data sharing agreements.
  • Linked data is more identifying than any source. Diagnoses, dates and locations combined can single someone out without touching the token. Linked data belongs in a trusted research environment or secure data environment, with output checking, not on a shared drive.
  • The key is the crown jewel. It lives in an HSM or key management service, held by the tokenisation service alone, never in the analytics environment. A leaked key lets anyone brute-force every token.
  • Plan for rotation. Changing the key changes every token and breaks longitudinal linkage. Rotation needs a controlled re-tokenisation run that carries existing person identifiers across.

Standards and libraries

HMAC is defined by an IETF RFC and a NIST FIPS. KMAC is defined in a NIST Special Publication alongside the other SHA-3 derived functions.

StandardDefinesPublished
RFC 2104HMAC: Keyed-Hashing for Message Authentication (Krawczyk, Bellare, Canetti)February 1997
FIPS 198-1The Keyed-Hash Message Authentication Code (HMAC)July 2008
NIST SP 800-224 (draft)HMAC specification and recommendations, set to replace FIPS 198-1June 2024 draft
NIST SP 800-185SHA-3 derived functions: cSHAKE, KMAC, TupleHash, ParallelHashDecember 2016

HMAC is available in every mainstream language. KMAC support is patchier, so check your stack before committing to it.

LibraryLanguageHMACKMAC
hmac modulePython (standard library)Yes, hmac.new / hmac.digestNo
PyCryptodomePythonYesKMAC128, KMAC256 with custom string
OpenSSL 3 EVP_MACC, and anything built on OpenSSLYesKMAC-128, KMAC-256, “custom” parameter
Bouncy CastleJava, C#YesKMAC-128, KMAC-256
System.Security.Cryptography.NETYes, HMACSHA256Kmac128, Kmac256 (.NET 9+), check IsSupported
Node.js cryptoJavaScriptYes, createHmacThrough createMac where the OpenSSL build provides it
crypto/hmacGoYes, hmac.New / hmac.EqualNot in the standard library

In Python, a per-project KMAC token is short with PyCryptodome:

from Crypto.Hash import KMAC128
mac = KMAC128.new(key=key, mac_len=16, custom=b"project-123")
mac.update(person_id)
token = mac.hexdigest()

When comparing tags, use a constant-time check such as Python’s hmac.compare_digest or Go’s hmac.Equal rather than ==.

HMAC or KMAC?

HMAC-SHA256 is the pragmatic default, because every platform supports it: every mainstream language and cryptography library, as the table above shows. KMAC is the cleaner design for multi-project tokenisation, since its customisation string handles domain separation for you. Choose it if your stack supports SHA-3 well.

Either way, the function is the easy part. What makes linkage trustworthy is everything around it:

  1. Normalise identifiers identically everywhere.
  2. Keep the key in one place, away from the analysts.
  3. Separate identifiers from clinical data in transit.
  4. Give every project its own pseudonyms.
  5. Treat the linked result as personal data, because it is.

More posts on data, security and health architecture at apicrazy.com.

Leave a Reply