DataMasque Portal

Deterministic Masking

What is Deterministic Masking?

Deterministic Masking allows the same entity to receive the same mask, both within and across documents.

For example:

Patient ID Original Text Masked Text
PATIENT_001 Tarzan is sleeping Jefferie is sleeping
PATIENT_001 Tarzan is happy. Jefferie is happy.
PATIENT_002 Jane is recovering Adin is recovering

All names within the same patient_id receive the same masked value.

This can be useful for:

  • preserving referential integrity
  • preserving consistency in datasets for training downstream AI/ML models

Deterministic Masking is not unique to unstructured masking, for more see Deterministic masking

Which Hash Source to Use?

For example, when masking PATIENT_FIRST_NAME, you might want consistency based on different criteria:

Use Case Recommended Approach Example
Same patient ID, same mask hash_columns with
patient_id
patient_id=001:
"John" → "Alice"
"Jane" → "Alice"
patient_id=002:
"John" → "Bob"
Same name, same mask hash_sources with
self: entity
"John" → "Alice" (everywhere)
"Jane" → "Carol" (everywhere)
Same doctor, same mask hash_sources with
label: "DOCTOR"
Text with Dr. "Smith":
"John" → "Alice"
"Jane" → "Alice"
Text with Dr. "Jones":
"John" → "Bob"
Same patient ID and name, same mask hash_columns with
patient_id
+ hash_sources with
self: entity
patient_id=001:
"John" → "Alice"
"Jane" → "Carol"
patient_id=002:
"John" → "Bob"

hash_columns

Use hash_columns to hash on column values from the current row. This works the same as deterministic masking for structured data.

hash_columns can be specified at both the task-level (applying to all columns) and the column-level (overriding the task-level default).

- column: raw_text
  hash_columns:
    - patient_id
  masks:
    - type: unstructured_text
      matchers:
        ai_detect:
          - label: "FIRST_NAME"
            use_preset: "FIRST_NAME"
patient_id raw_text masked_text
PATIENT_001 Tarzan sleeps well. Jefferie sleeps well.
PATIENT_001 Tarzan greets Jane Jefferie greets Jefferie
PATIENT_002 Tarzan is recovering Adin is recovering
PATIENT_002 Jane visited today Adin visited today

All FIRST_NAME entities with the same patient_id will receive the same masked value.

hash_sources

Use hash_sources to hash on entities detected within unstructured text. For the full specification of hash_sources parameters, see Ruleset YAML Specification.

self: entity

Use self: entity to hash on the detected entity itself.

- column: raw_text
  masks:
    - type: unstructured_text
      hash_sources:
        - self: entity
      matchers:
        ai_detect:
          - label: "FIRST_NAME"
            use_preset: "FIRST_NAME"
patient_id raw_text masked_text
PATIENT_001 Tarzan sleeps well. Bryor sleeps well.
PATIENT_001 Tarzan greets Jane Bryor greets Maeola
PATIENT_002 Tarzan is recovering Bryor is recovering
PATIENT_002 Jane visited today Maeola visited today

All FIRST_NAME entities with the same text will receive the same masked value. For example, "Tarzan" is always masked to "Bryor", and "Jane" is always masked to "Maeola".

Label-based hashing

Use label: "LABEL" to hash on specific entities within the text.

- column: raw_text
  masks:
    - type: unstructured_text
      hash_sources:
        - label: "PATIENT_ID"
      matchers:
        regex:
          - label: "PATIENT_ID"
            pattern: 'PATIENT_\d{3}'
        ai_detect:
          - label: "FIRST_NAME"
            use_preset: "FIRST_NAME"
raw_text masked_text
PATIENT_001: Tarzan sleeps well. Tarzan eats well. PATIENT_001: Bryor sleeps well. Bryor eats well.
PATIENT_001: Tarzan is recovering. PATIENT_001: Bryor is recovering.
PATIENT_002: Tarzan greets Jane. PATIENT_002: Khalik greets Khalik.

All FIRST_NAME entities will receive the same masked value based on the first PATIENT_ID found.

match and match_until

The optional match parameter can be used to index on the label to hash on, which can be useful in cases where there are multiple matches for the same label.

Extending this with the optional match_until parameter will allow you to define a range of labels to hash on.

hash_sources:
  - label: "PATIENT_ID"
    match: 0
    match_until: 1
Parameter(s) Hashes on
match: 0 First entity (default if not specified)
match: 1 Second entity
match: -1 Last entity
match: 0, match_until: 1 First two entities (concatenated)
match: -2, match_until: -1 Last two entities (concatenated)

Hashing on masked results

Adding state: masked to a label: hash source allows for hashing on a label's masked text.

This can be added to both mask-level and per-label hash_sources.

The following example masks ACCOUNT_NUMBER, using a hash derived from a masked CUSTOMER_ID.

masks:
  - label: "CUSTOMER_ID"
    masks:
      - type: imitate
  - label: "ACCOUNT_NUMBER"
    masks:
      - type: imitate
    hash_sources:
      - label: "CUSTOMER_ID"
        state: masked

All ACCOUNT_NUMBER entities which appear alongside the same CUSTOMER_ID, will receive the same masked value, derived from CUSTOMER_IDs masked text.

When set at the mask level, state: masked applies to every label without a per-label override.

The following example masks both ACCOUNT_NUMBER and INVOICE_NUMBER, using a hash derived from a CUSTOMER_ID via a mask level hash_sources.

hash_sources:
  - label: "CUSTOMER_ID"
    state: masked
masks:
  - label: "CUSTOMER_ID"
    masks:
      - type: imitate
    hash_sources:
      - self: entity
  - label: "ACCOUNT_NUMBER"
    masks:
      - type: imitate
  - label: "INVOICE_NUMBER"
    masks:
      - type: imitate

Note: To prevent a circular self-reference (masking CUSTOMER_ID using a hash derived from its own masked value) self: entity is used as a hash source for CUSTOMER_ID. This allows for masking CUSTOMER_ID with a hash derived from its pre-masked value.

ACCOUNT_NUMBER and INVOICE_NUMBER have no per-label hash_sources, so they inherit the mask-level hash_sources and hash on CUSTOMER_ID's masked text.

Behaviour:

  • match and match_until will select from the masked list in document order.
  • state: masked does not allow for circular references or chains longer than one hop.
    • A label referenced with state: masked can't itself use state: masked.

case_transform

Use the optional case_transform parameter to normalise the case of the entity text before hashing. Accepts upper or lower. Useful when entities appear with inconsistent capitalisation, so that case variants hash to the same masked value.

case_transform is available on both self: entity and label: forms.

hash_sources:
- self: entity
  case_transform: upper
raw_text masked_text
Tarzan sleeps well. Bryor sleeps well.
tarzan is recovering. Bryor is recovering.
TARZAN greets jane. Bryor greets Maeola.

Note: case_transform only normalises the input to the hash. It does not affect the case of the masked output — the output reflects whatever the underlying mask (e.g. from_file) produces.

Per-label hash_sources overrides

By default, hash_sources defined at the mask level apply to every detected entity. To configure hashing differently for individual labels — or to opt a single label out of entity-level hashing entirely — set hash_sources on the per-label mask entry.

A per-label hash_sources replaces the mask-level hash_sources for that label (no merging — same as how column-level hash_columns overrides task-level hash_columns).

Per-label config Behaviour
hash_sources: [...] Replaces the mask-level hash_sources for this label
hash_sources: [] No entity-level hashing for this label
hash_sources omitted Inherits the mask-level hash_sources

Rule-level hash_columns continues to combine in regardless of any per-label configuration — per-label hash_sources only controls entity-level sources.

- column: raw_text
  masks:
    - type: unstructured_text
      hash_sources:                 # mask-level default
        - self: entity
      matchers:
        regex:
          - label: "PATIENT_ID"
            pattern: 'PATIENT_\d{3}'
        ai_detect:
          - label: "FIRST_NAME"
            annotation_config:
              use_preset: "FIRST_NAME"
          - label: "DIAGNOSIS"
            annotation_config:
              use_preset: "DIAGNOSIS"
      masks:
        - label: "FIRST_NAME"
          hash_sources:              # overrides for FIRST_NAME only
            - label: "PATIENT_ID"
              match: 0
          masks:
            - type: from_file
              seed_file: name_seed
              seed_column: name
        - label: "DIAGNOSIS"
          hash_sources: []           # explicit opt-out — non-deterministic
          masks:
            - type: from_fixed
              value: "[DIAGNOSIS]"
        - label: "PATIENT_ID"        # no override — uses `self: entity`
          masks:
            - type: from_format_string
              value_format: "PATIENT_$$$"

A hash_sources entry that references a label: which isn't defined by any matcher is rejected when the masking task runs. This applies to both mask-level and per-label hash_sources.

Combining hashes

Chaining together multiple hash_columns and hash_sources will create a composite hash.

Important: The order will affect the final result.

- column: raw_text
  hash_columns:
    - patient_id
    - employee_id
  masks:
    - type: unstructured_text
      hash_sources:
        - label: "HOSPITAL_ID"
        - self: entity
      matchers:
        regex:
          - label: "HOSPITAL_ID"
            pattern: 'H_\d{3}'
        ai_detect:
          - label: "FIRST_NAME"
            use_preset: "FIRST_NAME"

In the example above, FIRST_NAME will be hashed (in order) on:

Hash Input Example
The patient_id P_001
The employee_id E_001
The first HOSPITAL_ID in the document H_001
The discovered entity text itself Tarzan

Limitations

Hash sources scope:

  • Hash sources only reference entities within the current document. You cannot hash on entities detected within other unstructured documents.
  • For a workaround and assuming the entities are easily extractable, consider extracting these entities as metadata into a separate column and then using hash_columns.

No JSON/XML path support:

  • Unlike file masking, hash_sources for unstructured text does not support json_path, xpath, or file_path parameters.
  • Workarounds: