Skip to content

Transforms ​

Transforms normalise a plaintext value before it is passed to the HMAC function. Normalisation ensures that logically equivalent inputs - differing only in case or whitespace - produce the same blind index fingerprint.

Without transforms, searching for "jane@example.com" would not match a record saved with "Jane@Example.com", even though they refer to the same address.

Configuring Transforms ​

Transforms are declared on the [BlindIndex] attribute as an ordered array of strings:

cs
[PersonalData]
[BlindIndex(
    IndexPropertyName = nameof(EmailHash),
    Transforms = ["lowercase", "trim"])]
public string Email { get; set; } = "";
anchor

Transforms are applied left to right. The example above first converts to lowercase, then strips surrounding whitespace.

The BlindIndexTransforms class has a constant for each built-in name (BlindIndexTransforms.Lowercase, BlindIndexTransforms.Fold, and so on) if you prefer them to string literals.

Transforms are part of the stored index

The index stores a hash of the transformed value, so the transform list is part of the persisted format. Adding, removing, or reordering a transform on a field that already has data makes every existing index value stop matching, with no error. Treat a transform change like an HMAC key rotation and run a recompute over the existing data. The same applies to a custom transform whose logic changes: version its name (my_fold_v2) rather than editing it in place.

Built-In Transforms ​

lowercase ​

Converts the entire value to lowercase using the invariant culture.

InputOutput
"Jane@Example.COM""jane@example.com"
"ACME Corp""acme corp"
"already lower""already lower"

Use on: email addresses, usernames, domain names.


trim ​

Removes leading and trailing whitespace (spaces, tabs, newlines).

InputOutput
" jane@example.com ""jane@example.com"
"\tjane\n""jane"
"no change""no change"

Use on: any field where trailing spaces might appear from user input or data imports.


alphanumeric ​

Removes all characters that are not ASCII letters (a-z, A-Z) or digits (0-9). Useful for normalising names or identifiers that might contain punctuation.

InputOutput
"O'Brien""OBrien"
"Smith-Jones""SmithJones"
"+1 (555) 867-5309""15558675309"
"Müller""Mller"

Accented letters are removed, not kept

alphanumeric is ASCII-only, so a letter such as ü, ø, or ß is deleted outright: "Müller" becomes "Mller" and no longer matches "Muller". For names, put fold first: ["lowercase", "fold", "alphanumeric"] turns "Müller-Łódź" into "mullerlodz".

Combine with lowercase for case-insensitive matching

alphanumeric alone does not change case. Use ["lowercase", "alphanumeric"] if you want case-insensitive matching.


digits ​

Retains only ASCII digit characters (0-9). All other characters are removed. Designed for phone numbers, tax IDs, and other numeric identifiers.

InputOutput
"+1 (555) 867-5309""15558675309"
"SSN: 123-45-6789""123456789"
"GB VAT 123 456 789""123456789"

last4 ​

Retains only the last 4 characters of the value after all other characters have been processed. Commonly used for partial credit card or SSN matching.

InputOutput
"4111111111111111""1111"
"123-45-6789""6789"
"AB12""AB12"
"AB""AB" (shorter than 4 - returned as-is)

Combine last4 with digits for card numbers

Use ["digits", "last4"] to strip formatting characters before taking the last four digits. This ensures "4111-1111-1111-1111" and "4111111111111111" produce the same result.


first_char ​

Retains only the first character of the value. Useful for bucketed or initial-based lookups.

InputOutput
"Jane""J"
"jane""j"
"""" (empty string is preserved)

Low cardinality warning

first_char produces at most 26 distinct values (plus digits and symbols). This is a very low-cardinality blind index and is susceptible to frequency analysis. See Security Considerations.


fold ​

Folds diacritics and letter variants to their base letters, for accent-insensitive matching on names and addresses.

InputOutput
"Müller""Muller"
"Søren""Soren"
"Straße""Strasse"
"Łukasz Wałęsa""Lukasz Walesa"
"IJsbrand""IJsbrand"
"Ærø""AEro"
cs
public class FoldedSurnamePatient
{
    [DataSubjectId]
    public string PatientId { get; set; } = "";

    // "Müller", "MULLER" and " muller " all produce the same index.
    [PersonalData]
    [BlindIndex(Transforms = [BlindIndexTransforms.Trim, BlindIndexTransforms.Lowercase, BlindIndexTransforms.Fold])]
    public string Surname { get; set; } = "";
    public string? SurnameIndex { get; set; }
}
anchor

What it covers:

  • Diacritics in the Latin, Greek, and Cyrillic blocks, using Unicode compatibility decomposition with the marks removed. Compatibility (NFKD) rather than canonical (NFD) decomposition is what splits the Dutch ij into ij, and the typographic ligatures fi and fl into fi and fl.
  • Letters with no decomposition, folded explicitly: ø to o, ß to ss, æ to ae, œ to oe, ł to l, đ and ð to d, þ to th, ħ to h, ı to i, and their capitals.
  • Precomposed and decomposed input alike: ü typed as one character and as u plus a combining diaeresis fold to the same value.

What it does not do:

  • Change case. Combine it with lowercase for case-insensitive matching.
  • Transliterate. Folding goes to the base letter, so ü becomes u, never the German ue. If Müller must also match Mueller, add a named custom transform before fold.
  • Touch other scripts. CJK, Hangul, Hebrew, Arabic, and similar text passes through unchanged.

fold works from a fixed table checked into Tayra rather than calling string.Normalize, so it produces the same output on every OS, ICU version, and .NET version, including containers that run in globalization-invariant mode, where string.Normalize leaves non-ASCII text unchanged. Existing mappings in the table never change between Tayra versions.


Transform Ordering ​

Transforms are applied in the order they are declared. Order matters.

Example: ["trim", "lowercase", "digits"]

Input:  "  +1 (555) 867-5309  "
  trim →  "+1 (555) 867-5309"
  lowercase → "+1 (555) 867-5309"  (no letters, no change)
  digits → "15558675309"

Example: ["digits", "last4"]

Input:  "4111-1111-1111-1111"
  digits → "4111111111111111"
  last4 → "1111"

Reversing the order would give last4 the formatted string first, which could produce a different result depending on the trailing characters.

Custom Transforms ​

There are two ways to add your own transform: register a named transform that attributes can reference, or pass an inline function to the fluent API.

Named custom transforms ​

Implement IBlindIndexTransform and register it with opts.BlindIndex.RegisterTransform(...). Any [BlindIndex], [ArrayBlindIndex], or [CompoundBlindIndex] attribute can then use its Name in Transforms, alongside the built-in names:

cs
/// <summary>
/// Expands German umlauts to their two-letter spelling (DIN 5007-2), so that "Müller" and
/// "Mueller" match. Expects lowercase input; register it before "fold" in the pipeline.
/// </summary>
public sealed class GermanUmlautTransform : IBlindIndexTransform
{
    public string Name => "german_umlauts";

    public string Apply(string input) => input
        .Replace("ä", "ae")
        .Replace("ö", "oe")
        .Replace("ü", "ue")
        .Replace("ß", "ss");
}

public class GermanSurnamePatient
{
    [DataSubjectId]
    public string PatientId { get; set; } = "";

    // "Müller" and "Mueller" match; "Muller" does not.
    [PersonalData]
    [BlindIndex(Transforms = ["trim", "lowercase", "german_umlauts", "fold"])]
    public string Surname { get; set; } = "";
    public string? SurnameIndex { get; set; }
}
anchor
cs
// Register a named transform once; any [BlindIndex], [ArrayBlindIndex] or
// [CompoundBlindIndex] attribute can then reference it by name.
using var umlautTayra = TayraHost.Create(opts =>
{
    opts.LicenseKey = licenseKey;
    opts.BlindIndex.RegisterTransform(new GermanUmlautTransform());
});

var germanPatient = new GermanSurnamePatient { PatientId = "p-2", Surname = "Müller" };
await umlautTayra.EncryptAsync(germanPatient);

var umlautSearchHash = await umlautTayra.ComputeBlindIndexAsync("Mueller", "SurnameIndex", typeof(GermanSurnamePatient));
Console.WriteLine($"  Mueller matches Müller: {umlautSearchHash == germanPatient.SurnameIndex}");
anchor

RegisterTransform works the same way with services.AddTayra(opts => ...). An attribute that names a transform nobody registered fails when Tayra first builds that type's metadata, with an error naming the missing transform. Registering a transform under a built-in name replaces the built-in.

Inline transforms (fluent API) ​

Use WithTransform() to add an inline function in the fluent API:

cs
// Inline custom transforms - no class or registration needed
var transformServices = new ServiceCollection();
var transformBuilder = transformServices.AddTayra(opts => opts.LicenseKey = licenseKey);
transformBuilder.Entity<IndexedCustomer>(e =>
{
    e.DataSubjectId(c => c.CustomerId);
    e.PersonalData(c => c.Email);
    e.BlindIndex(c => c.Email)
        .WithTransform(value => value.Split('@')[0]) // extract local part
        .WithLowercase()
        .StoredIn(c => c.EmailIndex);
});
anchor

An inline transform needs no class or registration, and composes with the built-in transforms in the pipeline. The fluent API also has a With...() method for each built-in, including WithFold().

Custom Transform Rules ​

  • The function must be a pure function - same input always produces the same output, on every machine that writes or queries the index. Avoid culture-sensitive APIs such as ToLower() without Invariant, and string.Normalize, which does nothing to non-ASCII text in globalization-invariant mode.
  • The function must not throw on an empty string.
  • Transforms should be fast (no I/O, no allocations if avoidable).

Transform Reference Summary ​

NameEffectTypical Use
lowercaseConverts to invariant lowercaseEmail, username
trimRemoves leading/trailing whitespaceAny user-input field
alphanumericKeeps only [a-zA-Z0-9]; removes accented lettersIdentifiers; names after fold
digitsKeeps only [0-9]Phone numbers, tax IDs
last4Keeps last 4 charactersCard numbers, SSN suffix
first_charKeeps first character onlyBucketed lookups
foldFolds diacritics and letter variants to base lettersNames, addresses

See Also ​