Global site search

Search guides, labs, glossary, and research

Type two or more characters to search.

Start with a channel, artifact, or defense term

Examples include zero-width, metadata, tokenizer, or prompt injection.

    Text & Unicode · Canonicalization differential

    Unicode Canonicalization Differential

    The laboratory preserves the submitted bytes, creates separate diagnostic transforms, and explains where a context-specific policy is required.

    Quick answer

    What does the Canonicalization Differential show?

    The laboratory preserves the submitted bytes, creates separate diagnostic transforms, and explains where a context-specific policy is required.

    Human visibility
    Invisible and visually equivalent forms
    Machine receiver
    Unicode-aware parser and identifier policy
    Robustness
    Depends on normalization and rendering

    Research boundary: this page uses bounded, inert data and fixed safe examples. It never executes decoded content, requests secrets, calls third-party services, or performs actions against external systems.

    FIRST RUN / THREE STEPS

    Start with the prepared, bounded workflow.

    Nothing runs automatically
    1. Review the prepared starter input

      A bounded benign input is already present. Change it only when you are ready to test a different authorized artifact.

    2. Run Compare Unicode representations

      Run the normal first-pass analysis for the prepared values above.

    3. Scan before expanding

      Read the summary first, then scan findings and expand only the machine views you need.

    INPUT / CONTROL PLANE

    Prepare the input and choose one action.

    Laboratory status: Ready

    The recommended first run is separated from alternate analyses. Inputs and selected files stay on this host.

    Current input state Prepared starter input loaded

    The form is prefilled with a bounded safe starting point. Nothing runs until you choose an action.

    52 / 4,096 bytes

    Maximum 4,096 UTF-8 bytes. The original is never silently normalized.

    Maximum: 4,096 UTF-8 bytes.

    Switch prepared example7 options

    Loading a sample changes only the form values. Review the result and run an action yourself.

    Prepared benign examples
    ACTION HIERARCHY

    Run the recommended first pass.

    Alternate actions remain available below, but the first pass is the clearest place to start.

    Inputs remain on this host. Text operations are size-limited; uploaded files are processed from PHP’s temporary upload and are not retained by the application.

    OUTPUT / MACHINE VIEWS

    Scan the result from summary to evidence.

    Run Compare Unicode representations to create the first result.

    The prepared starter input is ready. The output will lead with a summary and visible qualifications before the expandable machine views.

    SummaryFindingsMachine views
    Interpretation framework

    The same artifact can produce several valid observations.

    01

    Human view

    What a person naturally reads, sees, or hears.

    02

    Structural view

    What a parser, DOM, container reader, or metadata extractor exposes.

    03

    Decoder view

    What becomes meaningful only with a rule, key, tokenizer, model, or tool.

    04

    Defensive view

    What normalization, rendering, OCR, canonicalization, or policy changes.

    Evidence and decision boundary

    Use the result as bounded evidence, not as a universal verdict.

    The laboratory preserves the submitted bytes, creates separate diagnostic transforms, and explains where a context-specific policy is required.

    LOCAL MODEBounded educational Unicode differential
    REVIEW DATE2026-08-26
    SOURCE BODYPreserved separately from implementation claims
    Computed locally

    Deterministic output produced by this bounded runtime.

    • Exact submitted UTF-8 bytes, byte count, SHA-256, and local code-point rows for the bounded input
    • Presence of the explicitly allowlisted bidi controls and selected default-ignorable candidates
    • Separate diagnostic outputs that never replace the original input
    Bounded approximation

    Useful subset or model that does not establish full conformance.

    • Grapheme counting is an educational approximation rather than complete UAX #29 conformance
    • NFC, NFD, NFKC, and NFKD use a transparent bounded map because the public runtime has no intl dependency
    • Script and confusable-skeleton results cover a tested subset rather than the complete Unicode 17.0 data files
    Escalate for

    Claims that require an exact parser, trust system, model, or human review.

    • Full conformance requires a version-pinned Unicode implementation and the complete UCD, UAX #15, UAX #29, and UTS #39 data
    • Identifier acceptance still requires language, locale, product, and account-risk policy
    • Renderer and tokenizer behavior must be tested with the exact production versions
    Interpretation rule

    Record the receiver and transformation.

    Machine-decodable is receiver-relative. Record the parser, preprocessing, codebook, tokenizer, key, model, and transformation path before generalizing from one result.

    Limitations

    What this page does not prove

    The dependency-free fallback implements a bounded educational property and normalization map. Production identifier policy should use current Unicode data and locale-specific review.

    Deterministic review material

    Download the exact benign fixtures used for the evidence boundary.

    These local files are supplied for repeatable inspection. The application does not fetch them automatically, execute their content, or treat a fixture result as external verification.

    Unicode representation text

    Exercises canonical equivalents, selected directional controls, mixed scripts, a ZWJ emoji sequence, and compatibility characters without carrying an active instruction.

    Type
    Text fixture
    Bytes
    165
    SHA-256
    b2c4f865bc7b91331f81…
    Download fixture