Global site search

Search guides, labs, glossary, and research

Type two or more characters to search.

Start with a channel, artifact, or defense term

Examples include zero-width, metadata, tokenizer, or prompt injection.

    Text & Unicode · Identifier confusion

    Unicode Homoglyph & Mixed-Script Analyzer

    A glyph can look like Latin text while carrying Greek or Cyrillic code points. The analyzer reveals scripts, bytes, and a limited educational skeleton without automating domain spoofing.

    Quick answer

    What does the Homoglyphs & Confusables show?

    A glyph can look like Latin text while carrying Greek or Cyrillic code points. The analyzer reveals scripts, bytes, and a limited educational skeleton without automating domain spoofing.

    Human visibility
    Visually similar
    Machine receiver
    Code-point and script analyzer
    Robustness
    Survives ordinary rendering

    Research boundary: this page uses bounded, inert data and fixed safe examples. It never executes decoded content, requests secrets, calls third-party services, or performs actions against external systems.

    FIRST RUN / THREE STEPS

    Start with the prepared, bounded workflow.

    Nothing runs automatically
    1. Review the prepared starter input

      A bounded benign input is already present. Change it only when you are ready to test a different authorized artifact.

    2. Run Analyze look-alikes

      Run the normal first-pass analysis for the prepared values above.

    3. Scan before expanding

      Read the summary first, then scan findings and expand only the machine views you need.

    INPUT / CONTROL PLANE

    Prepare the input and choose one action.

    Laboratory status: Ready

    The recommended first run is separated from alternate analyses. Inputs and selected files stay on this host.

    Current input state Mixed-script code loaded

    These bounded starter values are ready to inspect. Review them before running the recommended action.

    12 / 512 bytes

    Enter one bounded, authorized value for this local analysis.

    Maximum: 512 UTF-8 bytes.

    Switch prepared example3 options

    Loading a sample changes only the form values. Review the result and run an action yourself.

    Prepared benign examples
    ACTION HIERARCHY

    Run the recommended first pass.

    Alternate actions remain available below, but the first pass is the clearest place to start.

    Other analyses1 action

    Inputs remain on this host. Text operations are size-limited; uploaded files are processed from PHP’s temporary upload and are not retained by the application.

    OUTPUT / MACHINE VIEWS

    Scan the result from summary to evidence.

    Run Analyze look-alikes to create the first result.

    The prepared starter input is ready. The output will lead with a summary and visible qualifications before the expandable machine views.

    SummaryFindingsMachine views
    Interpretation framework

    The same artifact can produce several valid observations.

    01

    Human view

    What a person naturally reads, sees, or hears.

    02

    Structural view

    What a parser, DOM, container reader, or metadata extractor exposes.

    03

    Decoder view

    What becomes meaningful only with a rule, key, tokenizer, model, or tool.

    04

    Defensive view

    What normalization, rendering, OCR, canonicalization, or policy changes.

    Evidence and decision boundary

    Use the result as bounded evidence, not as a universal verdict.

    A glyph can look like Latin text while carrying Greek or Cyrillic code points. The analyzer reveals scripts, bytes, and a limited educational skeleton without automating domain spoofing.

    Interpretation rule

    Record the receiver and transformation.

    Machine-decodable is receiver-relative. Record the parser, preprocessing, codebook, tokenizer, key, model, and transformation path before generalizing from one result.

    Limitations

    What this page does not prove

    This page uses a deliberately small local mapping for demonstration. Production checks should use standardized Unicode confusable data and locale-aware script policy.