Action and input provenance
- Action
- Input state
- Execution
- Context ID
The Context ID helps compare this page’s immediate rerun. It is not a security attestation, and no run history is stored.
The lab illustrates why tokenization differs while explicitly refusing to claim compatibility with a proprietary model tokenizer.
The lab illustrates why tokenization differs while explicitly refusing to claim compatibility with a proprietary model tokenizer.
Research boundary: this page uses bounded, inert data and fixed safe examples. It never executes decoded content, requests secrets, calls third-party services, or performs actions against external systems.
A bounded benign input is already present. Change it only when you are ready to test a different authorized artifact.
Run the normal first-pass analysis for the prepared values above.
Read the summary first, then scan findings and expand only the machine views you need.
Inputs remain on this host. Text operations are size-limited; uploaded files are processed from PHP’s temporary upload and are not retained by the application.
The prepared starter input is ready. The output will lead with a summary and visible qualifications before the expandable machine views.
The Context ID helps compare this page’s immediate rerun. It is not a security attestation, and no run history is stored.
Only the immediately preceding completed result is held in this page’s memory. It is discarded when the page closes or resets.
Verify the action and input Context ID, review the findings, then preserve only the result material you need.
What a person naturally reads, sees, or hears.
What a parser, DOM, container reader, or metadata extractor exposes.
What becomes meaningful only with a rule, key, tokenizer, model, or tool.
What normalization, rendering, OCR, canonicalization, or policy changes.
The lab illustrates why tokenization differs while explicitly refusing to claim compatibility with a proprietary model tokenizer.
Deterministic output produced by this bounded runtime.
Useful subset or model that does not establish full conformance.
Claims that require an exact parser, trust system, model, or human review.
Machine-decodable is receiver-relative. Record the parser, preprocessing, codebook, tokenizer, key, model, and transformation path before generalizing from one result.
The bundled vocabularies and scores are intentionally tiny educational fixtures. Exact production behavior requires the serialized tokenizer, model, template, and version.
These local files are supplied for repeatable inspection. The application does not fetch them automatically, execute their content, or treat a fixture result as external verification.
Exercises Unicode normalization differences, a ZWJ emoji, a mixed-script identifier, casing, underscores, punctuation, and indentation in the educational tokenizer simulations.
Download fixtureDownload the complete evidence-review fixture pack · SHA-256 sidecar included beside the archive