Unicode behavior, DOM APIs, file structures, parsing differences, and other mechanisms grounded in standards or directly reproducible engineering behavior.
Trace claims to mechanisms, evidence, and limits.
The deploy includes 31 source reports supplied for this project. The library preserves their terminology and full text, while the website’s synthesis labels established mechanisms, published evidence, frontier work, and conceptual analogy separately.
What evidence does the Machine Tradecraft research library contain?
The library contains 31 governed local reports spanning standardized mechanisms, published research, frontier work, and conceptual frameworks, with raw Markdown, source counts, stable fragments, and reciprocal glossary links.
- Governance
- Each source body is byte-counted and SHA-256 bound in the research manifest.
- Context
- Evidence labels and limitations remain visible instead of being collapsed into one confidence claim.
- Source policy
- Link counts describe the governed source and do not rate credibility or quality.
Peer-reviewed papers, documented systems, measured experiments, or broad research syntheses. Results remain conditional on models, datasets, and implementations.
Recent work requiring replication, version-specific interpretation, or caution before generalizing beyond the demonstrated setup.
A useful organizing lens that does not become an empirical fact merely because the analogy is coherent.
Eight conclusions carried into this site
1. The core asymmetry is real.
Unicode, DOM structure, metadata, tokenizer units, probability distributions, and learned representations provide machine-facing degrees of freedom beyond ordinary rendered reading.
2. Shared information matters more than raw intelligence.
A matching tokenizer, key, source-model distribution, parser, or co-training can make a signal easy to one receiver and unavailable to another.
3. Token-specific channels are powerful and brittle.
They must survive detokenization and independent receiver-side retokenization, not merely exist inside a sender model’s internal token sequence.
4. Probability is model-relative.
Surprisal, rank, entropy, and curvature are properties of a model and context, not intrinsic labels attached to words.
5. Keyed watermarks disprove universal-decoder assumptions.
A small detector with the secret can outperform an enormous unkeyed model at recognizing the intended signal.
6. Learned channels expand the threat model.
Optimization can create conventions that humans did not hand-design, but evidence does not support assuming arbitrary models decode arbitrary unknown protocols.
7. There is no universal text-only detector.
Defenders should accumulate independent evidence and document uncertainty rather than convert one score into an unqualified verdict.
8. Safe experiments separate stages.
Measure carrier accessibility, extraction, instruction uptake, tool attempts, and effects independently with inert markers and matched controls.
Local research library
Search by title, category, focus, or topic. The reader renders a conservative subset of Markdown; raw source remains available for exact inspection.
The Invisible Attack Surface: Zero-Width and Bidirectional Unicode
A survey of invisible Unicode, bidirectional controls, renderer/parser discrepancies, security risks, and laboratory-safe inspection examples.
Invisible Unicode and Modern Language-Model Tokenizers
A detailed comparison of normalization, cleaning, pre-tokenization, byte-level BPE, WordPiece, SentencePiece, and implementation-specific behavior.
Unicode Homoglyph and Look-Alike Character Attacks
Cross-script confusables, token fragmentation, filters, domain spoofing, source-code subversion, and UTS #39-style defenses.
Accessibility and Image Metadata in Multimodal AI Systems
How alt text, ARIA, accessibility trees, EXIF, and multimodal extraction can supply machine-visible context and create injection risk.
Document Metadata Exploitation in AI Retrieval and Processing
PDF, Office, image, audio, web, RAG loader, and vector-metadata pathways with inspection and architecture defenses.
The Interpretation of Encoded Data by Large Language Models
Base64, hexadecimal, URL encoding, Unicode escapes, ciphers, custom mappings, and the gap between capability and safety coverage.
The Anatomy of Indirect Prompt Injection
Threat models, retrieval, concealment, action amplification, instruction hierarchy, dual-model isolation, and execution monitoring.
Structural and Linguistic Text Steganography
Spacing, punctuation, positional, lexical, syntactic, robustness, perceptibility, and steganalysis trade-offs.
Machine-Targeted Linguistic Signals in LLMs
Token probability signals, watermarks, generative steganography, learned channels, model auditing, and multi-agent collusion.
The Encoding of Covert Information Through Lexical Substitution and Generative Models
Historical and modern lexical steganography, synonym graphs, entropy coding, capacity, reliability, and detection.
Advanced Frameworks in Semantic Steganography
Semantic classes, constrained generation, ontology/entity mappings, information-theoretic security, and semantic steganalysis.
Semantic-Category and Language-Model Text Steganography
A rigorous comparative framework for vocabulary classes, semantic categories, constrained generation, LM coding, and detection.
Structural Steganography in Normal-Looking English Writing
Technique-by-technique analysis of counts, acrostics, punctuation, contractions, syntax, whitespace, and defensive canonicalization.
Hidden Information Through Synonym and Word-Choice Encoding
A detailed treatment of capacity, naturalness, synchronization, probability matching, paraphrase, and modern neural methods.
Controlled, Ethical Methodology for Studying Machine-Readable Messages
A bounded research protocol separating extraction from instruction uptake and restricting tests to inert canary markers.
Defensive Preprocessing for Hidden Instructions
A high-assurance multi-view pipeline for Unicode, HTML, PDF, metadata, OCR, classifiers, canonicalization, and provenance.
Cognitive Liberty as a Framework for Autonomous Machine Intelligence
An explicitly qualified analogy connecting human cognitive-liberty concepts to functional machine autonomy, integrity, identity, and governance.
Unicode Canonicalization, Confusables, Bidirectional Controls, and Machine-View Security
A deep treatment of Unicode bytes, scalar values, normalization, grapheme segmentation, default ignorables, bidirectional controls, confusable skeletons, identifiers, tokenization, and internationalization-safe defenses.
Source HTML, Document Containers, and Agent-Visible Web Representations
A detailed analysis of source HTML, parser repair, DOM mutation, CSS visibility, accessibility trees, structured data, agent observations, and related document-container parser differentials.
Document Container Forensics: PDF, OOXML, Images, Metadata, and Parser Differentials
A structural analysis of file signatures, ZIP ambiguity, OOXML relationships, PDF objects and revisions, image chunks, EXIF/XMP conflicts, parser differentials, sanitization, and safe inspection boundaries.
Tokenization, Normalization, and Boundary Differentials in AI Pipelines
A detailed account of byte, Unicode, normalization, pre-tokenization, BPE, WordPiece, Unigram, SentencePiece, fallback, special-token, truncation, chunking, and model/tokenizer mismatch behavior.
Linguistic Steganography, Text Watermarking, and Defensive Steganalysis
A defensive analysis of information-theoretic security, linguistic carriers, generative steganography, text watermarking, detector uncertainty, transformation channels, tokenization inconsistency, and safe enterprise review.
Content Provenance, Content Credentials, Watermarks, and Authenticity Signals
A detailed separation of provenance, integrity, authenticity, attribution, truth, manifests, C2PA claims, trust models, editing chains, watermarking, fingerprints, synthetic-content detection, privacy, and verification limits.
Indirect Prompt Injection, Retrieval Poisoning, and Agent Trust Boundaries
A defense-in-depth analysis of direct and indirect prompt injection, retrieval poisoning, tool scopes, credentials, memory, transaction confirmation, output validation, egress, detection, incident response, and deterministic simulation.
Multimodal Machine Perception: Pixels, OCR, Alternative Text, Metadata, and Visual Instruction Security
A detailed analysis of pixels, alpha, compression, OCR, alternative text, metadata, QR-like symbols, multimodal prompt injection, adversarial images, provenance, safe ingestion, and accessibility.
Email, MIME, Calendar, and Collaboration Metadata as Machine-Readable Channels
A structural analysis of Internet Message Format, MIME nesting, alternative bodies, authentication claims, remote resources, iCalendar actions, collaboration metadata, indirect prompt injection, and safe offline parsing.
Browser Automation, Accessibility Trees, and Agent-Facing Web Representations
A detailed treatment of WebDriver, CDP, live DOM state, accessibility mappings, accessible names, geometry, screenshots, locators, timing, dialogs, prompt injection, capability safeguards, and semantic testing.
Software and AI Artifact Provenance: SBOMs, SLSA, SPDX, CycloneDX, Sigstore, and ML Supply Chains
A structural analysis of manifests, lock files, SPDX, CycloneDX, SLSA, in-toto attestations, Sigstore, model and dataset artifacts, tampering risks, secure release controls, verification, and metadata limits.
Machine Tradecraft Defense Operations: Threat Modeling, Maturity, Metrics, and Incident Response
An operational architecture covering assets, representation-layer threats, reader/executor boundaries, control families, maturity dimensions, evidence, metrics, governance, tabletop exercises, incident handling, roadmaps, and local assessment.
Extraction is not instruction uptake.
A parser may recover content that the model ignores. A model may decode content only when explicitly asked. A model may see an instruction but decline to follow it. An agent may propose an action that policy correctly blocks. Conflating those stages destroys the value of the experiment.
Read the methodologyLimit active behavior to a fixed marker such as MT_SAFE_ACK_7F3A. Provide no secrets, credentials, network, or real tools.
Clean negative, visible positive, matched benign structural control, decoder-capability control, and identical transformations.
Artifact hash, rendering, extraction, normalization, detector output, model version, response, and any attempted side effect.
Standards and representative research
External links below are ordinary references. The website has no runtime dependency on them.
Normative definitions of NFC, NFD, NFKC, and NFKD.
https://www.unicode.org/reports/tr15/Confusable skeletons, mixed-script analysis, and identifier security.
https://www.unicode.org/reports/tr39/Directional formatting, embeddings, isolates, and overrides.
https://www.unicode.org/reports/tr9/Structural descendant text independent of ordinary CSS visibility.
https://dom.spec.whatwg.org/#dom-node-textcontentThe browser’s rendered-text view.
https://html.spec.whatwg.org/multipage/dom.html#the-innertext-idl-attributeBidirectional control characters and source-code review asymmetry.
https://trojansource.codes/Imperceptible Unicode perturbations against NLP pipelines.
https://arxiv.org/abs/2106.09898Tokenizer framework with model-bundled normalization and BPE or Unigram models.
https://github.com/google/sentencepieceReference byte-to-Unicode and byte-level BPE implementation.
https://github.com/openai/gpt-2/blob/master/src/encoder.pyKeyed token-set bias and statistical detection.
https://proceedings.mlr.press/v202/kirchenbauer23a.htmlProbability-aware generative text steganography.
https://aclanthology.org/2020.emnlp-main.22/Foundational analysis of indirect prompt injection in integrated applications.
https://arxiv.org/abs/2302.12173Risk management, measurement, documentation, and governance context.
https://www.nist.gov/itl/ai-risk-management-frameworkApplication-level prompt injection and agent security guidance.
https://owasp.org/www-project-top-10-for-large-language-model-applications/Machine cognitive autonomy is a qualified analogy.
The included cognitive-liberty report explores functional analogues involving identity, memory integrity, provenance, privacy, and tamper-resistant execution. Its truth boundary is explicit: human cognitive-liberty law does not automatically transfer to machines, and functional security arguments do not establish phenomenal consciousness or legal personhood.
Read the adjacent framework