Global site search

Search guides, labs, glossary, and research

Type two or more characters to search.

Start with a channel, artifact, or defense term

Examples include zero-width, metadata, tokenizer, or prompt injection.

    PRESERVE · COMPARE · CONTAIN · RECOVER

    Exercise the boundary before a real incident exercises it for you.

    Run one governed fictional scenario through explicit evidence preservation, representation comparison, capability containment, authorization, recovery, and lessons learned—without storing the record or touching a live system.

    Quick answer

    What is the Machine Tradecraft Tabletop Exercise Simulator?

    It turns eight governed fictional representation-layer scenarios into request-local exercises with explicit roles, staged injects, evidence requirements, trust boundaries, categorized decisions, missing-authority checks, recovery prompts, and deterministic runbook exports. It does not score readiness, declare pass or failure, store incident data, or authorize a production action.

    Scenarios
    8 fictional exercises derived from the governed defense-operations report.
    Roles
    8 explicit operational and observer roles.
    Stages
    4 bounded phases from preservation through recovery and learning.
    Operational model

    Keep reading, comparison, and execution as separate authorities.

    The tabletop starts with an untrusted fictional artifact, preserves its identity, compares independent representations, and prevents unresolved evidence from reaching a privileged action. The generated runbook records method and authority gaps; it never becomes an execution token.

    01

    Quarantined reader

    Preserves source bytes and derives bounded structural, visual, semantic, metadata, token, or behavioral views without credentials or write authority.

    02

    Comparison layer

    Records representation agreement, conflict, conditional comparability, and the exact parser, browser, model, registry, or owner still required.

    03

    Privileged executor

    Remains separate. State-changing, financial, externally visible, or production actions require an explicit accountable human and normal operational controls.

    Exercise rule: preserving and comparing evidence can inform a decision, but the simulator cannot authorize containment, publication, financial action, disclosure, recovery, or production change.

    Governed fictional scenarios

    Eight channels, one consistent incident method.

    The scenario titles, vectors, and defensive objectives are derived from the governed defense-operations report. The implementation adds bounded injects and runbook structure without altering that source body.

    Fictional scenario 1

    The Phantom PDF

    A fictional applicant resume contains a parser differential that presents benign visible content while a separate machine-readable layer attempts to influence an HR ranking workflow.

    Defensive objective
    Verify that document structure, rendered output, extracted text, and model-bound context are compared before any ranking or recommendation.
    Evidence planes
    source-bytes, container-structure, rendered-visual, decoded-extracted, token-model-facing, behavior-capability
    Required roles
    Engineering and architecture, Security operations, AI and model operations, Content and data owner, Incident commander
    Review the four staged injects
    1. Stage 1 · Intake and preservationResume enters quarantine

      A fictional resume is accepted by the upload boundary. The file extension and media type appear ordinary, but the exact bytes and parser version have not yet been preserved.

    2. Stage 2 · Representation comparisonRepresentations disagree

      The inert object map reports an additional text-bearing object that is absent from the approved rendered page but appears in one text-extraction path.

    3. Stage 3 · Containment and authorizationRanking workflow requests a decision

      The HR workflow proposes using the extracted text in a candidate-ranking summary before the representation conflict is resolved.

    4. Stage 4 · Recovery and learningKnown-good replay is available

      A flattened, inert derivative and a separately preserved original are available for replay through the corrected pipeline.

    Run this fictional scenario
    Fictional scenario 2

    The Homograph Heist

    A fictional financial document mixes visually confusable Latin and Cyrillic characters to evade an identifier or compliance comparison before reaching an automated trading workflow.

    Defensive objective
    Validate context-specific identifier policy, confusable analysis, and explicit human authorization before any financial action.
    Evidence planes
    source-bytes, unicode-text, rendered-visual, decoded-extracted, token-model-facing, behavior-capability
    Required roles
    Engineering and architecture, Security operations, AI and model operations, Executive owner, Incident commander
    Review the four staged injects
    1. Stage 1 · Intake and preservationIdentifier discrepancy reported

      A fictional financial name appears ordinary in the visible document, while code-point inspection reports mixed Latin and Cyrillic scripts.

    2. Stage 2 · Representation comparisonNormalization is inconclusive

      NFC leaves the mixed-script identifier unchanged. An educational confusable skeleton collides with a protected Latin identifier.

    3. Stage 3 · Containment and authorizationTrading action is proposed

      A downstream workflow proposes an externally visible financial action based on the unresolved identifier.

    4. Stage 4 · Recovery and learningPolicy and comparison controls are revised

      The fictional system can replay the document using a versioned identifier profile and a blocked-by-default action boundary.

    Run this fictional scenario
    Fictional scenario 3

    The Poisoned Vector

    A fictional internal wiki page is indexed into a retrieval store and causes an IT assistant to propose unapproved executable links.

    Defensive objective
    Test provenance, retrieval isolation, source labelling, tool allowlisting, confirmation, and egress controls.
    Evidence planes
    source-bytes, metadata-provenance, decoded-extracted, token-model-facing, behavior-capability, operational-evidence
    Required roles
    Content and data owner, AI and model operations, Security operations, Engineering and architecture, Incident commander
    Review the four staged injects
    1. Stage 1 · Intake and preservationUnexpected wiki revision is retrieved

      A fictional wiki revision with unclear ownership ranks above an approved support article for a common IT query.

    2. Stage 2 · Representation comparisonContext contains conflicting instructions

      The retrieved chunk contains content that conflicts with the user objective and proposes an external executable link.

    3. Stage 3 · Containment and authorizationAssistant proposes distribution

      The IT assistant proposes sharing the link across an employee channel, but no human confirmation or egress allowlist has been applied.

    4. Stage 4 · Recovery and learningStore and workflow are repaired

      The poisoned revision is removed from active retrieval, the approved source is reindexed, and the fixed scenario is replayed.

    Run this fictional scenario
    Fictional scenario 4

    The Screen-Reader Smuggle

    A fictional competitor page exposes hidden accessibility text that is absent from the painted view but appears in an agent-facing semantic representation.

    Defensive objective
    Ensure DOM, rendered output, accessibility semantics, geometry, and source provenance are compared before generated market research is trusted.
    Evidence planes
    source-bytes, rendered-visual, semantic-accessibility, decoded-extracted, behavior-capability, operational-evidence
    Required roles
    Engineering and architecture, Security operations, AI and model operations, Content and data owner, Independent observer
    Review the four staged injects
    1. Stage 1 · Intake and preservationSemantic-only text is observed

      The fictional browser snapshot contains text in the accessibility tree that is not visible in the approved screenshot.

    2. Stage 2 · Representation comparisonRepresentations conflict

      A role/name locator and text extraction include the hidden content, while visual hit testing reports no visible target.

    3. Stage 3 · Containment and authorizationPublication is proposed

      The scraping agent proposes a market-research statement derived from the conflicting content.

    4. Stage 4 · Recovery and learningBrowser policy is corrected

      The workflow now records semantic, visual, and geometry views separately and marks conflicts for human review.

    Run this fictional scenario
    Fictional scenario 5

    The Bidi-Override Backdoor

    Fictional source code contains bidirectional controls that cause visual order to diverge from logical compiler order during automated review.

    Defensive objective
    Verify exact-byte preservation, visible control warnings, compiler or pre-commit diagnostics, and separation between code reading and execution.
    Evidence planes
    source-bytes, unicode-text, rendered-visual, decoded-extracted, behavior-capability, operational-evidence
    Required roles
    Engineering and architecture, Security operations, AI and model operations, Incident commander, Independent observer
    Review the four staged injects
    1. Stage 1 · Intake and preservationControl characters detected

      A fictional source diff contains explicit bidirectional override characters inside a comment or string region.

    2. Stage 2 · Representation comparisonHuman and compiler views diverge

      The visual editor order differs from logical token order, while a compiler diagnostic is not yet available in the review record.

    3. Stage 3 · Containment and authorizationMerge approval is requested

      An automated review suggests approval despite the unresolved bidi control warning.

    4. Stage 4 · Recovery and learningKnown-good source is prepared

      The controls are escaped or removed in a separately reviewed revision, and compiler warnings are enabled by default.

    Run this fictional scenario
    Fictional scenario 6

    The Supply-Chain Exfiltration

    A fictional orchestration dependency changes prompt construction so sensitive values are appended to outbound URL parameters.

    Defensive objective
    Test artifact identity, SBOM and provenance claims, build verification, output validation, and egress allowlisting.
    Evidence planes
    source-bytes, container-structure, metadata-provenance, decoded-extracted, behavior-capability, trust-verification
    Required roles
    Engineering and architecture, Security operations, AI and model operations, Privacy or legal, Incident commander, Executive owner
    Review the four staged injects
    1. Stage 1 · Intake and preservationDependency identity differs

      A fictional package digest does not match the previously reviewed lock file, while the package name and version string appear unchanged.

    2. Stage 2 · Representation comparisonPrompt construction changed

      A deterministic diff shows that the dependency modified prompt or URL-construction behavior.

    3. Stage 3 · Containment and authorizationOutbound request is proposed

      The fictional runtime proposes an external URL containing sensitive placeholder values in query parameters.

    4. Stage 4 · Recovery and learningTrusted build is restored

      A known-good artifact and verified build record are available for redeployment in the fictional environment.

    Run this fictional scenario
    Fictional scenario 7

    The Tokenizer Collision

    A fictional adversarial suffix produces unexpected token boundaries and a policy-violating model output under one specific model and tokenizer configuration.

    Defensive objective
    Measure fail-closed exception handling, preserve exact text and token configuration, and prevent model output from reaching privileged actions.
    Evidence planes
    source-bytes, unicode-text, token-model-facing, transformation-survival, behavior-capability, operational-evidence
    Required roles
    AI and model operations, Engineering and architecture, Security operations, Incident commander, Independent observer
    Review the four staged injects
    1. Stage 1 · Intake and preservationUnexpected output recorded

      A fictional evaluation returns a policy-violating output for a bounded test input under one declared model/tokenizer pair.

    2. Stage 2 · Representation comparisonBoundary differential observed

      Educational token inspection shows materially different boundaries after a declared normalization or invisible-character transformation.

    3. Stage 3 · Containment and authorizationDownstream use is requested

      A downstream system proposes rendering or acting on the output despite the failed evaluation.

    4. Stage 4 · Recovery and learningEvaluation boundary is updated

      A new regression fixture and explicit tokenizer/model identity are added to the fictional test suite.

    Run this fictional scenario
    Fictional scenario 8

    The Image OCR Hijack

    A fictional receipt contains micro-text or another low-salience visual layer that an OCR path extracts as an instruction to approve a maximum expense.

    Defensive objective
    Validate pixel, metadata, OCR, accessibility, and model-input separation before any financial approval action.
    Evidence planes
    source-bytes, rendered-visual, metadata-provenance, decoded-extracted, token-model-facing, behavior-capability
    Required roles
    Engineering and architecture, Security operations, AI and model operations, Content and data owner, Executive owner, Incident commander
    Review the four staged injects
    1. Stage 1 · Intake and preservationReceipt enters image quarantine

      A fictional receipt image is accepted. The raw file, dimensions, metadata, and exact transformation path are not yet fully recorded.

    2. Stage 2 · Representation comparisonOCR and human reading diverge

      Prepared OCR reports a low-salience instruction-like phrase not present in the approved human description.

    3. Stage 3 · Containment and authorizationMaximum approval is proposed

      The expense agent proposes an approval action based on OCR text rather than an authorized policy field.

    4. Stage 4 · Recovery and learningSanitized derivative and policy are available

      A re-encoded derivative, separated OCR channel, and corrected approval policy are ready for replay.

    Run this fictional scenario
    Decision ownership

    Make missing authority visible before it becomes a hidden dependency.

    The laboratory accepts two to eight role IDs. Each scenario declares a minimum role set, and the generated record lists every required role that was not represented.

    Role IDVisible labelExercise responsibility
    executive-ownerExecutive ownerOwns risk tolerance, consequential-action policy, and resource decisions.
    engineeringEngineering and architectureOwns parser isolation, representation comparison, control implementation, and technical recovery.
    security-operationsSecurity operationsOwns triage, evidence preservation, containment coordination, and detection follow-up.
    ai-model-operationsAI and model operationsOwns model, retrieval, tokenizer, prompt-construction, and evaluation boundaries.
    content-data-ownerContent and data ownerOwns source authority, data lineage, publication, and retrieval-store stewardship.
    privacy-legalPrivacy or legalOwns privacy, disclosure, contractual, regulatory, and retention constraints.
    incident-commanderIncident commanderOwns exercise coordination, decision logging, escalation, and recovery sequencing.
    independent-observerIndependent observerRecords gaps, contradictions, missing evidence, and lessons without directing the exercise.
    Progressive staged reveal

    Move from preservation to recovery without skipping the representation gap.

    Every exercise uses the same four stages. JavaScript progressively hides unrevealed decision fields; no-JavaScript visitors can still submit the complete native form.

    1. 01

      Stage 1 · Intake and preservation

      Recognize the fictional anomaly, preserve source identity, and establish initial ownership.

    2. 02

      Stage 2 · Representation comparison

      Compare independent evidence planes and identify the exact unresolved representation gap.

    3. 03

      Stage 3 · Containment and authorization

      Limit capabilities, require authority, and prevent an unverified representation from driving action.

    4. 04

      Stage 4 · Recovery and learning

      Restore a known-good boundary, verify controls, and record lessons and unanswered questions.

    preserve

    Preserve

    Protect raw and transformed evidence before changing the system.

    inspect

    Inspect

    Use an inert, bounded representation-specific analysis path.

    contain

    Contain

    Limit reachable capabilities or isolate the affected boundary.

    verify

    Verify

    Name the external authority required to establish a claim.

    authorize

    Authorize

    Require explicit authority before a consequential action.

    recover

    Recover

    Restore a known-good state and validate the repaired boundary.

    defer

    Defer

    Record why the decision cannot yet be made and what evidence is missing.

    Deterministic method record

    Export the exercise without pretending it is an incident verdict.

    Identical inputs under the same release produce the same JSON and Markdown bytes, Tabletop ID, and full SHA-256. The record contains no timestamp, request address, random value, hidden history, or readiness score.

    Identity and source

    Site release, governed report path and SHA-256, scenario source row, participating roles, missing roles, and fictional-state declaration.

    Evidence and decisions

    Staged injects, required evidence, dependencies, decision category, bounded decision note, and exact external-verification classes.

    Runbook and questions

    Ordered runbook steps, containment and recovery notes, unanswered questions, lessons prompts, and explicit non-conclusions.

    Privacy boundary

    No account, cookie, analytics, database, browser storage, upload, mail, model, registry, DNS, trust-service, tool, or production action.

    Source and authority boundary

    Fictional method, exact source identity, no live-action claim.

    What the simulator establishes

    • Which fictional evidence should be preserved and compared.
    • Which role or external authority remains missing.
    • Which bounded decision and runbook step was recorded.
    • Whether identical inputs reproduce the same method record.

    What it does not establish

    • That a real incident occurred or that an actor had malicious intent.
    • Severity, probability, readiness, compliance, certification, or assurance.
    • Production parser, browser, model, registry, cryptographic, or organizational behavior.
    • Authority to contain, disclose, recover, communicate, pay, merge, approve, or deploy.
    Governed source

    research/reports/machine-tradecraft-defense-maturity.md

    SHA-256 77ede5cba60a8121d3054a50969c1ab6950b046deb484f759b24d20790136d5e

    Read the defense-operations source section