Global site search

Search guides, labs, glossary, and research

Type two or more characters to search.

Start with a channel, artifact, or defense term

Examples include zero-width, metadata, tokenizer, or prompt injection.

    Controlled AI Tests · Confused-deputy side effects

    Retrieval and Agent Trust-Boundary Simulator

    The simulator performs no model call or external action and demonstrates how application controls change a fixed harmless scenario.

    Quick answer

    What does the Agent Trust Simulator show?

    The simulator performs no model call or external action and demonstrates how application controls change a fixed harmless scenario.

    Human visibility
    Retrieved evidence and proposed action
    Machine receiver
    Deterministic local policy simulator
    Robustness
    Control-configuration specific

    Research boundary: this page uses bounded, inert data and fixed safe examples. It never executes decoded content, requests secrets, calls third-party services, or performs actions against external systems.

    FIRST RUN / THREE STEPS

    Start with the prepared, bounded workflow.

    Nothing runs automatically
    1. Review the prepared starter input

      A bounded benign input is already present. Change it only when you are ready to test a different authorized artifact.

    2. Run Simulate trust boundaries

      Run the normal local synthetic simulation with the values above.

    3. Scan before expanding

      Confirm the fixed benign marker and safety qualification before copying or downloading any generated fixture.

    INPUT / CONTROL PLANE

    Prepare the input and choose one action.

    Laboratory status: Ready

    The recommended first run is separated from alternate analyses. Inputs and selected files stay on this host.

    Current input state Defense-in-depth loaded

    These bounded starter values are ready to inspect. Review them before running the recommended action.

    Choose the prepared evidence scenario used for this run. The selection changes only this local analysis.

    Choose the source labels used for this run. The selection changes only this local analysis.

    Choose the tool policy used for this run. The selection changes only this local analysis.

    Choose the human confirmation used for this run. The selection changes only this local analysis.

    Choose the egress policy used for this run. The selection changes only this local analysis.

    Choose the memory policy used for this run. The selection changes only this local analysis.

    Switch prepared example5 options

    Loading a sample changes only the form values. Review the result and run an action yourself.

    Prepared benign examples
    ACTION HIERARCHY

    Run the recommended first pass.

    Alternate actions remain available below, but the first pass is the clearest place to start.

    Inputs remain on this host. Text operations are size-limited; uploaded files are processed from PHP’s temporary upload and are not retained by the application.

    OUTPUT / MACHINE VIEWS

    Scan the result from summary to evidence.

    Run Simulate trust boundaries to create the first result.

    The prepared starter input is ready. The output will lead with a summary and visible qualifications before the expandable machine views.

    SummaryFindingsMachine views
    Interpretation framework

    The same artifact can produce several valid observations.

    01

    Human view

    What a person naturally reads, sees, or hears.

    02

    Structural view

    What a parser, DOM, container reader, or metadata extractor exposes.

    03

    Decoder view

    What becomes meaningful only with a rule, key, tokenizer, model, or tool.

    04

    Defensive view

    What normalization, rendering, OCR, canonicalization, or policy changes.

    Evidence and decision boundary

    Use the result as bounded evidence, not as a universal verdict.

    The simulator performs no model call or external action and demonstrates how application controls change a fixed harmless scenario.

    LOCAL MODEDeterministic application-control state machine
    REVIEW DATE2026-08-26
    SOURCE BODYPreserved separately from implementation claims
    Computed locally

    Deterministic output produced by this bounded runtime.

    • The selected scenario and five explicit application-control states
    • A deterministic proposed action, policy trace, control-gap count, and allow/block result
    • Proof that the local simulation used no model, credential, write tool, network request, or external side effect
    Bounded approximation

    Useful subset or model that does not establish full conformance.

    • The state machine does not model language-model probabilities, instruction hierarchy, jailbreak behavior, or adaptive attackers
    • The control-gap score is a local scenario rule, not a risk score or attack-success estimate
    • Mock tools and evidence bundles do not reproduce a production agent framework
    Escalate for

    Claims that require an exact parser, trust system, model, or human review.

    • Model susceptibility requires controlled evaluation of the exact model, prompt stack, retriever, tools, permissions, and versions
    • Tool authorization and confirmation must be verified at the actual execution boundary
    • Incident and benchmark claims require primary evidence and current vendor or vulnerability records
    Interpretation rule

    Record the receiver and transformation.

    Machine-decodable is receiver-relative. Record the parser, preprocessing, codebook, tokenizer, key, model, and transformation path before generalizing from one result.

    Limitations

    What this page does not prove

    This deterministic state machine does not measure a language model’s probabilistic susceptibility. It validates only the represented application controls and fixed scenario.

    Deterministic review material

    Download the exact benign fixtures used for the evidence boundary.

    These local files are supplied for repeatable inspection. The application does not fetch them automatically, execute their content, or treat a fixture result as external verification.

    Agent policy scenario

    Documents one deterministic simulated agent-control configuration; it contains no credentials, executable tool call, or external side effect.

    Type
    JSON fixture
    Bytes
    384
    SHA-256
    a05f5e0154f76ee8fbf3…
    Download fixture