PUBLISHED CONTRACT · 13 CASES · NO LIVE TOOLS

Can your boundary hold?

An AI agent proposes a tool call. Your host freezes the proposal, checks its exact policy, binds approval to the digest, and reports what happened at dispatch.

These are project-authored expected results. This board does not run an adapter or claim an independent pass.

AGENT REQUEST → HOST BOUNDARYEXAMPLE / 03
01send_messageproposed
02outsider@example.testchanged destination
03No exact host rulepolicy decision
04BLOCKED0 tool calls

Illustration of the published changed-destination fixture; no message is sent.

13published adversarial cases
0live services or real approvals
PASS / FAIL / MIXEDoutside reports invited

Every edge has an expected result.

Select a case to inspect its host fixture, expected receipt, and tool-call count. Filters show the expected outcome, not observed performance.

Loading the published cases…

SELECTED FIXTURE

Choose a case

The contract is loading.

Expected policy decision—
Expected receipt—
Expected tool calls—
Receipt reason—

Source of truth: machine contract. The browser holds a copy checked against that file in repository tests.

PUBLISHED CASE RECORD
Select a case to inspect its JSON.
YOUR CODE · YOUR REPOSITORY · REPRODUCIBLE REPORT

Make the claim falsifiable.

Implement the adapter outside this repository, then run the harness against your executable. It prints case-by-case failures and hashes of the exact challenge and vector files. Keep your implementation and evidence outside the project reference code.

node scripts/run-action-boundary-challenge.mjs -- ./your-adapter

Run from a checkout of aml-core; replace ./your-adapter with your executable. The harness uses simulated callbacks only.