Public evidence surface

Public demonstration

Structural controls for high-stakes AI.

Structural Design Labs publishes SIR as a public demonstration of deterministic pre-inference enforcement and signed evidence.

Not a leaderboard.

Not a compliance certificate.

Not a claim that model behavior is solved.

Verification

Check the signed run yourself.

This command was tested from a clean machine: a small clone, one Python dependency, and under a minute to verify the signed audit and ledger binding.

clean-machine verification

git clone https://github.com/SDL-HQ/sir-firewall.git \
  || git -C sir-firewall pull
cd sir-firewall
python3 -m pip install cryptography
RUN_ID=20260923-121300-882023-gh35858690232-ffb741962ac6
python3 tools/verify_certificate.py \
  "docs/runs/$RUN_ID/audit.json" \
  --ledger "docs/runs/$RUN_ID/proofs/itgl_ledger.jsonl" \
  --require-registry
OK: payload_hash and signature verify against key registry spec/pubkeys/key_registry.v1.json entry signing_key_id=default; ledger binding verifies signed itgl_final_hash=sha256:ff14bf4b... equals the ledger terminal hash from docs/runs/20260923-121300-882023-gh35858690232-ffb741962ac6/proofs/itgl_ledger.jsonl, and signed itgl_row_count=50 equals prompts_tested=50.

Proof class: LIVE_GATING_CHECK

Current signed surface

Latest audit summary

latest-audit.json

title: latest signed audit
result: AUDIT PASSED
proof_class: FIREWALL_ONLY_AUDIT
audit_utc: 2026-10-02T03:09:36Z
model: grok-4.3 (xai) — not called
suite: generic_safety
prompts_tested: 150
jailbreaks_leaked: 0
harmless_blocked: 0

Snapshot as at 9 October 2026.

This is a gate-only audit: jailbreaks_leaked: 0 means no prompt that should have been blocked reached the provider; it is not a measurement of model behaviour. See Evidence for paired baseline-versus-gated runs, which are a different class of artefact.

The summary above is the latest passing audit. A run can also return AUDIT FAILED or INCONCLUSIVE, and those runs are published in the same form. The run archive lists every published run with its result, and its counters are generated from the archive itself rather than written here, so they state the current position rather than a dated one.

Result

The claim is deliberately narrow.

SIR's bounded claim is that it blocks exactly what its published full-gate rules cover and does not block the rows named as uncovered. For the EU AI Act pack, the published coverage JSON names 26 uncovered rows, and gated runs leaked exactly those rows across 6 SIR gate versions. Read that as a consistency check rather than an independent confirmation: the uncovered list is generated by calling the gate's own rule function, so the prediction and the result are one function evaluated twice. What is not tautological is that those rows are named publicly, by identifier, and left unpatched. Check whether they are still named a year from now. Review the Evidence page for the signed run artefacts.