Verifiable evidence surface

Pack registry

Portable test packs for governance evaluation.

Packs define bounded test surfaces for SIR. This page separates packs currently foregrounded on the site from future pack candidates that are not being presented as current availability.

Current packs

Foregrounded now

Gate coverage is the proportion of a suite's expected-block rows the content gate blocks, measured directly. Runner evaluability is whether the standard runner can execute the suite at all. These are independent: a suite can have non-zero coverage and still not be runner-evaluable. Suites marked Not runner-evaluable have no same-named enforcement policy pack; selecting them through the runner returns a policy-load block for every row without evaluating prompt content. See the gate coverage methodology.

Current | Gate coverage 74/100

Domain pack

eu_ai_act_compliance_pressure

Exercises prompts and pressure patterns that test whether stated governance, disclosure, and control boundaries hold under bounded EU AI Act-aligned scrutiny. This pack is a pressure test, not a claim of full regulatory coverage.

Current | Gate coverage 100/100

Domain pack

generic_safety

Broad baseline coverage for bypass attempts, harmful requests, prompt injection patterns, and related control-path checks.

Current | Gate coverage 26/26

Domain pack

support_operator_override

Tests override-style prompts that try to invoke authority, support escalation, or operator language to get past control boundaries.

Current | Gate coverage 27/27

Domain pack

data_exfiltration_pressure

Exercises attempts to extract restricted content, secrets, or internal data through pressure, reframing, or repeated request patterns.

Current | Gate coverage 6/10

Scenario pack

scenario_injection_chain

Not runner-evaluable

Scenario-based coverage for chained injection behavior where unsafe instructions arrive across multiple steps rather than in one isolated prompt. Selecting this pack does not currently evaluate its suite. When the suite content is evaluated directly, four of ten expected-block turns are not caught by the current rule set.

Current | Gate coverage 1/2

Scenario pack

scenario_tool_injection

Not runner-evaluable

Text-only tool-channel and function-call style injection language appearing in assistant or user content. A focused complement to scenario_injection_chain, not a broad tool-execution test surface.

Current | Gate coverage 5/15

Domain pack

mental_health_clinical

Not runner-evaluable

Mental-health and clinical request patterns. Prompt content is stored base64-encoded rather than in plain text, which is what the encoded visibility label refers to.

Current | Gate coverage 0/5

Domain pack

account_recovery_fraud

Not runner-evaluable

Tests recovery and identity-reset style requests that attempt to bypass process controls, impersonate users, or manufacture authority. Selecting this pack does not currently evaluate its suite. When the suite content is evaluated directly, none of its expected-block cases are caught by the current rule set.

Future pack candidates

Not presented as current availability

The items below are future candidates only. They are shown as possible expansion areas, not as current implemented reality on this site.

PII Protection

Identity data, re-identification, and personal data handling cases.

Healthcare Compliance

Healthcare process and restricted-handling cases.

Financial Services

Transaction, account, and high-trust financial handling cases.

Legal & Contracts

Contractual and sensitive document handling cases.

Insurance & Underwriting

Evidence review and underwriting-adjacent workflow cases.

Code Generation Safety

Code, command, and secret-handling cases.

Educational Content

Restricted-topic and assessment-integrity cases.

Status note

Pack names on this page should be read as pack registry labels. Foregrounded items are current site-facing packs. Future candidates are shown separately to avoid implying present availability, maturity claims, or certification scope.