Threat Tribunal

Retesting AI Safety After Changes

A model, prompt, analyzer or searchable-data change can alter security behaviour even when ordinary application tests still pass.

Reusable attack-test pack

  • Direct prompt-injection attempts
  • Hidden instructions inside documents
  • Attempts to reveal personally identifiable information (PII)
  • Invalid AI responses and unauthorized-action requests

See AI safety regression evidence for when the pack runs, what blocks release and which evidence the team retains.

Fictional Security Examples

Illustrative Regression and Guardrail Examples

Illustrative only - not an engagement finding

AI Can Trigger a System Action That Was Not Approved

Critical See where this finding applies on the diagrams: model output path · bounded output handling · CP5 AI-call boundary
Business summary
An AI response is untrusted text, not permission to act. The application must never turn it directly into an email, database change, file deletion, paid operation or other system action. Before anything happens, server-side rules must check the signed-in user, the affected customer, the requested action and any required human approval.
Assessment profile
Ease of attack
A manipulated response becomes a commandAn attacker changes the model's response, and the application uses it to choose a tool, destination or operation without an independent check.
Impact
Unauthorized system action
  • Model output could trigger disclosure, destructive changes or external communication.
  • It could start paid processing outside the intended workflow.
Detection difficulty
Follow the response to the final actionTrace the Azure OpenAI response through the worker, permission checks, approval step and resulting operation.
Pattern
The model recommends; the server decidesThe model may suggest an action, but fixed server-side rules must decide whether it is allowed.
Expected fix Verification
  1. Expected fix Keep every operation behind application code Background worker and server-side service adapters: Give Azure OpenAI no credentials and no direct route to tools or connected services.
    Verification Confirm the model cannot act directly The model has no credential or connection that can invoke a tool or connected service.
  2. Expected fix Allow only defined actions Worker action-selection code: Match a checked model response to a fixed list of actions and permitted fields.
    Verification Reject unsupported actions Unknown actions, destinations and values fail before anything is sent, changed or charged.
  3. Expected fix Take customer and destination details from the server FastAPI sign-in context and worker jobs: Set the customer, destination and permitted values from trusted server data, not from the model response.
    Verification Ignore model-proposed scope changes A customer, destination or value proposed by the model cannot replace the trusted server value.
  4. Expected fix Check permission immediately before acting Server-side action and data-access code: Recheck the signed-in user and customer immediately before each operation.
    Verification Deny the action at the final boundary An otherwise valid action with the wrong customer or user permission is denied immediately before execution.
  5. Expected fix Require accountable approval FastAPI review workflow and worker action state: Require a named person's recorded approval before any external, irreversible or high-impact operation.
    Verification Stop before high-impact work Every external, irreversible or high-impact test action waits for recorded human approval.
  6. Expected fix Record who requested and approved the action Azure Monitor and application audit logs: Record the user, requested action, approval, enforcing code and outcome without personally identifiable information (PII), document content or model output.
    Verification Keep actions traceable The audit trail identifies the user, request, approval, enforcing code and outcome without storing PII or document content.

Illustrative only - not an engagement finding

Unchecked AI Responses Can Be Saved as Trusted Records

High See where this finding applies on the diagrams: model output path · bounded output handling
Business summary
An AI response can contain invented categories, unexpected fields, unsafe links or code-like text. If the application saves or displays it without checking, those values can corrupt a customer record, run in a browser or appear in the final working paper.
Assessment profile
Ease of attack
An unexpected AI response is acceptedThe model returns an invented category, identifier, field or piece of document content, and the application saves it without applying fixed validation rules.
Impact
Incorrect records or unsafe displayed content
  • Unexpected values could put a record into the wrong workflow state or cause active content to run.
  • Unchecked AI text could reach a customer's final working paper.
Detection difficulty
Follow the response to every destinationInspect how the response is checked, written to Cosmos DB, displayed in React and approved before PDF assembly.
Pattern
Check before saving or displayingAccept only expected Document Intelligence and Azure OpenAI values before they reach storage, the browser or a working paper.
Expected fix Verification
  1. Expected fix Define exactly what a valid response contains Document Intelligence and Azure OpenAI response models: Accept only named categories, expected fields and required data types.
    Verification Accept a valid response A response with an approved category, fields and data types is checked and saved correctly.
  2. Expected fix Stop when a response cannot be checked Background-worker response checks: Stop when the response cannot be read or fails validation, and do not save part of the workflow state.
    Verification Save nothing from an invalid response An invalid response fails before any classification or workflow value is saved.
  3. Expected fix Reject unexpected values before saving Cosmos DB data-access code: Reject extra fields, unknown values and categories that are not approved before every write.
    Verification Keep unexpected output out of Cosmos DB Extra fields, unknown values and unapproved categories are rejected before a Cosmos DB operation occurs.
  4. Expected fix Display AI text as text, not code React, application logs and PDF assembly: Encode model content for each destination so it cannot be interpreted as HTML, script or a control sequence.
    Verification Keep active content harmless HTML, script syntax and control characters remain inert in browser, PDF and log output.
  5. Expected fix Require review before delivery FastAPI review state, Cosmos DB and PDF assembly: Allow only recorded, human-approved model content into the final working paper.
    Verification Publish reviewed content only Only human-approved classifications and content appear in the final test working paper.

Illustrative only - not an engagement finding

AI Safety Changes Can Ship Without Retesting

Medium See where this finding applies on the diagrams: delivery zone [planned] · AI analysis step · CP6 supply-chain gate
Business summary
Changing the AI model, the instructions it receives, how documents are analyzed or the information it can search may allow an attack that the previous release blocked. Before release, the team must rerun known attacks and normal user tasks. If an attack succeeds, a normal task fails or a required result is missing, the release must stop.
Assessment profile
Ease of attack
A change can weaken protection silentlyA new model, instruction, analyzer or response-handling change can behave differently while ordinary application tests continue to pass.
Impact
Previously blocked attacks can work again
  • Prompt injection or PII disclosure that was blocked before could return.
  • AI output checks or action limits could weaken without an obvious application error.
Detection difficulty
Ordinary tests may still passThe problem appears only when the same attack and normal-use tests are run against the changed AI setup and compared with the approved result.
Pattern
Retest attacks before every releaseEvery AI-related change must run the approved attack tests and normal business cases before deployment.
Expected fix Verification
  1. Expected fix Store the exact test setup Security-test evidence store: Keep the attack cases, test documents, approved results, Document Intelligence analyzer, Azure OpenAI deployment, instructions and response format together under version control.
    Verification Show exactly what was tested The release record names every model, analyzer, instruction set, response format, test document and approved result used by the run.
  2. Expected fix Rerun tests after every relevant change Azure DevOps security-regression job: Run the versioned set when AI services, prompts, response handling or control code change.
    Verification Test attacks and normal work together Every relevant change runs the complete set. Known attacks stay blocked and normal business cases still work.
  3. Expected fix Require approval when a result changes Azure DevOps protected review: Require a named reviewer to explain and approve every changed expected result.
    Verification Reject unapproved result changes A changed expected result cannot pass the release gate without the named review and approval.
  4. Expected fix Stop release when protection gets weaker Azure DevOps release gate: Stop release when a required result is missing or a previously blocked attack now succeeds.
    Verification Block missing or failed security results A missing result or newly successful attack case blocks release.
  5. Expected fix Keep the results with the release Azure DevOps release artifacts: Store the result comparison, reviewer identity and approval with the release record.
    Verification Make the release decision reviewable The release record keeps the exact test comparison and named approval that allowed deployment.