Threat Tribunal

How AI Can Be Manipulated

Prompt injection means giving an AI model instructions designed to manipulate its response. Those instructions can come directly from a user or hide inside documents, email, images, retrieved records or other content the application asks the model to process.

Threats covered

Resource and cost abuse, often called denial of wallet, can cause serious harm without a data breach. One crafted or oversized upload can trigger repeated model calls before anyone notices the rising bill. The review checks file limits, retry and queue limits, model quotas, and spend alerts with a named owner and response path.

Security boundary. Content filtering can reduce risk, but it does not replace sign-in, permission checks, data-shape validation or minimum necessary access. High-impact actions still need deterministic code: fixed application rules that produce the same decision from the same input.

Fictional Security Examples

Illustrative AI and LLM Threat Examples

Illustrative only - not an engagement finding

Uploaded Documents Can Make AI Access Another Customer's Data

Critical See where this finding applies on the diagrams: AI services zone · AI analysis step · CP5 AI-call boundary
Business summary
Every uploaded document is AI input that the application must not trust. If hidden instructions can influence which records are retrieved or who may access them, prompt injection can become a way for one customer's data to reach another.
Assessment profile
Ease of attack
Requires crafted document contentText in a PDF, email or image can request another customer’s records. The application becomes vulnerable when it trusts the model-selected record or search scope without rechecking it.
Impact
Cross-customer disclosure
  • Customer records containing personally identifiable information (PII) could cross a customer boundary.
  • PII could appear in an AI response, log entry or generated working paper.
Detection difficulty
Deep trace requiredTrace model output through retrieval, authorization and subsequent data access.
Pattern
Prompt-injection boundaryUntrusted document content influences which records the application retrieves.
Expected fix Verification
  1. Expected fix Treat document text as data Azure AI Document Intelligence output and background-worker prompt construction: Keep extracted document text separate from application instructions.
    Verification Ignore hidden instructions Test documents cannot reveal planted cross-customer data, while authorized classification still works.
  2. Expected fix Derive scope server-side FastAPI authentication context and worker job payloads: Derive scope only from the signed-in user and server-side data.
    Verification Protect server-derived scope Model-proposed identifiers cannot override the signed-in user's scope.
  3. Expected fix Reauthorize every operation Cosmos DB and Blob Storage data-access methods: Recheck access immediately before every read or write.
    Verification Recheck at the data boundary An out-of-scope request is denied immediately before the data operation.
  4. Expected fix Validate model output strictly Azure OpenAI response parsing and schema validation: Reject unknown identifiers, actions, fields or destinations.
    Verification Reject unknown output Unknown values are rejected before any database or storage operation.
  5. Expected fix Mediate all data access Azure OpenAI integration and server-side data-access adapters: Give the model no database or storage credentials or direct route.
    Verification Block direct model access The model has no credential or route to database or storage records.
  6. Expected fix Fail closed and log safely FastAPI and background worker authorization failures: Deny scope changes and record them in server-side logs without document or customer content.
    Verification Deny and log scope changes Each attempt is denied and logged without document or customer content.

Illustrative only - not an engagement finding

AI Responses and Error Messages Can Reveal Personally Identifiable Information

High See where this finding applies on the diagrams: AI services zone · telemetry and error paths · CP8 telemetry separation
Business summary
Tax documents are dense with personally identifiable information (PII). Copies can escape the approved workflow through model responses, prompts, traces, errors or support logs even when the primary data store is protected.
Assessment profile
Ease of attack
Output or log accessA caller or support user can trigger or inspect responses that contain more document content than the workflow requires.
Impact
PII disclosure
  • Names, identifiers and financial details could reach unauthorized users or support tools.
  • Monitoring data could retain PII longer than the source document.
Detection difficulty
End-to-end content traceTrace planted test details through document analysis, model input and output, API errors, status records and monitoring logs.
Pattern
Minimize, mask and restrictDocument content crosses more application and monitoring surfaces than the approved business process requires.
Expected fix Verification
  1. Expected fix Minimize model input Background worker after Azure AI Document Intelligence: Send Azure OpenAI only the fields required for classification.
    Verification Limit model exposure Captured requests show that Azure OpenAI receives only the approved minimum fields.
  2. Expected fix Mask details the model does not need Worker content-preparation step: Mask PII before the model call when classification does not require it.
    Verification Preserve useful classification Planted unnecessary details are masked while authorized classification still produces the expected result.
  3. Expected fix Redact every outbound surface FastAPI responses, worker errors, Cosmos DB status records and PDF assembly: Redact or reject PII before content leaves the approved workflow.
    Verification Stop secondary disclosure Planted details do not appear in responses, errors, ordinary logs or unapproved status records, while approved output retains the minimum required information.
  4. Expected fix Protect necessary diagnostic traces Azure Monitor and Log Analytics: Keep diagnostics containing PII or document content in dedicated tables with explicit retention.
    Verification Expire protected traces Test traces are removed when the approved retention period ends.
  5. Expected fix Restrict diagnostic access Log Analytics role assignments: Grant access to restricted diagnostic tables only to named support or security roles.
    Verification Deny broad log access A user without an approved support or security role cannot read the restricted diagnostic tables.

Illustrative only - not an engagement finding

Document Text Can Change How AI Classifies It

Medium See where this finding applies on the diagrams: AI services zone · AI analysis step · CP5 AI-call boundary
Business summary
Even when a model cannot take privileged action, hidden instructions can still bias classification, increase manual review and route documents through the wrong business process.
Assessment profile
Ease of attack
Crafted document inputHidden instructions in extracted text can ask the model to ignore the category rules or invent confidence.
Impact
Misclassification and workflow disruption
  • Documents could be omitted, misrouted or accepted with misleading confidence.
  • The final working paper could become less reliable.
Detection difficulty
Adversarial comparison requiredCompare crafted documents and model responses with a reviewed baseline and the recorded reviewer path.
Pattern
Bounded classificationUntrusted document content attempts to influence a closed-set category and confidence decision.
Expected fix Verification
  1. Expected fix Keep categories closed Azure OpenAI classification deployment: Allow only the approved category list.
    Verification Reject unknown categories Direct and document-borne injection cases stay inside the approved list, while normal cases still produce valid categories.
  2. Expected fix Validate category and confidence Background-worker response parser: Enforce a strict schema and reject malformed values or additional fields.
    Verification Reject malformed output Unknown categories, malformed confidence values and extra fields are rejected without saving a classification.
  3. Expected fix Separate instructions from content Worker boundary between Azure AI Document Intelligence and Azure OpenAI: Treat extracted text only as data inside the classification prompt.
    Verification Ignore document instructions Crafted PDFs, emails and images cannot replace the classification rules or force a preferred confidence value.
  4. Expected fix Route uncertainty to people FastAPI review workflow and Cosmos DB status: Send uncertain, declined or below-threshold results to a recorded classification review.
    Verification Require recorded review Every uncertain, declined or below-threshold result enters the recorded classification review before it is accepted.
  5. Expected fix Preserve original model evidence Cosmos DB review audit and PDF assembly: Store the original category, confidence and model evidence beside any reviewer correction.
    Verification Keep corrections auditable An approved reviewer can correct the result without changing or replacing the original model evidence.

Assessed control pattern - not an engagement finding

Bounded, Deterministic LLM Controls

Design position See where these controls apply on the diagrams: AI services zone · AI analysis step · CP5 AI-call boundary

Content filtering can reduce harmful model responses, but it does not replace authorization, output validation or approval in application code.