How I assess a production GenAI document platform across its source code,
cloud footprint and AI/LLM attack surface: two complementary reviews, every
conclusion tied to reproducible evidence.
3 frontier models in adversarial roles2 independent evidence streams10 control-domain chapters32 worked fictional examples
How AI/LLM threats, bounded model controls, adversarial regression,
application boundaries, safe file and output handling, and software
supply-chain risks are assessed through source and repository evidence.
How identities, networks, configuration, data services, containers,
release controls and operational telemetry are assessed from the
perspective of the running cloud environment.
Relationship. The Azure review assesses running
configuration, service relationships and observable operational
controls using safe, read-only evidence. The application review
assesses source code, infrastructure as code, dependencies and delivery
artifacts. Areas of overlap are intentionally compared. Agreement
between the two evidence streams strengthens confidence, while
differences reflect their distinct scopes and evidence boundaries.
Assessment Methodology
I ran the assessment with Threat Tribunal, my multi-model security
harness: complementary proposal, challenge and arbitration roles across
Fable, Claude Opus and GPT-5.6. No model-generated claim became
delivery work without human review and supporting evidence. The work
produced one report grounded in the deployed Azure footprint and another
grounded in application source and repository artifacts.
Threat Tribunal as a method illustration. It contains no client
information; the same pipeline applies to any engagement.
AI and LLM Security Focus
Document-processing applications create a distinctive trust problem:
user prompts, retrieved records and uploaded files can all become model
input. A document that looks ordinary to a person can contain instructions
intended for the model.
Design position. Treat the model as an untrusted,
probabilistic component inside a bounded application. Enforce
authorization, schemas, tool permissions and approval outside the model.
Reports are useful only when teams can act on them. Human-approved
remediation findings are automatically mapped into Azure DevOps Features and User
Stories with ownership, acceptance criteria, dependencies and verification
evidence. My evidence-tracked, idempotent publisher synchronizes the mapped
backlog.
Agree the systems, evidence boundaries and standards in view.
Map
Reconstruct the as-built architecture and trust boundaries.
Assess
Run Threat Tribunal across both evidence streams.
Validate
Human review of every model-proposed observation.
Deliver
Governed backlog, verification evidence and the delivery roadmap.
Designed as a one-month engagement.
Public Edition and Disclosure
External-sharing boundary. This purpose-built public
edition is intended for external sharing. It contains no client names,
tenant or subscription identifiers, resource names, internal URLs,
work-item references, source paths, commit identifiers, screenshots,
sample data, credentials, or combinations of details that could
reconstruct a specific finding.
Both reports include explicitly labelled fictional severity examples
to demonstrate risk communication, remediation design and verification.
They are not engagement findings and do not describe the reviewed
environment.
Their cross-reference links point only to newly authored demonstration
source, Azure and tool-output examples. No client evidence destination
exists in the public editions.
Limitations
This publication explains the assessment scope, methodology, control
domains and generalized security lessons. It does not reproduce
client-specific findings, severity ratings, evidence, remediation status
or current production configuration. It is not a penetration-test
report, certification, assurance opinion or statement that any example
weakness was present. Severity labels in both reports apply only to
explicitly labelled fictional examples. Inclusion of a topic means it
was assessed, not that a finding existed.
About the Author
I'm a Senior Azure AI and Software Architect. I build agentic AI
platforms on Azure and run AI/LLM security and cost optimization
assessments of production GenAI systems. One rule anchors both: code
owns deterministic facts, and LLMs are bound to reasoning, annotation
and judgment.
What an engagement delivers
Two evidence-grounded security reviews: live cloud footprint and application source
An evidence register: every conclusion tied to a reproducible source
A prioritized remediation backlog in Azure DevOps, human-approved before any work item exists
A secure-delivery roadmap that encodes the rules as checks on every build