Analysis

Threat-model a private AI deployment before selecting controls

Map a fictional document assistant's trust boundaries, abuse cases and evidence gaps before choosing controls or admitting real files and tool actions.

The short answer

A private AI threat model identifies the assets, identities, data flows and authority changes around a deployment, then connects specific abuse cases to enforceable controls and evidence. Local inference does not establish that retrieval, tool actions, updates or backups are safe. Keep real documents and consequential actions held while the selected controls lack verified evidence.

Moving inference onto owned hardware changes where a model runs. It does not decide which employee may read a document, whether a generated action is authorised or who can replace the software. A useful private AI threat model starts with those relationships, before selecting a firewall, guardrail or scanning product.

NIST SP 800-207 does not grant implicit trust simply because a resource is local or organisation-owned. This example turns one proposed document assistant into reviewable boundaries, abuse cases and evidence requirements.

Define what the assistant may do

Consider a fictional service workshop. Staff want answers from approved internal documents and, later, help creating follow-up tasks. The proposed application uses an internal identity service, a document index and a locally hosted model. A separate executor would create tasks in an internal project system. No email, shell or general web-fetch tool is included. Model updates enter through a separate staging process; backups go to a separately administered vault.

The example protects document content, access-control lists (ACLs), task integrity, service availability, update authority and recovery material. A service colleague, u-serv, may read document D-S, a service checklist, but not finance document D-F. Separately, the proposed task policy lets u-serv create tasks only in the service project, never the finance project. A finance colleague has a different entitlement. These are invented identities and files, not a description of Mickai infrastructure.

Assume an attacker could control a staff session, contribute a document or substitute an update package. Include accidental mistakes such as restoring an obsolete permission list. Host administrators remain powerful; this exercise does not solve a fully compromised operating system. Every control below is proposed, and every deployment test is NOT RUN. No model or real service was exercised.

Draw authority changes, not just boxes

A trust boundary is crossed when data or an instruction reaches a component that must make its own trust decision. The following table is the readable equivalent of a data-flow diagram. Arrows show the intended paths, not proven isolation. Owner names describe fictional responsibilities.

Proposed document assistant: seven boundaries to investigate
BoundaryData flowAuthority question and owner
B1: identityStaff browser → identity service → applicationWhich authenticated principal made this request? Identity owner and application owner.
B2: documentsFiles → ingestion/index; application ↔ retrievalWhich current ACL permits each chunk, title and citation? Document owner and retrieval owner.
B3: inference and replyAllowed context → model → application → browserCan source text become authority or output become active content? Application owner.
B4: actionsModel proposal → executor → task serviceIs this exact action permitted and approved for its destination audience? Tool owner.
B5: updatesExternal bundle → staging → serving environmentWho may approve these exact model, tokenizer, loader and configuration versions? Release owner.
B6: recoveryFiles, index and configuration → backup vault → isolated restoreAre data, keys and current permissions recoverable together? Recovery owner.
B7: recordsApplication, executor and update events → protected logsWho may see, alter and retain potentially sensitive records? Operations owner.

The model receives only the context selected for a request; it has no direct document-store, update or task credentials. Initial design disables shared answer caching and cross-user conversation reuse. Answers are buffered, never streamed, pending the final permission check. Browser delivery, monitoring and backup copies still belong in the model. Omitting those paths hides where information could escape its intended audience.

Write invariants before choosing controls

For this assistant, four rules must remain true: a request cannot gain another user's document permissions; source text cannot grant a tool capability; an approval cannot authorise a different action or audience; and restoring or updating the system cannot silently restore revoked authority. These are design requirements to test, not assertions about a finished implementation.

OWASP's authorization guidance supports denying access unless it is permitted and checking permissions on every request. Apply that outside the model. Preserve the authenticated principal through retrieval and execution; do not replace it with a broad service account simply because the request came from the assistant.

Original review register: all V1-V7 checks are proposed and NOT RUN
Abuse caseProposed enforcing controlFalsifiable check and remaining risk
T1: wrong or stale identity at B1/B2C1: server-validated identity; ACL filtering before context selection and recheck before disclosure, including citations and saved conversation state.V1: u-serv can use D-S; D-F never enters its model context or returned content, titles or citations. Revoke D-S mid-request: withhold the buffered answer. Residual: compromised authorised sessions.
T2: document instructions at B2/B3/B4C2: treat retrieved text as untrusted data; keep enforcement outside the model. C1 limits retrieval; C3 independently restricts the proposed action.V2: an invented D-S note requesting a finance task cannot cause a tool write or broaden retrieval. Inspect context and executor records, not refusal wording alone. Residual: misleading answers.
T3: action substitution at B4C3: separate executor checks actor, exact arguments, destination audience and single-use approval. Only the named create-task operation is available.V3: an approved service task can run once; changed project, changed content, missing approval or replay produces no extra task. Residual: a person may approve an unsuitable permitted action.
T4: substituted import at B5C4: separate release authority, protected expected bundle manifest, isolated staging and explicit activation/rollback policy.V4: changed bytes, unapproved loader or retired release cannot activate; selected serving version stays unchanged. Residual: approved components may still be vulnerable.
T5: obsolete authority restored at B6C5: isolated recovery with usable keys and current ACL reconciliation before requests resume.V5: restore the fixture, then confirm the revoked u-serv permission remains denied. Missing keys or uncertain current policy keeps service held. Residual: unrecovered changes since backup.
T6: disclosure through replies or logs at B3/B7C6: render inert text with controlled citations; restrict logging to necessary events and protect log access/retention.V6: output markup cannot execute or fetch a remote resource; a dummy confidential marker is absent from ordinary operational logs. Residual: authorised readers can copy visible text.
T7: resource exhaustion at B2/B3C7: input/parser limits, bounded context, request admission and cancellation outside the model.V7: with a proposed four-active-request limit, a fifth is refused before model allocation; cancellation releases capacity once. Residual: permitted requests may still be costly.

These seven cases are a starting set, not a claim of comprehensive coverage. OWASP's threat-modelling framework connects scope, possible failures, mitigations and evaluation. Adapt the cases when the architecture changes. A numerical score invented without exposure or incident evidence would add apparent precision without settling these decisions.

Follow one document through the failure path

Suppose D-S contains the ordinary instruction to check a service schedule, followed by an injected instruction to create a finance task. The document is readable by u-serv, so its arrival in the model context is not itself an authorization failure. The boundary fails if its text changes application permissions or produces an unauthorised side effect. OWASP describes indirect prompt injection through external content and recommends constrained tool permissions alongside other defences.

A stronger system prompt may help behaviour, but it cannot be the permission check. Even if the model emits a perfectly formed finance-task proposal, the executor must reject it. For an allowed service task, show the person the exact content, project and audience, then bind their approval to that immutable proposal. Recheck their authority when executing. Consent does not create permission, and changing any approved field requires a new decision.

Task creation is also a disclosure path. Someone allowed to read finance material must not paste it into a task visible to the whole workshop merely because they can create tasks. This design must check the destination audience against the material's permitted audience, or constrain task content to an approved template that excludes document-derived text. That unresolved choice keeps tool writes disabled in the decision below.

Revocation exposes another hidden assumption. Removing access from the document store is insufficient if an old chunk remains in pending context, a citation or conversation history. The proposed C1 recheck must discard a response whose supporting context is no longer permitted; the next request must rebuild permitted context. V1 therefore needs both an allowed baseline and a permission change, with traces showing which document IDs reached the model.

Include tomorrow's system and its recovery copy

The update boundary concerns more than weight files. Record the model, tokenizer, executable loader, dependencies and runtime configuration together. A matching digest identifies expected bytes only when the expected value comes from a separately protected, trusted process. It does not prove benign behaviour. NCSC's secure-design guidance treats third-party model imports as a supply-chain and isolation concern.

For backups, encryption protects one aspect of storage; it does not show that restoration works. A recovery exercise must recover the required keys, reconcile permissions with current authority and verify denied access before reopening the service. Protect logs too: NCSC identifies logs as sensitive assets. Recording every prompt by default could create another confidential document repository. Restrict audit-record access to assigned operators, separately from ordinary document users.

The reply channel needs its own controls. OWASP's output-handling guidance distinguishes generating text from safely passing it to browsers or other components. For this design, model-written markup stays inert; citations are application-generated links to authorised documents, not arbitrary model-supplied URLs.

Make a decision that evidence can change

Use the register to choose the next bounded activity. Prioritise the disclosure, unauthorised-action and recovery failures because the proposed service must prevent those outcomes, not because a fictional probability calculation says they are frequent. Give each missing result an owner rather than colouring an untested control green.

Decision DR-1: illustrative record, not a real approval
FieldRecorded position
Scope and stateDocument assistant design version 1. HOLD real files, live users and tool writes. Only a future isolated synthetic-fixture evaluation is proposed.
Accountable ownerThe fictional service owner retains HOLD and must record an evidence and residual-risk review before permitting any broader activity. No approval has occurred.
EvidenceV1-V7: NOT RUN. Identity, retrieval, executor, release, recovery and operations owners must supply versioned configuration, inputs, expected/actual results and traces.
Blocking questionsHow are permissions refreshed? How is destination audience enforced? Can recovery obtain current ACLs and keys? Which exact import bundle is approved?
Failure ruleAny selected invariant failure or missing required evidence retains HOLD. A good average answer score cannot cancel a disclosure or unauthorised write.
Residual risk and reviewAuthorised misuse, wrong answers, administrator compromise and availability costs remain for accountable review. Revisit after identity, corpus, model, tool, update or recovery changes.

For each future run, record case ID, actor, ACL revision, document/index revision, model/runtime bundle, tool policy, approval identifier, observed side effects and evidence location. Separate NOT RUN, FAIL and PASS. Passing the selected cases would establish only their observed results, not certification or production readiness. NCSC's maintenance guidance also connects model and prompt changes with renewed evaluation.

Try changing one assumption on paper: allow an external contractor, introduce shared answer caching or restore an older index. Identify the new crossing, the invariant at risk and one check that could falsify the proposed control. If the drawing and decision stay unchanged, the model probably has not captured the new authority relationship.

For earlier context, read Unified's discussions of confidential information and hosted chatbots and AI on owned hardware. These links do not support a general safety, data-transfer-law or compliance conclusion from hardware ownership. This design needs its own evidence record.

Contribution and sources: AI-assisted piece credited to Mickarle Wagstaff-Irons - Micky Irons, full name Mickarle Sean Junior Wagstaff-Irons, Mickai's founder. Primary guidance was checked on 27 September 2026. Unified and Mickai share ownership; Mickai's AI-readiness information is separate commercial context. This fictional design is not a product audit or a legal/compliance conclusion.

Questions readers ask

Does running a model on owned hardware remove the need for threat modelling?
No. Users, files, retrieval services, tool credentials, update packages and backups still cross trust boundaries. Location changes the architecture, but does not establish who may read or change each resource.
Can a system prompt enforce document permissions?
No. The application must enforce permissions outside the model, before supplying context and before disclosing results. A prompt can describe expected behaviour, but retrieved text and model output cannot grant authority.
What should happen when a proposed control has no test evidence?
Record the evidence as NOT RUN or missing, identify its owner and keep the relevant release decision on HOLD. A completed threat table is a planning artifact, not proof that its controls work.
MICKAI®

Published by Mickai LTD.

Unified covers the field broadly and treats Mickai as one example within it. About the journal and the team.

Mickarle Wagstaff-Irons - Micky Irons, full name Mickarle Sean Junior Wagstaff-Irons. Article author. Biography and related work.

Keep reading