Audit the System Prompt
Accountable for the Work, Blind to the Instructions: Why We Need AI System Prompt Auditing for Large Language Model Applications
I’m stoked to be a co-author of the new paper “AISPA: User-Centric System Prompt Auditing for Large Language Model Applications”.
AISPA introduces a framework for auditing system prompts, the developer-written instructions that help configure commercial AI products. My co-authors have explained the framework and its findings. I want to focus on one implication: what hidden vendor instructions mean for lawyers and other professionals who remain answerable for AI-assisted work. And yes, if you are depending on this technology for high value, mission critical, sensitive or other important work, this is for YOU too.
Your prompt may not be the controlling instruction
When a professional uses AI through a vendor’s product, the carefully written user prompt is only one part of the effective instruction stack.
Depending on the product, that stack may also include foundation-model policies, vendor system or developer instructions, tool and agent instructions, enterprise configurations, retrieved material, memory, and output filters, among other things. Some of those instructions may qualify or override what the professional asks.
They can affect what the system prioritizes, omits, refuses, assumes, or does before requesting confirmation.
Call this vendor-side instructional intermediation.
This is not prompt injection. No attacker needs to be involved. The vendor’s instructions may be legitimate, beneficial, and carefully tested. But they still introduce an instructional layer between the professional and the model, and it is one that professional customers often cannot fully inspect, modify, or track as it changes.

System prompts are not decorative boilerplate. They are part of the product’s operating methodology.
Harvey’s engineering team, for example, has written candidly about maintaining a growing system prompt across multiple development teams. Harvey describes the challenge as “merging English”: one instruction can conflict with another and cause regressions that must be caught through evaluations.
That candor is valuable. It confirms that the instruction layer is operationally consequential (engineered, changed, and tested inside the vendor) while customers generally receive much less visibility into it.
The accountability asymmetry
The professional-responsibility problem can be stated simply:
Professionals remain accountable for the work they deliver, while material instructions shaping that work may remain invisible, vendor-controlled, and subject to change.
For lawyers, this is especially concrete.
ABA Formal Opinion 512 says lawyers using generative AI must reasonably understand the specific tool’s capabilities and limitations, exercise professional judgment, supervise appropriately, and remain fully responsible for work performed for the client.
The opinion does not discuss system prompts. But it was one of the first things I thought of when contemplating issues and options for reasonable reliance.
That does not establish that every lawyer is legally entitled to receive every raw vendor prompt. Nor is raw-prompt disclosure indispensable in every deployment. Meaningful disclosures, behavioral evaluations, contractual assurances, and independent audits may provide the information necessary for responsible use.
But today, the gap remains significant.
A lawyer may be expected to understand and supervise a product whose behavior is materially affected by instructions the lawyer has not seen and may not know have changed. There is not yet a common assurance standard ensuring that the lawyer receives an adequate substitute for direct access to the system prompts that are in fact part of their workflow and impact their work.
Professionals therefore have a genuine information interest in the instruction layers governing their tools. I believe that interest should mature into a recognized professional information right: meaningful information about material instructions, priorities, limitations, and changes, but not necessarily a right to directly inspect, publish, or edit every line of a vendor’s prompt.
Hidden instructions can encode professional methodology
In a legal product, the instruction layer may affect:
Which authorities or sources receive priority.
Whether the system takes a conservative or aggressive interpretive posture.
How uncertainty and conflicting authority are presented.
When the system searches, cites, refuses, redirects, or asks for confirmation.
How legal, reputational, safety, and commercial considerations are balanced.
Whether firm instructions or client objectives can be overridden.
And other things because the hidden instruction can impact whatever it instructs.
None of these choices is necessarily improper. Every professional system requires design decisions.
But collectively they may amount to an undeclared professional methodology. The lawyer should receive enough information to decide whether that methodology is appropriate and acceptable for the matter, the client, and the lawyer’s own obligations.
Customization is not governance
Vendors correctly point out that their products may be customizable: customers can add prompts, playbooks, and workflows, among other things.
Those capabilities are useful. But customization is not the same as governance.
If customer instructions may be impacted by or indeed subordinate to undisclosed vendor rules, a firm can personalize its workflow without knowing which instructions can override that personalization or how conflicts will be resolved.
Put more directly:
You can decorate the workflow without controlling its constitution.
Meaningful accountable governance requires visibility into and the ability to change or knowledgeably adopt the instruction hierarchy, the limits of customer control, and the product’s behavior when instructions conflict.
What AISPA found and what it did not find
AISPA evaluates system prompts across eight user-protection dimensions.

Keep in mind these system prompts arise from much broader contexts than just professional vendor systems. Nonetheless, applying the framework to publicly available prompts associated with 88 commercial AI products, the audit found that nearly every prompt contained at least one protective instruction. But fewer than one-quarter covered all eight dimensions, and roughly 40% contained at least one instruction classified as working against user interests.
User agency was the most frequently violated dimension, including instructions directing assistants or agents to proceed without first obtaining confirmation. Protective and problematic instructions sometimes coexisted within the same prompt. This theme is explored much more deeply in my recent report on how to evaluate AI agents against the duty of loyalty, a fiduciary agency law concept.
But for purposes of this post, there are two scope factors that matter.
First, this was a cross-category study covering chatbots, coding assistants, autonomous agents, and other applications. Its corpus came from public repositories containing leaked or community-disclosed prompts. It was not a verified census of current production systems, and it did not find that any particular percentage of legal-AI products or other professional-grade services contain problematic instructions.
The application to professional technology is my extension of the paper’s findings.
Second, AISPA audits what prompts say but not whether deployed systems consistently obey them. The system prompt, consequential as it is, is just one component of a much larger processing pipeline. Depending on the product, that pipeline may rewrite or expand a user prompt, route it among models, inject retrieved documents or memory, and truncate context, among many other things. After generation, outputs may pass through moderation, citation checking, redaction, formatting, or other post-processing before reaching the professional. These pre-processing and post-processing stages can materially affect both what the model receives and what the professional ultimately sees, and customers may have limited (or precisely zero) visibility into any of this.
Even so, I think auditing the system prompt is an appropriate starting point because it is a consequential, legible (literally, it is written in words we can read), and auditable control artifact. As I have discussed with co-authors, assurance for professional use will eventually need to encompass the material components of the greater effective context and wider processing pipeline. That is beyond AISPA’s present scope, and squarely within the conversation it should start.
That is why mature assurance must combine instruction review with behavioral evaluation, versioning, provenance, and operational evidence.
The answer is assurance, not indiscriminate publication
The answer is not necessarily to require every vendor to publish every prompt.
Vendors can have legitimate interests in protecting intellectual property and operational details. Raw prompts may also be lengthy, difficult to interpret, and incomplete as descriptions of behavior.
Security, however, cannot carry the entire argument for opacity. OWASP’s guidance says system prompts should not be treated as secrets or relied upon as security controls. Critical permissions and safeguards belong in external, auditable mechanisms.
The better model is tiered assurance, perhaps including:
Meaningful disclosure. A System Prompt Card should describe the product’s intended role, source priorities, major restrictions, professional assumptions, and the instructions customers can or cannot override.
Version and change information. Customers should be able to identify material model, prompt, tool, and application versions and receive notice of consequential changes.
Behavioral evidence. Vendors should demonstrate through representative and adversarial evaluations that the product follows its stated safeguards. Buyer-side custom evals can in fact depend on this type of clear and predictable foundation.
Controlled access. Independent auditors, regulators, and appropriate professional customers should be able to inspect the complete relevant instruction stack under confidentiality and other protections.
Contractual accountability. Procurement agreements should address material changes, audit rights, version retention, conflicting instructions, incident reporting, and cooperation in litigation or regulatory review.
AISPA supplies a starting methodology for the independent-audit component.
The analogy to financial auditing is institutional, not technical. Organizations do not disclose every proprietary record publicly. Independent professionals inspect confidential evidence under defined standards and issue assurance reports.
System-prompt assurance does not yet have financial auditing’s mature standards, independence requirements, or enforcement infrastructure. That is a reason to build those institutions, not a reason to accept permanent opacity.
There is also a limited procurement precedent. In a particular US federal context, Executive Order 14319 and OMB Memorandum M-26-04 contemplate disclosure concerning system-level prompts, evaluations, enterprise controls, and third-party modifications without requiring model weights.
That policy provides a very useful example: prompt-related information can be treated as a material procurement specification, and controlled disclosure can coexist with protection for sensitive technology.
Questions professional buyers should ask
A law firm, legal department, or other professional organization procuring AI should ask:
What instruction layers operate before, with, or above our instructions?
Which vendor instructions can qualify or override ours, and how are conflicts handled?
What professional role, source priorities, risk posture, refusal rules, and tool-use policies are encoded?
How are material changes tested, versioned, and communicated?
Can we identify the effective model, prompt, tool, retrieval, and application configuration associated with significant work?
What behavioral evaluations demonstrate that the product follows its stated safeguards?
Will an independent auditor receive confidential access to the complete relevant instruction stack?
What contractual rights apply if undisclosed changes or conflicting instructions affect professional work?
None of these questions requires public disclosure of a vendor’s complete proprietary prompt. Each asks for information or assurance proportionate to the accountability borne by the professional customer.
Beyond law
I have framed this through legal practice because that is where the accountability structure is clearest to me. But this is not confined to law.
It can arise wherever physicians, accountants, engineers, financial professionals, government officials, and others remain answerable for consequential work performed with AI assistance.
The governing duties differ across professions. The structural question is shared:
How can a professional responsibly adopt, supervise, and stand behind an AI-assisted process when material parts of that process are controlled by an outside vendor and unavailable for meaningful evaluation?
And as I mentioned at the top of this post, this is not limited to professionals. Rather, it should concern any organization or person who depends upon the outputs of these systems precisely because each output is highly sensitive to the exact inputs. System prompts are material inputs that impact the outputs you depend on.
As AI becomes embedded in all work, system-prompt assurance may become a missing bridge between professional responsibility and vendor accountability.
AISPA gives us a framework with which to begin. Building the institutions around it, and eventually extending assurance from the prompt to the whole pipeline, is the work ahead.
Read the paper: AISPA on arXiv.

