AI Washing In Legal Tech: 7 Vendor Questions To Ask

AI Washing In Legal Tech
AI Washing In Legal Tech

AI commonly appears across legal technology proposals, demonstrations and product roadmaps in current times. Yet the label can describe very different things from a focused classification model to a third-party generative AI service or a rules-based feature presented as intelligent automation, with capability that remains largely aspirational. For General Counsel and legal leaders, this creates a legal technology evaluation problem.

A polished demonstration may show that a feature can produce an impressive output. It does not establish how reliably the feature performs on your work, what happens to your data, or who remains accountable when the output is wrong. Effective legal ops vendor selection requires a move from labels to evidence.

These seven questions help buyers test whether an AI claim represents a governed capability that is fit for a defined purpose.

1. What Does the AI Actually Do?

Ask the vendor to describe the capability in operational terms. What input does it receive? What task does it perform? What output does it create? Which model or service sits behind it? Is the capability in production, in limited release or on the roadmap?

The answer should distinguish the AI component from surrounding workflow automation. That distinction matters because automation can still deliver value, but buyers need to know what they are assessing and paying for. A vendor that cannot explain the use case without reverting to broad claims about efficiency has not provided enough information to make a sound decision.

2. What Evidence Supports the Performance Claim?

If a vendor claims greater accuracy, faster completion or improved productivity, ask for the basis of comparison, the test conditions and the results. A credible vendor should be able to discuss limitations as readily as strengths.

Testing should also reflect the work the legal team intends to perform. Results produced on clean demonstration data may not translate well to the imperfect organisational reality.

The US Federal Trade Commission and regulators in other jurisdictions have taken enforcement action against deceptive AI claims, reinforcing the straightforward business principle that marketing claims should be supported by evidence.

3. Where Does Our Data Go?

Map the full data path. What information is sent to the AI system? Where is it processed and stored? What is retained, logged or used to train? Which sub processors are used and where are they located? What deletion, residency and security arrangements apply?

These questions are particularly important for in-house legal teams use of AI because matter data may contain personal information, commercially sensitive material or privileged communications. General assurances about enterprise security are not a substitute for a clear data-flow description and contract terms that reflect it.

4. What Are the Known Limitations and Failure Modes?

Ask when the feature should not be used and how it can fail. Relevant answers may include unsupported document types, weak performance on specific inputs, incorrect classifications, fabricated content, or reduced accuracy outside the tested context.

Australia’s current Guidance for AI Adoption calls for supply chain transparency regarding test methods, known limitations, risks, mitigations, data processes, and privacy and cybersecurity practices. This provides a useful benchmark for vendor due diligence even where the guidance is voluntary.

5. Where Does Human Judgement Sit?

Identify the decisions the system supports and the points at which a person reviews, approves, or overrides its output. The control should match the risk of the use case. A low-impact suggestion may require a lighter checkpoint than an output that informs legal advice, external communication, or a material risk decision.

Human review should be designed into the workflow, with responsibility assigned to a suitable role. It should not appear as a broad disclaimer after the product has shaped the decision. Legal accountability remains with the organisation and its people.

6. How Will the Capability Be Monitored and Changed?

AI services can change due to model updates, new data, altered prompts, configuration changes, or the replacement of an upstream provider. Ask what the vendor monitors, what performance information customers receive, and how material changes are communicated.

The NIST AI Risk Management Framework treats risk management as a lifecycle activity spanning design, deployment, use, testing and evaluation. Procurement should reflect the same principle. Contracts and governance processes should address incident notification, audit information, version changes, service levels, remediation and exit arrangements.

7. Can We Test the Claim in Our Operating Environment?

A buyer should be able to validate the feature against a defined business problem before scaling it. Set a narrow use case, representative scenarios, measurable acceptance criteria and a clear review period. Include negative tests and exceptions, not only the straightforward examples used in a demonstration.

Legal operations can then assess more than output quality. The evaluation should consider user effort, review time, workflow fit, traceability, permissions and the operational response when the system is uncertain or incorrect.

A Better Standard for Legal Technology Evaluation

AI washing in legal tech is best addressed through disciplined questions rather than scepticism alone. The goal is to establish whether a capability solves a defined problem, performs to an acceptable standard and can operate within the organisation’s data, risk and governance requirements.

For the General Counsel, this approach reduces the risk of buying a label rather than an outcome. For Legal operations, it creates a defensible record of why a tool was selected, how it will be controlled and what evidence will be used to judge its value.

Share

Share