A checklist for the whole category
Fourteen questions to ask any legal AI vendor.
We make one of these products. That is the first thing to know about this page, because it should change how you read it.
We wrote it because we would rather compete on whether we answer these questions well than on whose marketing sounds more confident. Our own answers are written down beside each one, so you can hold us to them exactly as you hold every other vendor to theirs.
It names no other product. We do not run them, so we are not in a position to tell you what they do — and anything we wrote about a competitor today would start going out of date the next time they ship. A question does not have that problem. It cannot be unfair to anybody, it does not expire, and the answer comes from the person who actually knows.
Every question below is one we think separates a system that will hold up on a matter from one that will not. None of them is about how good the output is. All of them are about what the system actually does — work performed, artifacts produced, checks run — because that is the part a firm can verify instead of taking on trust.
How to use it
Ask for a demonstration, not an answer
Every question here is answerable by showing you something on a screen. That is the property we selected for. A vendor who can only describe the answer is telling you about a roadmap; a vendor who can show you is telling you about a product. The distinction is worth being blunt about, because both sound identical in a meeting.
- Ask it about a run, not about the system. "Show me a matter where the checking flagged something" beats "do you check citations". The first has an answer that exists or doesn't.
- Ask what the output looked like when it went badly. Every one of these systems has a worst case. A vendor who will show you theirs is worth more than one who won't, and you learn more from it than from the demo they rehearsed.
- Take the same list to everyone, including us. The value is in the comparison, and it only works if the questions don't change between rooms.
- Where an answer is "no", that is often fine. Several of these describe deliberate design trades rather than defects. What matters is that the vendor knows which one they made and will say so.
Each HiCourt answer says what the system does today, and marks anything not built yet as “not yet”. Nothing marked that way is available until it is.
The fourteen
Grouped into the four things worth knowing
How the work gets done · what happens when it is wrong · what you can see afterwards · what happens to your firm's material.
A · How the work gets done
-
Can I hand it the matter, or do I have to break the matter into questions myself?
Why it separates products. There is a real difference between a system you brief and a system you interrogate. Both are useful. But if the lawyer has to decompose the work before the tool can start, the decomposition is the expensive part and it has not been done for you. Ask what happens to the documents, too: whether the relevant material arrives automatically, and whether the lawyer can force a specific document to stay in front of the system regardless of length.
HiCourt Hand it the matter. Pick a task from a menu of everyday legal work (research, assessments, summaries, chronologies, comparing two versions, proofreading a filing, drafting, preparing for a hearing and more), say what you need in your own words and drop in the files. Each file is read and indexed before a run can use it, and for research and analysis the firm’s permitted past work is recalled automatically. A document you tick goes into the run; if it is too long to include whole, the run says what was left out. A lawyer reviews, verifies and signs the work.
-
Does it look for what hurts my position, or only for what supports it?
Why it separates products. A system tends to answer a question the way it was asked. Frame the matter as why we win and you get a brief for that position, and the authority that cuts against it — the thing you most need to know before the other side tells you — is the thing least likely to appear. Ask to see a run where the output raised a case against the client’s position without being asked to.
HiCourt Yes. The work is written to look for what cuts against your position as well as what supports it, and the cases it cites are then tested against the claims they are cited for, whichever side they help.
-
When it draws on the firm’s past work, does it tell me which of it was ever reviewed?
Why it separates products. A firm’s past work is the most useful input it has and the most dangerous one. A memo nobody reviewed, or one a partner sent back, looks exactly like one that was signed, and a system that quietly reuses it is repeating whatever was wrong with it. Ask whether prior work arrives with its status attached, and whether it is ever treated as authority in place of checking the law again.
HiCourt Yes. For research and analysis, permitted past work is recalled with a label: reviewed or not, clean or with known problems, and where it came from. It informs the work; it never replaces checking the law afresh.
B · What happens when it is wrong
-
Is the thing that checks the work the same thing that wrote it?
Why it separates products. A system reviewing its own output is doing something genuinely useful, and it is not the same thing as an independent check. The distinction to listen for is whether the checker had sight of the draft and whether it is the same configuration that produced it. Self-review catches incoherence. It is structurally weaker at catching a mistake the writer was confident about, because the confidence travels with it.
HiCourt No. The cases the work cites are looked up in the published case law, then a separate checking stage that wrote none of the work reads the opinion itself and tests each case against the claim it is cited for. Uncertain and missing results stay visible rather than being smoothed over.
-
Is a citation's existence settled by a lookup or by a model? And is a model ever asked to confirm a citation rather than reproduce it?
Why it separates products. Two separate things hide in one question. First: no language model queries a reporter, so asking models whether a case exists samples recollection — the same faculty that produced the bad citation. Existence has to be settled against a source. Second, and easier to miss: the shape of the prompt matters. "Does this case exist?" invites agreement, and an agreeable model says yes. "Reproduce this case's details" does not. Ask to see the actual prompt.
HiCourt A lookup. Each case citation is checked against the published case law, and a case that cannot be found is flagged, never assumed; no AI is asked whether a case is real. The lookup covers case citations in reporter form. Not yet: quotation and pinpoint checks, whether a case is still good law, and statutes, regulations, rules and treatises.
-
When something comes back wrong, does it get fixed, or does it get a label?
Why it separates products. This is the question we would ask first if we were buying. Flagging a problem and repairing it are different products wearing similar words, and the difference is enormous in practice: a flag hands the work back to your lawyer, a repair does not. Neither is wrong — but a system that flags is a review tool and a system that repairs is a drafting tool, and they should not cost the same or be evaluated the same way. A pilot must assess the difference on the finished work.
HiCourt Fixed. Anything flagged goes back to be corrected or withdrawn, and the fix is checked again. The exception is by design: a citation check of a draft you paste in returns a table and never rewrites your draft. Findings on the finished document stay visible for your lawyer and never trigger a quiet rewrite.
-
What ends the checking, and does the output tell me which ending it hit?
Why it separates products. Anything that loops has to stop, and there are only a few ways: it converges, it hits a budget, or it hits a fixed number of rounds. All three are legitimate. What matters is whether you can tell them apart afterwards, because "we checked until we found nothing" and "we ran out" are completely different statements about a document and look identical if the output does not distinguish them.
HiCourt The repair loop is bounded, and every run records which way it ended. A finding on the finished document stays with your lawyer and never triggers another rewrite. A citation check of a pasted draft has no repair loop.
-
Does verification run on the draft, or on the document I actually receive?
Why it separates products. This one is easy to miss and hard to unsee. If a system verifies its working material and then generates the final document from it, the final document has not been verified — because generating it is another inference, and an inference can introduce a citation that was never in its input or state a holding more strongly than the source supports. Grounding the input is not the same as checking the output. Ask which artifact the last check ran against.
HiCourt The document you receive. After it is written, its cited claims are checked again, and a separate check asks whether it does what you asked. Every claim carries a status, including “not checked”, and those statuses travel with the download. Findings never trigger a rewrite.
C · What you can see afterwards
-
Where the law is genuinely unsettled, does the document say so?
Why it separates products. Some questions have more than one defensible reading. The question is whether the product shows you that or resolves it quietly. Resolving it quietly produces a cleaner document, which is a real benefit and also the entire risk: a memo that reads as settled on a question that is genuinely open is worse than one that says so, and you cannot tell from the prose which you are holding.
HiCourt Points the work could not settle, and flags still open, are carried into the document and marked rather than smoothed over. Whether a point of law is truly settled is still your lawyer’s call.
-
Can I map a sentence in the output back to the authority it rests on?
Why it separates products. Provenance is either carried through every stage or it is reconstructed at the end, and reconstruction is guesswork dressed as a citation. There is also a practical consequence worth asking about: if provenance survives, re-checking after a summarising step can be narrow — only the new claims. If it does not, the vendor has to choose between re-running everything and re-running nothing, and the second is cheaper.
HiCourt A sentence that cites a case is checked against that case in the finished document, and every sentence carries a status, so a claim with nothing behind it is marked as open rather than dressed up. Not yet: a source carried for every sentence through every step of the work.
-
What settled this citation, and as of when?
Why it separates products. A green tick is not a fact unless it names its source and its date. Law changes; a citation that was good in March can be bad in November, and a status with no as-of date silently claims to be timeless. Ask what happens when the same document is re-run next quarter and something comes back differently — a system that treats that as a contradiction has misunderstood its own output.
HiCourt Each run records separately what the lookup found and what the checking concluded, and when. Missing source or date evidence stays missing. A run’s date is when the work was checked, not an opinion that the case is still good law, which is not checked yet.
-
Six months from now, what can the firm reconstruct about why a citation was relied on?
Why it separates products. The moment this matters is the moment nobody wants to be in: someone is asking why a filing said what it said, and the person who ran the work has left. What you want to exist is a per-run record — what ran, against what, what came back, what was flagged and what was done about it. A chat history is not that, and neither is a document with citations in it.
HiCourt Each run keeps a protected record of the request, the files it used, what came back, what was flagged and what was done about it, readable later only by people still allowed to see it. Every finished run is also filed into the firm’s memory, and a lawyer’s signed review makes a new version.
D · What happens to your firm's material
-
Is the ethical wall enforced before scoring, after scoring, or in the interface? And do derived documents inherit it?
Why it separates products. We would push hardest on this one, in anyone's product including ours. A retrieval index does not respect an ethical wall by nature — ask it an ordinary question and it returns whatever is nearest, which may be from a matter the person asking is screened off from. They have done nothing wrong, which is exactly why the control cannot live in people's judgement. The follow-up that catches the most: does a past memo quoting a screened matter inherit the screen? Derived work looks like new work and is actually a copy.
HiCourt Before scoring. Permissions are resolved first, so a run only searches what the person asking may see, and a result from anywhere else is rejected before it is opened. Past work built from a matter carries that matter’s permissions with it, and a sealed matter’s documents are kept out of other matters’ searches.
-
Where does our material rest, is that decided per matter or once for the whole firm, and what leaves during a run?
Why it separates products. Two questions that firms hear as one, and they have different answers. Where files rest is about the standing collection: who holds the keys, what a breach reaches, what there is to produce if someone asks. What travels during a run is about the moment work happens, and storage location does not bound it. Ask both separately, and be suspicious of any answer that lets the first quietly stand in for the second.
HiCourt Once for the whole firm: each firm runs in its own dedicated hosted environment, with its own server, database, storage and encryption keys. During a run, drafting and checking use commercial AI, so the work’s text is sent out for processing; nothing is sent that the person asking may not see. Not yet: keeping the library on a machine in your office or on each person’s computer.
Deliberately not on the list
Four questions that sound rigorous and are not
These get asked in evaluations constantly. We think each one is worse than it looks, and it is only fair to say which of them we have an interest in you dropping.
- skip it "What's your accuracy rate?" A number with no denominator is not a measurement. On what task set, graded by whom, against what baseline, and would a competitor score on the same rubric with the same graders? Without those, the figure measures the confidence of the person quoting it. Better question: who graded it, and were they blind to which system produced the answer?
- skip it "Is it hallucination-free?" Nobody can answer this truthfully in the affirmative, so the answers sort into people who say yes and people who explain. The explanation is the useful part. Better question: question 5, and then question 6 — what settles existence, and what happens to the ones that come back wrong.
- skip it "How many hours will this save us?" It is a claim about your firm made by someone who has not seen your matters, your staffing or your review habits. It is also the easiest number to produce and the hardest to check. Better question: run it on a matter you have already closed, and compare it to what you actually did.
- read this one knowing our interest "How many firms use it?" Adoption is evidence that a product sells, which is genuinely worth something and is not evidence that it is right. Plenty of widely-used software is wrong in ways its users have not audited. We have an obvious interest in you discounting this question, because we have no customers — so weigh it accordingly. We still think a firm that runs its own pilot on a closed matter learns more than one that counts logos.
And the questions to ask us hardest
A checklist written by a vendor is worth reading only if it also points at the vendor. So here is where we are weakest, stated before you find it.
- "Does this exist yet?" Yes, privately, for one test firm, across a menu of everyday legal tasks from research memos to chronologies. Access is by invitation. Nothing on this page means it has been qualified for your jurisdiction or your practice.
- "Which corpus and which citator?" Case citations are looked up in a published case-law database, not a licensed citator, and its coverage still has to be confirmed for each firm’s jurisdictions. Whether a case is still good law, and quotation and pinpoint checks, are not available yet.
- "So our client's text leaves the building?" During a run, yes: drafting and checking use commercial AI, so the work’s text is sent out for processing. Confirming the terms that govern it, and setting them out for each firm before it sends client material, is a condition of release.
- "What does all this checking cost, in time and money?" Each run records its own usage. How long a run takes end to end, what it saves and how its finished work compares have not been measured yet.
- "You wrote the questions. Aren't they shaped to your answers?" Fair, and worth saying plainly: we chose questions we think we answer well. The defence is not that we were neutral — it is that every one of them is checkable, so a vendor who answers better than us will show you, and you will have learned something either way.
Tell us which questions are wrong
This list is a draft of what we think matters, and the people who know are the ones who have run these evaluations. What did we miss, and which of these would you throw out? That is genuinely the most useful thing you could send us.