Practice·July 29, 2026 · 4 min read

Low confidence is an instruction, not a hedge

Our classifier can say it doesn't know. That single field changes what happens to an item next — and it's a habit worth borrowing the day you're reading a draft this product, or anyone else's AI, handed you.

By Connor Griggs — Regulatory & Quality Strategist

The model that reads a raw FDA document for this product is not asked whether the item matters. It is asked something narrower and more checkable: what changed, who it touches, and how sure it is of its own reading. That third answer — confidence — is not decoration on the other two. It is an instruction, and it is worth understanding exactly what it triggers, because the same habit applies the next time any AI tool hands you a draft and you have to decide how hard to check it.

What “low” is actually told to mean

The system prompt does not leave confidence to the model’s judgment about how confident sounds. It defines the trigger: use low when the source text is thin — a bare title, a one-line abstract — when the document is ambiguous about scope, or when the model is unsure the document is even device-related. And it states the consequence in the same sentence, so the model knows what the flag is for: a low here routes the item to closer human attention, which is exactly what it is for. Confidence is not the model hedging its own prose. It is the model telling the review queue where to look harder.

The tag list gets the same discipline

The same prompt applies an identical rule to structured tagging: pull values only from controlled vocabularies, tag only what the document actually implicates, and treat an empty list as a correct answer rather than a gap to fill — a recall of an orthopedic screw has no therapeutic area worth tagging, and forcing one in would be a fabrication wearing a field name. The instruction is explicit that over-tagging is worse than under-tagging, because a tag nobody can trust makes every filter built on it worthless. And when the correct value genuinely isn’t in the controlled list, the model is told to supply its best answer anyway rather than silently drop it — it gets flagged for review, not discarded, so an ontology gap becomes visible instead of quietly absorbed into the wrong bucket.

A model that never admits uncertainty hasn’t gotten better at reading. It has gotten better at sounding sure.

The habit this is built to produce

None of this changes what the model is allowed to decide — it still cannot assign urgency or write a recommended action; that authority stays with the human reviewer, full stop. What confidence changes is where that reviewer’s attention goes first. A high-confidence draft on a routine recall is a fast read: check it against the source, confirm the summary matches, move on. A low-confidence draft is a different task entirely — it is a signal to stop trusting the paraphrase and read the primary document yourself before deciding what, if anything, the item says. Treating both drafts the same defeats the entire point of asking the model to disclose the difference.

That habit travels well beyond this product. Any AI tool that hands you a summary, whether it is ours or a competitor’s, is making an implicit confidence claim every time it writes a clean, declarative sentence. The useful question is not whether the sentence reads well — fluent and wrong is the failure mode that matters here. It is whether the tool will tell you, in its own output, when it isn’t sure. If it never does, that is not evidence it is always right.

What this isn’t

None of this is a claim that a high-confidence draft is safe to act on unread, or that a low-confidence one is wrong — confidence describes how hard the item is to read, not whether it is material to you. Every drafted summary in this product still passes a human reviewer before anyone sees it, confidence field or not. This is regulatory intelligence about how the drafting works, never regulatory advice about what a given item means for your device. See our editorial standards for the rest of what the pipeline does and doesn’t let a model decide.

Regulatory intelligence, not regulatory advice. This post describes method and published FDA records as of its date; decisions about a specific device belong with your regulatory professional.

Method
A Class I device, a Class I recall
Practice
21 CFR 820 didn't move. Its contents did.
Method
The product code that doesn't exist yet
Practice
The classification posts. The 483 behind it doesn't.
Method
A recall has three dates, and the pipeline had to pick one
Practice
The count is real. The rate is not.
Method
The firm on the record is not the firm on the box
Method
The same company, spelled three ways
Practice
A device that was never a medical device
Method
FDA's warning letters, addressed by column number
Practice
Your regulation has a decimal. FDA's watch doesn't.
Method
Three letters is too short to search for
Practice
Most warning letters never close
Method
The guidance that skipped the draft
Practice
Ongoing, as of when?
Method
The least interesting fact in a 510(k)
Practice
No recall arrives with a product code attached
Practice
The deadline that doesn't email you
Method
The warning letter has two dates
Method
How to monitor FDA without drowning
Practice
Your predicate was recalled. Now what?
Method
Why no item reaches you without a human