At a House subcommittee discussion about using large language models to personalize legislative information, a witness warned that current AI systems can produce inaccurate outputs and argued that human oversight and system design are essential to preserve accuracy and trust. Moderator R opened the exchange by asking how to leverage LLMs to customize legislative material while ensuring accuracy and public confidence.
The witness (transcript label: SPEAKER 6) said accuracy has multiple meanings—exact wording versus outcome—and cited a Stanford study, saying that even top-tier commercial legal AI "hallucinated 17 to 33% of queries, even when it was given the data." The witness recommended human-centered design features such as mandatory user notifications that surface items the model could not compute and checkpoints that require manual confirmation before action: "You're not allowed to move on this journey until you have confirmed A, B, C, D," the witness said, describing an approach that puts factual accountability on users and system designers.
The witness also contrasted fully generative systems with citation-forward tools, naming Perplexity as an example that attaches sources to outputs so users can verify them. In response, a questioner (transcript label: SPEAKER 1) asked whether reliance on LLMs would ever remove the need for human review; the witness replied that it depends on engineering, and reiterated that traceable sourcing and enforced UX guardrails can reduce—but not eliminate—the need for human judgment. The exchange ended with the questioner beginning an analogy to writing standards.