Mode
Text Size
Log in / Sign up

Safety user interface bundles improve verification intentions in older adults using generative AISafety prompts boost older adults' AI verification

AI-generated summary of the cited source, checked by automated accuracy review. How we work

Key Takeaway
Consider using safety UI bundles to improve verification intentions and trust calibration in older adults using AI.

The study investigated the impact of a safety user interface bundle on older adults interacting with generative AI tools. The intervention included generic source-label cues and an uncertainty and verification nudge. Researchers measured several outcomes including verification intention, trust calibration, perceived trustworthiness, and cognitive load to determine how these design elements influenced user behavior.

The results indicated that the safety interface bundle led to a significant increase in verification intentions compared to a baseline chat interface. Furthermore, participants showed improved trust calibration, which is associated with a lower risk of overreliance on AI outputs. While perceived trustworthiness was slightly lower in the intervention group, comprehension and cognitive load remained stable across both conditions.

A primary limitation noted by the authors is that because the interventions were delivered as a bundle of cues rather than individual elements, the specific impact of any single component cannot be isolated. Additionally, the study's findings are based on a vignette survey among a specific demographic. These results suggest that interface-level strategies may help older adults navigate AI tools more critically, but further research is needed to isolate effective design components.

A new study tested whether a simple interface change could help older adults be more careful with AI chatbots. The researchers created a "safety UI bundle" that included generic source labels and a nudge reminding users to verify information. They compared this to a standard chat interface in a randomized experiment with 200 older Chinese adults aged 60 and older, recruited from community sites, clinic waiting areas, and WeChat groups.

People using the safety interface reported a higher intention to verify the AI's answers (average 4.72 vs. 4.41 on a 7-point scale). They also showed better trust calibration, meaning they were less likely to over-trust the AI. However, the change did not significantly affect how much they relied on the AI, and their actual verification behavior (clicking to expand source info) was higher but not statistically significant.

The safety interface did not hurt comprehension, usability, or cognitive load. Perceived trustworthiness was slightly lower, but that may be a positive sign of healthy skepticism. No safety concerns were reported, but the study did not track adverse events.

This was a small, short-term experiment using screenshots, not a real-world test. The intervention was tested as a bundle, so we can't say which specific element worked. Also, the results may not apply to other age groups or cultures. Still, the findings suggest that simple design tweaks could help older adults use AI more safely.

What this means for you:
Simple interface cues can encourage older adults to verify AI answers, but more research is needed.

Common questions

What did the study test?

The study tested a safety user interface bundle with generic source labels and a verification nudge, compared to a standard chat interface. It involved 200 older Chinese adults aged 60 and older. The goal was to see if these cues could encourage people to verify AI-generated information.

Did the safety interface reduce trust in AI?

Perceived trustworthiness was slightly lower in the safety interface group (average 5.20 vs. 5.39 on a 7-point scale). This was a modest change and might be a positive sign, as it suggests people were less likely to over-trust the AI. Trust calibration improved, meaning better judgment about when to trust.

Is this study proof that the interface works?

No. The study was a randomized experiment, but it used screenshots and tested the interface as a bundle. The authors say they cannot claim any single cue caused the effects. Also, the study only included older Chinese adults, so results may not apply to other groups.

Study Details

Study typeRct
Sample sizen = 1
EvidenceLevel 2
Follow-up720.0 mo
PublishedAug 2026
View Original Abstract ↓
BACKGROUND: Generative AI chat systems are increasingly used for everyday information seeking, but plausible errors and omissions can mislead users when outputs are accepted without scrutiny. Interface-level safety cues may help users calibrate trust and engage in verification; yet, evidence in older Chinese adults remains limited. OBJECTIVE: This study aimed to test whether adding a safety user interface (UI) bundle to a generative AI chat interface increases verification intention among older Chinese adults and to examine selected secondary outcomes, including reliance intention, trust calibration, perceived trustworthiness, comprehension, usability/readability, cognitive load, and a behavioral proxy of verification. METHODS: We conducted a cross-sectional survey with an embedded randomized UI vignette experiment between May 22, 2025, and September 3, 2025. Chinese adults aged ≥60 years were recruited through community sites, outpatient clinic waiting areas, and WeChat (Tencent Holdings Ltd) groups, and randomized 1:1 to view screenshots of a baseline chat UI or a safety UI bundle containing generic source-label cues, and an uncertainty and verification nudge. Each participant completed 2 scenarios (service/travel decision and general well-being related to sleep/fatigue), followed by measures of verification intention (primary), reliance intention, trust calibration index, comprehension (0-8), perceived trustworthiness, usability/readability, cognitive load (0-10), manipulation checks, and a behavioral proxy (expanding optional "source information"). Analyses used intention-to-treat regression models with covariate adjustment. RESULTS: Of 214 consenting respondents who started the survey, 200 were included in the analysis (100 per arm). The safety UI bundle increased verification intention (mean 4.72, SD 0.63 vs 4.41, SD 0.59 on a 7-point scale; adjusted β=0.293, 95% CI 0.128-0.457; P<.001). Reliance intention did not increase (mean 4.97, SD 0.54 vs 5.03, SD 0.58; adjusted β=-0.105, 95% CI -0.239 to 0.029; P=.13). Trust calibration improved (trust calibration index: mean -0.29, SD 1.43 vs 0.29, SD 1.43; adjusted β=-0.567, 95% CI -1.005 to -0.129; P=.01). Expansion of optional source information was numerically higher, although the adjusted CI included the null (42% vs 27%; adjusted odds ratio [OR]=1.76, 95% CI 0.95-3.27; P=.07). Comprehension remained high and similar across arms (mean 6.33, SD 1.14 vs 6.32, SD 1.08; adjusted β=-0.132, 95% CI -0.428 to 0.163; P=.38). Perceived trustworthiness was modestly lower in the Safety UI arm (mean 5.20, SD 0.61 vs 5.39, SD 0.66; adjusted β=-0.199, 95% CI -0.382 to -0.016; P=.03). Usability/readability was unchanged, and cognitive load did not increase. Manipulation checks indicated higher cue recognition in the Safety UI arm. CONCLUSIONS: In a randomized static-vignette survey of older Chinese adults, a brief safety UI bundle was associated with higher verification intention and a trust calibration index consistent with lower overreliance risk, without detectable reductions in comprehension or usability/readability. Because the intervention was tested as a bundle using screenshots and generic source labels, findings should be interpreted as evidence for a practical interface-level strategy rather than proof that any single cue caused the observed effects.
Free Newsletter

Clinical research that matters. Delivered to your inbox.

Join thousands of clinicians and researchers. No spam, unsubscribe anytime.