Mode
Text Size
Log in / Sign up

ChatGPT-5.5 Instant shows higher accuracy and safety scores than DeepSeek-V3 for HPV inquiriesChatGPT shows higher accuracy than DeepSeek for HPV health questions

AI-generated summary of the cited source, checked by automated accuracy review. How we work

Key Takeaway
Note that ChatGPT-5.5 Instant outperforms DeepSeek-V3 in accuracy and safety for HPV inquiries, but both require oversight.

This methodological study evaluates the performance of two large language models, ChatGPT-5.5 Instant and DeepSeek-V3, in providing information regarding high-risk human papillomavirus (HPV) infection. The study assessed 90 AI-generated responses across dimensions including accuracy, safety, completeness, and guideline concordance.

ChatGPT-5.5 Instant outperformed DeepSeek-V3 in several metrics. Specifically, ChatGPT-5.5 Instant showed higher scores in accuracy (95% CI, 0.11 to 0.39; FDR-adjusted p = 0.010), safety (95% CI, 0.07 to 0.47; p = 0.010), and completeness (95% CI, 0.07 to 0.37; p = 0.020). The composite score for ChatGPT-5.5 Instant was also higher (95% CI, 0.04 to 0.31; p = 0.027). However, the primary outcome of responses without major error or potential harm (93.3% for ChatGPT-5.5 Instant vs. 84.4% for DeepSeek-V3) did not reach statistical significance (p = 0.344).

Limitations included exploratory question-level risks identified in specific scenarios such as HPV16/18 positivity, normal cytology, partner management, and pregnancy. The authors conclude that while ChatGPT-5.5 Instant performed better in several metrics, both models require guideline-based clinical oversight for patient interactions.

How this fits prior evidence

This study addresses a gap in the evaluation of AI tools for patient education regarding HPV. While prior evidence suggests that interactive digital assistants may improve short-term HPV vaccine booking and cervical cancer screening uptake, this study specifically evaluates the accuracy and safety of the underlying AI models used in such tools. It confirms that while ChatGPT-5.5 Instant shows higher accuracy and safety scores than DeepSeek-V3, both models still present risks in specific clinical scenarios like pregnancy or partner management.

When patients search for answers about high-risk human papillomavirus (HPV) infections, they often turn to AI tools. Because these questions involve serious health concerns, it is vital to know which tools provide the most reliable information. Researchers compared two popular AI models, ChatGPT-5.5 Instant and DeepSeek-V3, to see how they handled patient questions about HPV.

The study found that ChatGPT-5.5 Instant performed better than DeepSeek-V3 in several key areas. Specifically, ChatGPT scored higher in accuracy, safety, and completeness. While both models were tested on 450 records, ChatGPT provided more complete answers and showed higher safety scores in the comparison.

However, there are important limits to keep in mind. While ChatGPT performed better in several categories, the difference in the primary safety metric—responses without major errors or potential harm—was not statistically significant. Additionally, both models showed risks when answering specific questions about pregnancy, partner management, and certain test results. Because of these risks, experts say both AI tools still require a doctor's oversight.

What this means for you:
ChatGPT-5.5 Instant outperformed DeepSeek-V3 in accuracy and completeness for HPV-related health questions.

Common questions

Which AI model is safer for HPV questions?

ChatGPT-5.5 Instant showed higher safety scores than DeepSeek-V3 in this study. However, the difference in the primary safety metric—responses without major errors or potential harm—was not statistically significant. Both models still require clinical oversight from a doctor.

How accurate are these AI models for HPV information?

ChatGPT-5.5 Instant was found to be more accurate than DeepSeek-V3. While it performed better in accuracy and completeness, both models showed risks when answering specific questions about pregnancy and partner management.

Can I rely on AI for my HPV diagnosis?

No, you should not rely solely on AI for medical advice. While ChatGPT-5.5 Instant scored higher in accuracy and completeness than DeepSeek-V3, both models require guideline-based clinical oversight from a healthcare professional.

Study Details

Study typeGuideline
EvidenceLevel 5
PublishedJul 2026
View Original Abstract ↓
BackgroundAfter diagnosis of high-risk human papillomavirus (HPV) infection, patients often seek guidance on cancer risk, colposcopy, follow-up, treatment, partner management, pregnancy, and anxiety. Large language models (LLMs) are increasingly used for medical consultation, but their response quality and safety in this setting require evaluation.MethodsIn this single-center, expert-rated methodological study, 15 patient-style questions based on common outpatient consultations were submitted to ChatGPT-5.5 Instant and DeepSeek-V3. Each model generated three independent responses per question, yielding 90 artificial intelligence (AI)-generated responses. Five blinded gynecology experts independently evaluated all responses, producing 450 expert-rating records. Accuracy, safety, guideline concordance, completeness, and understandability were rated on a 1–5 Likert scale. Experts also assessed major error and potential harm. The primary outcome was the proportion of responses without major error or potential harm.ResultsChatGPT-5.5 Instant had higher response-level scores than DeepSeek-V3 for accuracy (mean difference, 0.25; 95% CI, 0.11–0.39; FDR-adjusted p = 0.010), safety (0.27; 0.07–0.47; p = 0.010), completeness (0.22; 0.07–0.37; p = 0.020), and composite score (0.18; 0.04–0.31; p = 0.027). The difference in guideline concordance did not remain significant after multiplicity correction (0.16; −0.01 to 0.33; FDR-adjusted p = 0.057), and understandability was similar. Responses without major error or potential harm occurred in 42/45 (93.3%; 95% CI, 81.7–98.6%) ChatGPT responses and 38/45 (84.4%; 70.5–93.5%) DeepSeek-V3 responses (risk difference, 8.9 percentage points; 95% CI, −13.3 to 31.1; p = 0.344). Exploratory question-level risks were observed in HPV16/18 positivity, normal cytology, partner management, and pregnancy scenarios.ConclusionBoth two models demonstrated generally strong performance across the evaluated quality domains. ChatGPT achieved higher ratings in several quality domains, but response-level binary safety differences were imprecise and not statistically significant. Both models require guideline-based clinical oversight.
Free Newsletter

Clinical research that matters. Delivered to your inbox.

Join thousands of clinicians and researchers. No spam, unsubscribe anytime.