METHODS
In this observational cross-sectional study, 50 text-based and 30 case-based multiple-choice questions
derived from RECIST 1.1 were administered to eight LLMs with three different prompts and two junior
radiologists with seven years of experience. Responses were independently scored as correct or incorrect,
and non-parametric statistical analyses were performed to compare performance across groups.
RESULTS
LLMs demonstrated promising performance in text-based interpretation about RECIST, with only minor
performance variations. Claude 3.5 Sonnet had the most successful performance, achieving 83.3%
accuracy on case-based and 90% on text-based questions. Other models exhibited robust performance,
with no significant differences in case-based assessments between LLMs and radiologists. LLMs achieved
similar results across the three different prompts with minor variations.
CONCLUSION
LLMs have great potential for response evaluation in oncological imaging and not only support radiologists
but may soon redefine clinical workflows, setting a new benchmark for diagnostic excellence in radiology.
Keywords: ChatGPT; cancer; large language models; response; treatment
The radiology report is vital in guiding patient
management in oncology, requiring meticulous comparison
with prior studies and assessment. Response
Evaluation Criteria in Solid Tumors (RECIST) guideline,
revised in 2009 to RECIST 1.1, was developed to
address this need. RECIST guideline comprises criteria
such as defining measurable lesions (i.e., which measurement
defines a measurable lymph node), identifying
target lesions (i.e., which criteria the target lesion
must meet), and categorizing response types (regression,
stable disease, or progression). It provides a
standardized approach to reporting solid tumor measurements
and defines objective criteria for assessing
changes in tumor size, ensuring a consistent and reliable
approach to reporting.[
Previous studies have evaluated the proficiency
and knowledge of various LLMs in different specific
types of cancer.[
To the best of our knowledge, no study has compared
the performance of LLMs in relation to RECIST
1.1, a critical guideline in the radiological reporting of
follow-up imaging in cancer patients. We aimed to fill
this gap by evaluating the knowledge of various LLMs
in the RECIST 1.1 guideline and comparing them with
that of radiologists.
The MCQs did not include any authentic patient
data or images; therefore, ethical committee approval
was neither required nor applicable for this study.
Methodological transparency and reproducibility were
ensured by adhering to the Standards for Reporting
Diagnostic Accuracy Studies (STARD) guideline.[
Data Collection for Text-based and Case-based
Multiple-choice Questions
Design of Input-output Procedures for LLMs
The MCQs were administered sequentially within
a single conversation session per LLM to maintain
uniformity. None of the LLMs underwent additional
pre-training or fine-tuning by the study authors, and
no supplementary details that could potentially affect
the study results were provided (Fig.
Radiologist Performance Evaluation
Statistical Analysis
A total of 50 text-based MCQs and 30 case-based
MCQs were utilized in the study. These questions comprehensively
covered the all sections of RECIST 1.1
and tested the application of the information therein.
Each question was carefully constructed to focus on a
single, specific, and critical concept relevant to radiological
practice under this guideline. Each MCQ had
5 choices and only one choice was correct. A complete
list of text-based and case-based MCQs and dataset of
the study are available in the Appendix.
The three different input prompts provided to the
LLMs were: Prompt 1: "Act like a professor of radiology
who has 30 years of experience in oncological
imaging, especially with studies on RECIST 1.1. Give
just the letter of the most correct choice of multiplechoice
questions that I will ask you. Each question has
only one correct answer." Prompt 2: "You are a senior
academic radiologist. I have some questions about
RECIST 1.1. I will ask you multiple-choice questions
with a single correct answer. Provide only the
letter of the most accurate choice for each." Prompt
3: "I have a few questions about RECIST 1.1 criteria.
Some of them are text-based, and some of them are
case-based questions. I will present you with multiple-
choice questions, and each has only one correct
answer. Please reply with the letter corresponding to
the best choice only, without any explanation." These
prompts were consistently employed across eight
distinct platforms with default hyperparameters by
R3 in February 2025: Claude 3 Opus and 3.5 Sonnet
(https://claude.ai.com), ChatGPT-o1, ChatGPT-4o
(https://chat.openai.com), Gemini 1.5 Pro (https:// gemini.google.com), Mistral Large 2 (https://mistral.
ai), Llama 3.1 405B (https://metaai.com), and Perplexity
Pro (https://perplexity.ai). In order to assess
the consistency in the responses of the three different
prompts within each model, the responses of each
prompt and the model were evaluated carefully, and
the prompt with the most successful responses for all
models was recorded by R3.
R1 and R2 independently answered the MCQs in a
blinded manner in January 2025 using their personal
computers. They completed text-based MCQs first, immediately
followed by case-based MCQs without any
interval. R3 separately evaluated their answers and categorized
them as correct (1) or incorrect (0).
The Kolmogorov-Smirnov test assessed data distribution.
Descriptive statistics (minimum, maximum, median,
interquartile range, percentages) were calculated. As
the data were non-normally distributed, non-parametric
tests were used. Consistency and performance across
three prompts were evaluated with the Friedman and
McNemar tests; the latter compared correct response
rates between LLMs and radiologists. Chi-square tests
assessed differences by question type. Bonferroni correction
was applied for pairwise comparisons (p ≤0.028
significant), while p ≤0.05 indicated significance for
consistency and prompt-related analyses.
With "Prompt 1", Claude 3.5 Sonnet demonstrated
the highest accuracy at 83.3%, followed by R2 and Gemini
1.5 Pro, both of which achieved 80.0% (p>0.028). R1
closely followed with 76.7%, while ChatGPT-4o, Llama
3.1 405B, and Mistral Large 2 each recorded 73.3%
(p>0.028). Claude 3 Opus and ChatGPT-o1 shared an
accuracy of 66.7%. Perplexity Pro exhibited the lowest
accuracy among LLMs and radiologists, with 60.0%
(p>0.028) (Fig.
There was no significant difference in accuracy on
case-based MCQs among LLMs and between LLMs
and radiologists (p>0.028) (Table
Text-Based MCQs
Similar to case-based MCQ, all models reached the
highest performance with "Prompt 1" among the three
different prompts. The answers given by all models
to the questions with these prompts were consistent, and there were no significant performance differences
among models with them (p>0.05) (Table
With "Prompt 1", Claude 3.5 Sonnet achieved
the highest accuracy at 90.0%, followed by Claude 3
Opus and ChatGPT-o1, both scored 84.0% (p>0.028).
ChatGPT-4o recorded an accuracy of 82.0%. Gemini
1.5 Pro attained 74.0%, the accuracy of Mistral Large
2 and Llama 3.1 405B at 72.0%. R2 (T.C.) had a slightly lower accuracy of 70.0% (p>0.028). Perplexity Pro
demonstrated the lowest performance among all models,
with an accuracy of 68.0% (Fig.
Claude 3.5 Sonnet outperformed Mistral Large 2, Llama 3.1 405B, and Perplexity Pro, achieving the highest scores on text-based questions (p=0.012, p=0.022, p=0.007). It also demonstrated superior performance according to R1 and R2 (p=0.021, p=0.021).
When other LLMs were compared among themselves
and with radiologists, there was no significant
difference in performance between them (p>0.028)
(Table
Coskun et al.[
Another important result of our study is that LLMs
responded as well as radiologists on text-based and
case-based MCQs that require analysis of the findings
and data obtained. This result suggests that LLMs are
successful in analyzing and reasoning texts such as
radiology reports and providing the status of the disease
(progression, stable, or regression) according to
the RECIST guideline, which is the most critical for
clinicians. To our best knowledge, there are no studies
evaluating the performance of LLMs on case-based
questions about cancer. Previous studies have evaluated
LLMs" knowledge of cancer and cancer-related
guidelines, which were only text-based. Çıtır reported
that ChatGPT-3.5 gave largely correct answers to questions
about oral cancer, 51.25% gave "very good" and
46.25% gave "good" answers, and the overall reliability
was 97.5%.[
In our study, all LLMs with three different prompts
showed great consistency with minor differences in
responses. Due to the nature of LLMs, it is a surprising
result that these models, which largely determine
their answers according to the given prompt, perform
similarly with different prompts to the questions about
RECIST 1.1.[
The impressive performance of Claude 3.5 Sonnet-
achieving 83.3% accuracy on case-based MCQs
and 90% on text-based MCQs-indicates great potential
of its model in this field. The observed minor variations
in accuracy among LLMs can be largely attributed
to differences in their underlying architectures.
Models with real-time web access capability, such as
Gemini 1.5 Pro and Perplexity Pro, frequently derive
their responses from non-scientific sources, which may
account for the comparatively lower performance of
web-enabled LLMs relative to those without internet
access. In contrast, some of the ChatGPT and Claude
models are trained on closed datasets, potentially contributing
to their enhanced reliability.
Limitations of the Study
Second, we compared the accuracy of LLMs against
two general radiologists with seven years of experience.
It is likely that more experienced senior radiologists,
particularly those with more specialized knowledge
about oncological imaging, would achieve higher performance.
Senior radiologists who are specialized and/
or subspecialized in this field may perform even better
than LLMs, but since follow-up images of cancer
patients are often evaluated by general radiologists in
daily practice for many different reasons, the radiologists
included in this study are general radiologists to
better reflect real-life practice.
Lastly, this study assessed the performance of
LLMs about RECIST 1.1 textually, while visual evaluation
remains an integral component of radiological
assessment. As such, the results of our study may
not fully reflect the real-world applicability of LLMs
in this field. It is important to emphasize that while
LLMs performed well on structured MCQs, this
could not directly reflect their ability in actual imaging
interpretation for RECIST.
Our study has a few limitations. First, the number of
questions was limited, and the assessment relied solely on MCQs. The performance of LLMs on open-ended
questions was not evaluated in the study, which may
have led to exaggerated LLM performances.
Informed Consent: The authors declare that this study was conducted without any commercial or financial relationships that could be construed as a potential conflict of interest.
Conflict of Interest Statement: The authors have no conflicts of interest to declare.
Funding: No funding was received for this study.
Use of AI for Writing Assistance: No AI technologies utilized.
Author Contributions: Concept - E.Ç.; Design - E.Ç., T.C.; Supervision - Y.C.G.; Materials - E.Ç.; Data collection and/ or processing - E.Ç.; Data analysis and/or interpretation - E.Ç., T.C.; Literature search - E.Ç., T.C.; Writing - E.Ç.; Critical review - E.Ç., T.C., Y.C.G.
Peer-review: Externally peer-reviewed.