METHODS
We evaluated the effectiveness of GOOGLE Gemini 2.5 Pro Preview 03-25 in analyzing drug reimbursement
policies within the SUT. A structured prompt was developed to ensure responses
strictly adhered to regulatory text. A total of 80 oncology-related test questions, covering multiple
cancer types, were used to assess the model"s accuracy. Responses were categorized as correct and
complete, correct but incomplete, or incorrect. Performance metrics, including precision, recall,
and F1-score, were calculated. An iterative prompt engineering process was employed to optimize
model performance.
RESULTS
The LLM provided completely correct responses to 77 (96.3%) of 80 test cases and correct but incomplete
responses to 3 (3.7%), with no incorrect answers. Performance metrics demonstrated high accuracy
(precision: 1.00, recall: 0.96, F1-score: 0.98). The model successfully processed medical terminology
variations but showed limitations in synthesizing implicit reimbursement rules.
CONCLUSION
LLMs demonstrate strong potential for interpreting cancer drug reimbursement regulations, reducing
administrative burden for oncologists. Future refinements should address inference limitations to enhance
regulatory compliance support in clinical practice.
Keywords: Artificial intelligence; cancer; large language models; oncology policy
Artificial intelligence (AI) has demonstrated significant
potential in oncology, improving diagnostic
accuracy, treatment decision-making, and workflow efficiency.[
This study aims to evaluate the effectiveness of
LLMs in analyzing drug utilization principles in cancer
treatment according to the SUT regulations. Specifically,
we will investigate the capacity of LLMs to interpret
complex regulatory text and provide relevant guidance
on drug eligibility and reimbursement criteria within
the context of cancer treatment.
Large Language Models and Prompt Development
Testing and Evaluation of Model Responses
Support, defined as the number of instances per category,
was also recorded to ensure balanced representation
across different cancer types.
Iterative Refinement via Prompt Engineering
Statistical Analysis
We employed the advanced LLM, GOOGLE Gemini
2.5 Pro Preview 03?25, to process and analyze the SUT regulations. The model was selected based on its
demonstrated capabilities in natural language understanding
and complex regulatory interpretation, large
context capability up to 2 million tokens.[
To evaluate the LLMs" ability to extract and apply SUT
regulations, a set of test questions was developed by independent
experts in medical oncology. The test questions
were based on real-world clinical scenarios and
covered various tumor types, including lung, breast,
gastrointestinal, gynecological, genitourinary, central
nervous system, melanoma and skin, sarcoma, lymphoma,
multiple myeloma, and other cancers. Each
question was tested in an isolated session to prevent
contextual memory effects. The generated responses
were then compared against the SUT text to assess
their accuracy. LLM responses were categorized as: (1)
Correct and complete, (2) correct but incomplete, or
(3) incorrect. We calculated precision, recall, and F1-
score to quantitatively assess performance. These metrics
were computed as follows:
A key aspect of this study was the iterative refinement
of the prompt to improve model accuracy. When a response
was classified as incorrect or incomplete, an error
analysis was conducted to identify potential sources
of misunderstanding. We then updated the prompt
by incorporating specific clarifications to address these
issues. This process was repeated iteratively until the
model demonstrated optimal performance across multiple
test scenarios (Fig.
LLMs: Large language models; SUT: Health practice regulation.
Descriptive findings were reported as frequencies and
percentages. The Python sklearn.metrics module was
utilized for calculating classification metrics, while
matplotlib.pyplot and seaborn were employed to visualize
the confusion matrix.
The LLM demonstrated a high degree of accuracy
in extracting and interpreting drug reimbursement
criteria from the SUT regulations. Out of 80
test questions, the model provided completely correct
responses to 77 (96.3%) and correct but incomplete
responses to 3 (3.7%), with no instances (0%) of incorrect answers (Fig.
The model's ability to correctly interpret and extract
drug eligibility criteria from the SUT aligns with prior
studies evaluating LLMs in clinical decision support.
Benary et al.[
Despite these strengths, three specific test cases were
answered incompletely. Notably, when asked about
first-line treatment options for metastatic lung cancer
without driver mutations, the model omitted some
chemotherapies that could be used without explicit indication
restrictions. Similarly, in a question concerning
high-risk bone metastases in metastatic hormonesensitive
prostate cancer, the model correctly identified
denosumab but failed to acknowledge zoledronic acid,
despite its inclusion under broader reimbursement
conditions. A third case regarding metastatic laryngeal
squamous cell carcinoma highlighted the model"s tendency
to focus on explicitly mentioned therapies (cetuximab)
while neglecting chemotherapy options that
could be inferred from related SUT provisions. These
limitations were likely due to our prompt design, which
strictly instructed the model to rely only on SUT text
without making medical inferences. While this prevented
hallucinations, it also restricted the model"s ability to
synthesize related rules. This tradeoff is consistent with
findings from AI studies in radiology and oncology,
where models performed well when retrieving direct
information but struggled with complex reasoning.[
Our study also tested the model"s ability to handle
variations in medical terminology. We deliberately
rephrased key terms (e.g., "liver cancer" instead of "hepatocellular carcinoma," "cerbB2" instead of
"HER2") and used common abbreviations to mimic
real-world oncology practice. The model successfully
recognized these variations, suggesting that it can
adapt to different ways oncologists phrase reimbursement-
related queries. This contrasts with earlier LLM
studies where models had difficulty processing medical
language inconsistencies.[
The integration of AI into oncology decision-making
has been met with both enthusiasm and caution.
While AI holds significant promise for improving efficiency
in clinical workflows, concerns remain regarding
reliability, accountability, and regulatory compliance.[
A significant challenge in cancer drug reimbursement
worldwide is the misalignment between national
policies and international treatment guidelines, often
resulting in delays in patient access to novel therapies.
[
Informed Consent: Not applicable.
Conflict of Interest Statement: The authors have no conflicts of interest to declare.
Funding: The authors declared that this study received no financial support.
Use of AI for Writing Assistance: The authors used ChatGPT (GPT-4o, OpenAI) for language editing only. All content was authored by the authors without AI-generated material.
Author Contributions: Concept - R.I., Z.A.; Design - Z.A., M.K.; Supervision - Z.A.; Funding - Z.A.; Materials - R.I., Z.A., A.F., M.N.R.; Data collection and/or processing - A.O., Ö.A., R.I.; Data analysis and/or interpretation - R.I., Z.A.; Literature search - R.I.; Writing - R.I.; Critical review - Z.A., O.A.
Peer-review: Externally peer-reviewed.