Artificial Intelligence in Surgical Documentation: A Comparative Evaluation of Large Language Model-Generated and Surgeon-Generated Operative Documentation

Authors

  • Dr Honeypalsinh H Maharaul
  • Dr Ketul Shah
  • Dr Anant Shah
  • Dr Yuvraj Kothari
  • Dr Vipul Gurjar

Keywords:

artificial intelligence; large language models; surgical documentation; operative notes; ambient scribes; documentation burden

Abstract

Objective To evaluate the role of artificial intelligence — particularly large language models and ambient AI scribes — in surgical documentation; to compare AI-generated with surgeon-generated documentation for readability, accuracy, and completeness; and to synthesise the evidence on benefits, limitations, and implementation requirements of these technologies.

Design Cross-sectional comparative evaluation of AI-generated versus surgeon-generated documentation, with structured synthesis of published evidence.

Setting Department of General Surgery, SBKS Medical Institute & Research Centre, Sumandeep Vidyapeeth, Piparia, Waghodia, Vadodara, Gujarat, India.

Participants De-identified operative report transcripts and preoperative briefing notes assessed against expert-reviewed benchmarks by a panel of board-certified surgical reviewers. [Sample size to be completed by the authors.]

Main outcome measures Readability, accuracy, and completeness of generated documentation; documentation efficiency and clinician workload; and error profiles (omission, addition, and factual errors).

Results AI documentation tools demonstrated substantial efficiency gains: an LLM generated operation notes for laparoscopic appendicectomy in seconds with 78.8% (23.66/30) coverage of GIRFT guideline items(1), and ambient AI scribes were associated with perceived efficiency improvement in 75.8% of surgeons with high baseline documentation burden (p<0.001(2). In a pilot of ambient AI in surgical outpatient care, burnout fell from 67% to 33% and NASA-TLX mental demand from 14 to 5 (3). Accuracy concerns persisted: omission errors constituted 71% of all errors across four commercial scribe tools (4), clinicians reported hallucinations such as misgendering and unsolicited diagnoses (5) , and structured ambient notes, while more complete, required physician oversight (6). Video-based AI operative reports achieved higher accuracy than surgeon-written reports (7), and an AI computer-vision platform documented the Critical View of Safety with high intra-rater reliability (Cohen's kappa >0.8) (8).

Conclusions AI and LLM technologies can reduce documentation burden and improve the completeness and structure of surgical records, but current systems remain subject to omission errors and factual inaccuracies. Safe integration requires mandatory clinician review, guardrails for high-risk documentation, external validation, and robust governance frameworks. AI should augment — not replace — surgeon judgment in documentation (9,10).

Downloads

Published

2026-09-05

How to Cite

Maharaul, D. H. H., Shah, D. K., Shah, D. A., Kothari, D. Y., & Gurjar, D. V. (2026). Artificial Intelligence in Surgical Documentation: A Comparative Evaluation of Large Language Model-Generated and Surgeon-Generated Operative Documentation. International Journal of Artificial Intelligence and Machine Learning, 6(9s), 1143–1148. Retrieved from https://www.svedbergopen.com/index.php/ijaiml/article/view/1574