A robotic hand holding a feather quill pen in vintage engraving style. (Photo Credit: Bilkis Islam, AdobeStock)

Artificial Intelligence (AI) scribes have been adopted rapidly throughout the U.S. healthcare system with little reflection on use. This novel technology records the audio of a clinical visit, creates a written transcript from this recording, and then drafts a clinical note in the electronic medical record (EMR) for a clinician to review. AI scribes have been hailed as a promising strategy to reduce documentation burden and resulting burnout—serious concerns that compromise clinician quality of life and contribute to attrition from the field. However, enhancing workflow efficiency by implementing AI scribes in the clinic is not without its costs. A central concern that deserves broader recognition is the inability of this technology to capture the nuances of human speech and to act as a critical interpreter. Of course, no note perfectly captures every detail of the clinical encounter. This inevitability calls to mind the Italian expression traduttore, traditore: every translator is, in some sense, a traitor. If so, the AI scribe is an especially treacherous one.

Translation theory is a field of study that engages how best to translate the text of one language into another, given the inexorable lexical gaps between them. Two sometimes-competing modes of translation are formal equivalence, a word-for-word rendition of the text, and dynamic equivalence, which aims to capture its essence, including context and sentiment, features of language that a rote rendering would misrepresent. A good translation is both precise and dynamic, faithful to original intent and an act of critical interpretation. In the balancing act of translating clinical encounters into representative notes that speak to these imperatives, AI scribes teeter off the wire.

Three follies of this technology are especially salient to this concern. First, AI scribes have structural features that undermine documentary accuracy. Ambient documentation systems have been shown to hallucinate and to interpolate inaccurate information into the medical record, including listing medications that a patient is not taking or diagnoses that she does not have. The failure of AI scribes on the front of literalism can thus pose real hazards to patient health. Second, AI scribes are structurally circumscribed when it comes to the communication they can capture. For instance, they are unable to register nonverbal communication, such as appearance, body language, and gait, all of which are important visual cues of disease presentation that a clinician might reasonably choose to include in a note. Likewise, AI systems are unreliable interpreters when it comes to figurative speech, like sarcasm, idiom, and metaphor. This limitation might result in an AI scribe rendering speech too literally—in a way that is nonsensical or even contrary to the speaker’s intended meaning. Finally, in addition to these issues of inadequate capture, AI scribes may also pose a threat to patients by capturing too much: that is, by recording personal information that a discerning clinician might elect to exclude from the EMR, such as immigration status or ongoing domestic abuse. AI scribes lack interpretive nuance and discretion, and this constraint may incur harms for patients downstream.

While clinicians are expected to revise the content that AI scribes produce, the very time pressure that prompted the development of this technology militates against careful review. Similarly, the intrusion of cognitive biases raises questions about how reliable a corrective mechanism this process of review actually is. For instance, automation complacency and the anchoring effects of a default note may cause clinicians to be unduly deferential to the AI-authored draft. Furthermore, delegating note drafting to an AI scribe may also undermine clinicians’ opportunity for information synthesis and reflection on the case at hand, accruing what scholars have called a “cognitive debt” and diminishing our critical faculties over time.

Large language models (LLMs), the technology underlying AI scribes, are advancing rapidly. Training these models on more diverse datasets, expanding their context windows, and incorporating ambient visual documentation may help address the problems of underinclusive and overinclusive capture described above. Yet as these structural innovations improve the accuracy and scope of AI-generated documentation, they may also foster greater confidence in—and dependence on—AI scribes, further compromising communicative parsing and clinical ratiocination. The resulting erosion of clinicians’ own translational capacities—the essential art of doctor-patient communication and clinical documentation—may prove the greatest betrayal of all.

Conflicts of Interest: I have no conflicts of interest to disclose.


Ursula Francis.

Ursula Francis, JD, PhD, MSc, graduated from the Master of Science in Bioethics at the Center for Bioethics at Harvard Medical School in 2025. Her background is in law (JD, The University of Chicago Law School) and literature (PhD Classics, Columbia University). She is currently a Clinical Ethics Fellow at the University of Chicago MacLean Center for Clinical Medical Ethics.

The Bioethics Blog, hosted by Harvard Medical School Center for Bioethics, features short, conversational pieces from Center alumni and faculty on pressing issues in bioethics. Posts are curated by alumni editors Rebecca Li, Carolyn Baker Ringel, and Risa Jampel. Through these contributions, we aim to share ethical insights, invite further dialogue, and connect bioethicists across the globe working in related areas of research and practice. The views and opinions expressed in this article are those of the author and do not necessarily reflect the views, policies, or positions of the Center for Bioethics at Harvard Medical School. Please contact bioethics_alumni@hms.harvard.edu for more information about getting involved in this project.