top of page

Ambient Voice Technology in the NHS

Writer: Steve Franklin
Steve Franklin
10 minutes ago
19 min read

When better technology creates weaker controls

Why safer adoption of AI scribes requires Human Factors, inclusive design and dynamic clinical governance.


Ambient voice technology promises something that healthcare has sought for years: more attention for the patient and less time spent documenting the consultation.


An ambient voice technology system, often described as an ambient scribe or AI scribe, listens to a clinical conversation and converts it into a draft clinical note, letter, code or other structured output. NHS England is encouraging organisations to deploy the technology at pace as part of a wider move towards digitally enabled care. Its published evidence reports a 23.5% increase in direct patient interaction time, an 8.2% reduction in overall appointment length and, in emergency departments, a 13.4% increase in patients seen per shift during a sponsored evaluation [1, 2].


These are attractive benefits. Clinical documentation consumes scarce professional time, contributes to workload and can pull attention away from the person in the room. Used well, ambient voice technology could support better interaction, faster documentation and reduced administrative burden. However, the scale and speed of adoption create an important governance question:

What happens when the technology improves, but the controls around its use deteriorate?


The answer matters because the risk profile of ambient voice technology is unlikely to remain static. Some hazards may reduce as speech-recognition and generative AI models improve. Other risks may grow as users become familiar with the technology, trust it more readily and devote less attention to checking its output.


The central patient-safety challenge may therefore not be whether the technology works. It may be whether the surrounding system continues to behave as though it can sometimes fail.


The new UK Government toolkit makes risk appetite explicit

On 8 September 2026, the UK Government published its AI Risk Management Toolkit. Although written for government and the wider public sector rather than specifically for clinical technology, its central message is directly relevant to the NHS: responsible adoption requires active risk identification, treatment, monitoring and reporting throughout the AI lifecycle, not a one-off assessment before go-live [15].


Most importantly, the toolkit reiterates the need to be clear about risk appetite. It says that AI risk appetite should, where possible, be set at organisational level, aligned with the wider organisational risk appetite and signed off at board level. That is not governance decoration. Without a clear and measurable appetite, teams do not know which risks may be accepted in pursuit of benefit, which need further treatment and which should stop deployment [15].


A vague ambition to “use AI safely” is not a risk appetite. Boards need to say where they are willing to take a good risk, where they are cautious and where the consequences demand very little tolerance.


The toolkit also recognises that there is no final version of an AI system to validate once and forget. Models, data, interfaces, deployment contexts and user behaviour change. Treatments can lose efficacy. Risk management therefore has to continue from use-case selection through procurement, deployment, monitoring, material change and eventual retirement [15].


Questions the toolkit should prompt boards and governance leads to ask

  • Who is ultimately accountable for the AI solution and its outputs, and do they have sufficient authority, information, people and budget?

  • What is our appetite across clinical safety, fairness, transparency, accountability, technical robustness, security, legal compliance, contestability and financial exposure?

  • Who is affected by the output, including people who never directly interact with the tool, and how can they challenge it or seek rectification and redress?

  • What performance must the system maintain in the real context of use, including less-than-ideal conditions, unexpected data and natural drift?

  • If a person is expected to provide human-in-the-loop control, do they have the expertise, source information, time, interface and authority needed to intervene effectively?

  • How will fairness be tested during scaled deployment, and how will the organisation identify unequal or unlawful outcomes that aggregate performance data may conceal?

  • What monitoring, incident response, bypass and switch-off arrangements are in place, and what evidence will trigger them?

  • What existing governance and risk controls can be used, and where could AI create gaps between established committees, policies and accountabilities?


NHS England is moving from experimentation towards adoption at scale

NHS England describes ambient scribes as products that can convert speech into structured clinical documentation and states that NHS organisations are being directed to deploy ambient voice technology at pace. A national self-certified supplier registry has been created to help organisations compare suppliers and support local procurement and assurance [1, 2].


In July 2026, NHS England and the Medicines and Healthcare products Regulatory Agency clarified the regulatory position. Products intended solely to transcribe or summarise clinical conversations, draft correspondence or suggest clinical codes for review by a clinician are not necessarily medical devices under the current regulatory framework. Products intended to support diagnosis, treatment or automated clinical action remain subject to medical-device regulation [3].


This boundary is legally and operationally important, but it must not be mistaken for a boundary between safe and unsafe technology. A clinical note may shape diagnosis, prescribing, referral, follow-up and the account available to the next professional. An incorrect summary does not need to make a treatment decision autonomously to influence one.


The Royal College of Physicians and others have expressed concern that excluding some ambient voice products from medical-device regulation could create gaps in oversight. The BMJ has highlighted concerns about omissions, hallucinated content, coding errors and the transfer of responsibility to clinicians who may not be trained to assess how these products generate their outputs [4].

“Not a medical device” should therefore never be interpreted as “not a source of clinical risk”.


The risk profile is dynamic, not fixed

A conventional risk assessment can encourage people to think of risk as a property of the product: the system has a known accuracy rate, a set of documented hazards and a defined residual risk.


In practice, risk emerges from the interaction between the technology, clinician, patient, consultation, acoustic environment, interface, clinical record, local workflow, workload, interruptions, training, organisational expectations and quality of monitoring.


Ken Catchpole’s work emphasises the influence of tasks, teams, technology, environment and organisational conditions on clinical performance. This directs attention away from treating error as an isolated failure of either the clinician or the product. The relevant unit of analysis is the clinical work system into which the technology is introduced [5].


This means an ambient scribe can become technically more accurate while the overall sociotechnical system becomes less safe.


At first, clinicians may check every generated note carefully. They may compare the output with their memory of the consultation, correct language and verify medication details. The clinician is functioning as a strong compensating control.

If the output is usually accurate, however, checking may gradually become shorter and more superficial. Verification can shift from active comparison to recognition: the note appears plausible, contains the expected headings and broadly reflects the consultation. The clinician therefore accepts it.


The system may have improved, but the strength of the human control has deteriorated.


The control paradox

This produces an automation control paradox:

The more reliable the system appears to become, the less reliably people may check it.


Initially, the technology and its human checking process may operate as two partially independent protections. As trust develops, the protections can become increasingly dependent. The clinician’s judgement begins with the output produced by the system, and the check becomes an assessment of whether that output looks reasonable.


The generated text is no longer just an object being reviewed. It becomes the starting point, structure and anchor for the review.


Daniel Kahneman’s work on fast and slow thinking helps explain this. A fluent, coherent and professionally formatted note is easier to process than a fragmented or obviously poor one. Fluency can create a feeling of confidence and familiarity, even where clinically important information has been omitted. The reviewer may engage in rapid pattern recognition rather than effortful reconstruction of the original consultation.


Kahneman and Gary Klein concluded that professional intuition is most dependable where the environment contains sufficiently regular cues and people receive adequate opportunities to learn from timely, reliable feedback. They also warn that subjective confidence is not itself a reliable indicator of judgement accuracy [6].


Ambient documentation creates a difficult environment for developing reliable checking expertise. Many errors may be rare, subtle or only discovered much later. An omitted allergy, misunderstood dose or inaccurately attributed symptom may not generate immediate feedback. The clinician can therefore gain extensive experience of approving AI-generated notes without gaining equally strong experience of recognising the conditions under which the system fails.

Repeated use may build confidence faster than it builds expertise.


From checking the route to following the blue line

Think back to the early years of satnavs and Google Maps. Many of us would inspect the whole route before setting off. Did it use the road we expected? Was it sending us through somewhere implausible? If the route felt wrong, we kept a paper map or local knowledge close at hand.


More than a decade later, that behaviour has changed. Most people glance at the arrival time, press start and follow the blue line. The technology has improved, but familiarity has also changed the checking task. We have not simply become better at verifying navigation. In many situations, we have stopped treating verification as necessary.


AI-generated information may follow the same path, only faster. Today, users may approach an AI-generated note, briefing or board paper cautiously. As the prose becomes more fluent and the outputs are usually useful, active verification may quietly become a scan for obvious oddness. Then the source material is opened less often. Eventually, the generated account may become the thing against which reality is judged.


This is where Sidney Dekker’s work on drift into failure is especially useful. Complex systems do not always fail because someone makes one dramatic, reckless choice. Under pressure to deliver success with finite resources and competing goals, small locally sensible adaptations can accumulate until the system is operating close to a boundary nobody intended to approach [16].

In this context, drift may look like a review completed after the next appointment, then a batch approval at the end of clinic, then checking only medicines and diagnoses, then accepting the note because previous ones were right. None of those steps necessarily feels dangerous on the day. Together, they can transform an apparently strong control into a ritual.


Checking is a task and must be designed as one

“Clinician review required” is not, by itself, an adequate safety control.

A requirement to check transfers work to a person, but it does not specify what must be checked, against which source, at what point in the workflow, with how much uninterrupted time, how discrepancies should be resolved, which fields are safety-critical, how uncertainty should be displayed, when the output should be rejected or how the organisation will know whether meaningful checking occurred.


From a Human Factors perspective, checking must be treated as a clinical task with cognitive demands, environmental conditions and predictable failure modes.

A clinician reviewing the note immediately after a consultation may still remember the conversation. However, they may also be under pressure to see the next patient. If review is delayed, workload may be lower, but memory of the discussion is likely to have degraded. If the source audio is not retained or readily accessible, the reviewer may be comparing the AI summary only with their recollection rather than an independent source.


The assurance statement “the clinician remains responsible” does not solve this design problem. Responsibility cannot compensate reliably for poor feedback, excessive workload, weak interface design or loss of access to the original information.


Ambient technology changes the choice architecture of documentation

Richard Thaler and Cass Sunstein describe how choice architecture influences behaviour through defaults, presentation, feedback, mappings and the expectation that people will make errors. Decisions are never made in a neutral environment. The designer of the environment influences what people notice and what they are likely to do [7].


Ambient voice technology changes the default in clinical documentation.

Without an ambient scribe, the clinician constructs the note. With an ambient scribe, the clinician is presented with an existing draft and asked to edit or approve it. The default has shifted from a blank record to an apparently complete account.


This matters because accepting a plausible draft is cognitively easier than rejecting and reconstructing it. The system becomes an anchor. Information present in the draft is made salient; information that has been omitted is, by definition, less visible.


The interface may reinforce this effect. A prominent “approve” button, an unobtrusive “edit” function and no visual distinction between directly transcribed content and AI-generated inference all shape behaviour. If approval is required before moving to the next task, operational pressure can turn acceptance into the path of least resistance.


Good design should make the safest behaviour the easiest behaviour. It should also expect error. For ambient documentation, this could mean highlighting medicines, allergies, diagnosis, laterality and follow-up for explicit verification; distinguishing transcription from inference; presenting uncertainty; requiring active confirmation of critical fields; making correction and rejection easy; and providing feedback on recurrent error patterns.


Work-as-imagined will not be work-as-done

A deployment plan may state that every clinician will review every document fully before it enters the clinical record. That is work-as-imagined and work-as-prescribed.


Steven Shorrock’s work distinguishes these representations from work-as-done: the variable, adaptive and context-sensitive activity that takes place in real operational settings. Shorrock warns that procedures, metrics, reports and managerial assumptions are proxies for actual work rather than complete accounts of it [8].


Under real conditions, clinicians may scan rather than read, check selected sections, review documents while managing interruptions, rely on memory, assume familiar templates indicate accuracy, approve several notes together or create informal workarounds.


These behaviours should not automatically be interpreted as carelessness. They may represent adaptations to workload, time pressure and interface design.

Shorrock’s description of the efficiency-thoroughness trade-off is particularly relevant. Staff are expected to work efficiently, but after an unwanted event they may be criticised for not being sufficiently thorough. Organisations should not build productivity benefits on reduced documentation time and then treat the clinician’s thoroughness as an unlimited residual safeguard [8].


If the business case assumes faster consultations, higher throughput and reduced administrative burden, the clinical safety case must examine whether those same benefits reduce the time and attention available for checking.


Who is most likely to be misunderstood?

Caroline Criado Perez’s Invisible Women describes how technologies and services can reproduce existing inequalities when the underlying data treat one group as the default. She uses voice recognition as an example of technology that may perform less effectively for women where training and design data skew towards male voices [9].


The problem extends beyond a binary comparison between men and women. Speech varies with regional dialect, ethnicity, cultural background, first language, age, pitch, speech rate, socioeconomic background, disability, illness, fatigue, clinical vocabulary and environmental noise.


Research on automatic speech recognition has identified performance differences associated with gender, age, regional accent, non-native accent and speech impairment. These differences arise not only from who is present in training data but from pitch, articulation, language variation and the way systems are tested and optimised [10, 11].


In the NHS, regional variation is not a peripheral issue. A system evaluated predominantly using standard or southern English voices may not transfer uniformly to consultations in Birmingham, Liverpool, Newcastle, Glasgow, Belfast or rural and coastal communities. Even within a region, vocabulary and pronunciation can vary substantially by age, ethnicity and community.


The relevant question is not “Does the product understand British English?” It is:

Which people, speaking in which ways, under which clinical and environmental conditions, does it understand reliably?


What does the emerging clinical evidence show?

A 2026 study in BMJ Health & Care Informatics evaluated an AI scribe using simulated primary-care consultations, different communication styles, international English accents and several speech impairments. In that controlled single-system study, performance was broadly stable across the tested communication styles and most accents, although omissions were the most common error and were slightly higher for Chinese-accented and Indian-accented doctors. A particular phonological impairment produced a marked reduction in recognition performance [12].


This is more nuanced than a simple claim that ambient systems cannot understand accents. The study provides some reassurance for the system and conditions tested, but its authors still recommend clinician-in-the-loop verification, subgroup performance monitoring and predefined switch-off criteria where performance deteriorates [12].


It also demonstrates why local assurance matters. A product’s overall accuracy can remain high while clinically significant performance differences persist for particular subgroups, consultation types or speech characteristics.

A systematic review protocol published in BMJ Open reflects the relative immaturity of the evidence base. While ambient AI scribes promise gains in efficiency and reduced documentation burden, questions remain about documentation quality, accuracy, completeness, bias and the comparability of current studies [13].


National adoption should therefore be accompanied by national and local learning, not by an assumption that procurement from an accepted supplier completes the safety case.


The record has a life beyond the person who wrote it

A recent US court exchange offers a powerful, if uncomfortable, illustration. During the 2026 Lindsay Clancy trial in Massachusetts, psychiatrist Dr Jennifer Tufts was challenged about wording in a clinical note. In response to the disputed interpretation, she said: “I don’t care what it says. I know what I meant.” The case was not about an AI-generated note, and no inference should be drawn here about the wider care provided. The narrower lesson is about documentation: once a record is relied upon by another clinician, patient, investigator or court, the author’s private intention cannot repair ambiguous words on the page [18, 19].


AI makes that problem more complex. The clinician may know what they intended to communicate, but the system may summarise, compress or restructure the conversation into something different. If the clinician approves the output without detecting that difference, the record acquires authority from their name while its wording may partly reflect the model’s interpretation. Years later, “I know what I meant” may be even less persuasive because the central question will be whether the approved record accurately communicated it.


This is why provenance and edit history matter. Organisations should be able to distinguish what was directly heard, what was inferred or generated, what the clinician changed, what they actively verified and what entered the record unchanged. The purpose is not defensive bureaucracy. It is to preserve the reliability of the clinical narrative for everyone who must subsequently act on it.


Accuracy is not one number

A global accuracy percentage can obscure the errors that matter.

Consider two documents. One contains several minor stylistic inaccuracies but captures the medicines, symptoms, diagnosis and plan correctly. The other is linguistically polished but omits a red-flag symptom or changes a medication dose. The second may appear more accurate under some aggregate measures while creating substantially greater clinical risk.


Ambient scribe evaluation should distinguish between word-recognition errors, clinically important omissions, factual inaccuracies, hallucinated information, incorrect speaker attribution, loss of negation, temporal errors, coding errors, medication errors, laterality errors and inappropriate conversion of tentative discussion into a definitive diagnosis.


The BMJ research found that omissions predominated over factual inaccuracies and hallucinations in the system evaluated. This is significant because omissions are often harder to detect than false statements. A false statement can look wrong. Missing information leaves no visible prompt for the reviewer [12].


Copilot turns this from a product question into an organisational question

The same issues extend beyond ambient scribing. NHS England announced in June 2026 that Microsoft 365 Copilot would be made available to 505,000 clinicians and support staff, with rollout expected by October 2026. The stated uses include clinical and patient letters, discharge processes, service-data analysis, rotas, bed management, meeting minutes, HR, finance, procurement, board papers and organisational analysis. The preceding trial involved more than 30,000 staff across 90 NHS organisations and reported an average saving of 43 minutes per user per day [17].


That potential benefit is substantial. It also means that AI-generated text will not sit in one neatly bounded clinical application. It may appear across the information ecosystem through which trusts govern themselves and deliver care. A plausible error in a discharge draft and a plausible error in a board paper have different immediate consequences, but both can influence decisions, priorities and the account the organisation later relies upon.


Trusts therefore need more than a licence allocation and a generic acceptable-use policy. They need a use-case inventory, proportionate impact and risk assessment, information-governance controls, clear rules for patient and confidential information, role-based training, monitoring of benefits and harms, and governance for locally built agents. The risk appetite may legitimately differ between drafting an internal agenda, summarising a complaint, analysing finance data and producing text that enters a patient record.


Inclusive design also matters at scale. Access to Copilot does not create equal benefit if the training, workflows and interfaces assume a confident, office-based, full-time user with standard equipment, protected time and no accessibility needs. Trusts should evaluate who gains time, who inherits additional checking work, whose roles are transformed and whether women, disabled staff, part-time staff and other groups experience the technology differently. Women’s ergonomics is relevant here because job design, exposure, equipment, caring patterns and unpaid or invisible coordination work are not evenly distributed.

At trust scale, the risk is not simply that Copilot occasionally produces a wrong answer. It is that AI-generated material becomes normal organisational evidence before the organisation has designed how people should question it.


Does DCB0160 apply to ambient voice technology?

In most NHS deployments, yes.


Whether an ambient scribe is classified as a medical device is distinct from whether it is a health IT system that can introduce clinical risk. DCB0129 establishes clinical risk-management requirements for manufacturers of health IT systems, while DCB0160 establishes requirements for health and care organisations deploying and using such systems [14].


Therefore, an ambient voice product does not need to qualify as a medical device for DCB0160 considerations to arise. If its use can affect the completeness, accuracy, availability or interpretation of information used in clinical care, the deploying organisation needs a systematic clinical-risk-management process.

The organisation should not simply inherit the supplier’s conclusions. DCB0129 concerns how the manufacturer identifies, evaluates and controls clinical hazards associated with the product. DCB0160 concerns how the deploying organisation manages clinical risk within its own services, workflows, interfaces, staff groups and patient population [14].


A supplier may demonstrate that its product operates safely under intended conditions. The NHS organisation must determine whether those intended conditions exist locally.


At the time of writing, NHS England is reviewing DCB0129 and DCB0160 to ensure they remain relevant to AI, cyber risk, greater digitisation and changing clinical workflows. The public consultation opened on 29 June 2026 and is due to close on 11 September 2026 [14].


A supplier registry is not a local safety case

The NHS England AVT Supplier Registry may help organisations understand supplier readiness and reduce duplication in procurement assurance. However, it is described as a self-certified registry and does not remove the deploying organisation’s responsibility to ensure that the chosen product is safe and appropriate for local use [1, 2].


Local clinical safety work should examine where the product will be used, who will use it, whose speech must be recognised, which outputs will be produced, where those outputs will go, how they will be checked, how performance will be monitored and when use will be restricted or stopped.


Controls have a half-life

Clinical safety cases often describe controls as if they remain equally effective once introduced. Behavioural controls rarely do.


Training is forgotten. Workflows drift. Staffing changes. The technology is updated. Users become more familiar and less cautious. A warning message becomes routine background. A mandatory confirmation becomes a click performed without reflection.


Organisations should therefore consider the control half-life of ambient voice technology: the period over which a control remains sufficiently effective before it requires testing, reinforcement or redesign.


This is particularly important where the safety argument depends on human vigilance. “The clinician will check” is not a permanent property of the system. It is a hypothesis about future behaviour under operating conditions.


The control must be monitored through direct observation, user research, output sampling and discussion with clinicians and patients. This reflects the distinction between work-as-prescribed and work-as-done, and Catchpole’s emphasis on studying real clinical work rather than relying solely on how processes are expected to operate [5, 8].


What boards and digital leaders should ask

Boards, clinical safety officers, CCIOs, CIOs and clinical governance teams should ask the following questions.


  1. What clinical decisions could be influenced by the generated output, even if the product does not make those decisions itself? Map direct and indirect influences on care.

  2. What is the independent source against which the clinician checks the output? A check without an accessible source may become a plausibility review.

  3. Which errors are most likely to be missed because they are omissions rather than visible inaccuracies? Design monitoring around clinically significant failure modes.

  4. How long is allocated for meaningful review, and is that compatible with the productivity case? Do not count the same time saving twice.

  5. How does performance vary across sex, ethnicity, age, regional accent, first language, disability and speech impairment? Aggregate metrics may conceal unequal performance.

  6. What happens when the product, model, prompt, interface or integration changes? Treat material changes as triggers for reassessment.

  7. How will the hazard log and clinical safety case remain live after deployment? Safety assurance must continue beyond go-live.

  8. What evidence would trigger restriction, redesign or suspension of use? Define escalation and switch-off criteria in advance.

  9. How will patients and clinicians contribute to monitoring work as it is actually done? Combine performance data with lived operational experience.


From clinician in the loop to a safe system around the clinician

Putting a person somewhere in the process does not automatically create a stronger system. The control lies in the interaction: what the person can see, what source they can compare, how uncertainty is represented, when they review, what else competes for attention, whether they can reject the output without penalty, and whether the organisation learns from the corrections they make. A person who is rushed, poorly informed or unable to intervene is not a meaningful control. They are simply present.


“Human in the loop” has become a reassuring phrase in AI governance. It suggests that a person remains available to identify and correct technological error.


But being present in the process does not mean being supported to exercise effective control.


The future safety of ambient voice technology will depend on whether organisations design usable interfaces, meaningful verification tasks, representative testing, clear feedback, protected review time, responsive monitoring, accessible reporting, strong digital clinical governance and a learning system capable of adapting as both technology and human behaviour change.

Ambient voice technology may become increasingly accurate. That is welcome. However, technical improvement should never be used as a reason to make the wider system less attentive to failure.


The safest organisations will not ask only, “How well does the AI perform?” They will also ask how the AI has changed the work, what controls still function in practice, for whom the system works less well, and what new risks are created as confidence grows.


That is the contribution of Human Factors and Ergonomics. It helps organisations move beyond product assurance and examine the changing relationship between people, technology, tasks, environments and governance.

The ambition should not simply be to put a clinician in the loop. It should be to design a safe, equitable and adaptive system around them.


Is your ambient voice technology programme designed around real clinical work?

Promethean HD supports NHS and independent healthcare organisations to assess and implement digital technology using Human Factors, systems thinking and proportionate clinical governance. We can help organisations examine changing clinical work, identify hazards and unintended consequences, test usability and checking arrangements, evaluate performance across staff and patient groups, connect DCB0160 activity with operational reality, and provide independent governance and board assurance.



Continue exploring the Promethean HD Knowledge Hub

For related practical thinking, explore our Digital Health and Useability articles, our Inclusive Design & Work Wellbeing resources, and our Women’s Work Design and Ergonomics category. You may also find value in our Human Factors in Healthcare knowledge hub, our article on weak signals and the absence of evidence, and “Who Controls the Past?” on the power of the clinical narrative.

If your trust is moving from AI pilot to wider adoption, Promethean HD can support a proportionate board and governance review of risk appetite, use cases, work-as-done, inclusive design, clinical safety interfaces, monitoring and benefits realisation.



References

1. NHS England. (2026). Adopting ambient scribing products in health and care settings. Available online

2. NHS England. (2026). Guidance on the use of AI-enabled ambient scribing products in health and care settings, Version 3. Available online

3. NHS England and MHRA. (2026). Medical device regulation for ambient voice technology products. Available online

4. Armstrong, S. (2026). AI scribes: UK regulator issues update on use of technology by doctors. BMJ, 394, e100540. Available online

5. Catchpole, K. Human Factors, patient safety and healthcare systems: research profile and publications. Available online

6. Kahneman, D., and Klein, G. (2009). Conditions for intuitive expertise: a failure to disagree. American Psychologist, 64(6), 515-526. Available online

7. Thaler, R. H., Sunstein, C. R., and Balz, J. P. (2012). Choice Architecture. In The Behavioral Foundations of Public Policy. Available online

8. Shorrock, S. (2020-2023). Work-as-imagined, work-as-measured and work-as-judged. Humanistic Systems. Available online

9. Criado Perez, C. (2019). Invisible Women: Exposing Data Bias in a World Designed for Men. London: Chatto & Windus. Available online

10. Holliday, N. R., and Reed, P. E. (2025). Gender and racial bias issues in a commercial “tone of voice” analysis system. PLOS ONE, 20(2), e0314470. Available online

11. Feng, S., Kudina, O., Halpern, B. M., and Scharenborg, O. (2021). Quantifying Bias in Automatic Speech Recognition. Available online

12. Draper, T. C., Leake, J., Cox, T., et al. (2026). AI-generated clinical summaries: errors and susceptibility to speech and speaker variability. BMJ Health & Care Informatics, 33, e101918. Available online

13. Garcia Sanchez, C., Goer, V., Kharko, A., et al. (2026). Use of ambient AI scribe in physicians’ clinical documentation: a protocol for a systematic review. BMJ Open, 16, e115562. Available online

14. NHS England. (2026). Review of digital clinical safety standards: DCB0129 and DCB0160. Available online

15. Government Digital Service. (2026). AI Risk Management Toolkit: guidance. Published 8 September 2026. Available online: https://www.gov.uk/government/publications/ai-risk-management-toolkit/ai-risk-management-toolkit-guidance

16. Dekker, S. (2011). Drift into Failure: From Hunting Broken Components to Understanding Complex Systems. CRC Press. DOI: https://doi.org/10.1201/9781315257396

17. NHS England. (2026). 500,000 NHS staff to get new artificial intelligence tools to help free up more time for patients. Published 8 June 2026. Available online: https://www.england.nhs.uk/2026/06/500000-nhs-staff-to-get-new-artificial-intelligence-tools-to-help-free-up-more-time-for-patients/

18. NBC10 Boston. (2026). Recap: Lindsay Clancy’s psychiatrist cross examined, nurse describes treating her. Published 10 August 2026; updated 12 August 2026. Available online: https://www.nbcboston.com/news/local/lindsay-clancy-trial-day-10-live-stream-live-updates/3994777/

19. Chobany, M. (2026). “I Know What I Meant”: The Ethical Responsibility of Clinical Documentation. Bioethics Today, 18 August 2026. Available online: https://bioethicstoday.org/blog/in-the-news-i-know-what-i-meant-the-ethical-responsibility-of-clinical-documentation/

Comments


bottom of page