Skip to content

Hero image: Women clinicians and engineers discussing a digital health workflow beside a patient in natural daylight

All insights

Technology

AI and Connected Technologies in Women's Health: A Responsible Path to Clinical Value

11 min readEvidence synthesis
Read the evidence

The question in focus

A clinical and governance framework for combining AI, sensors, interoperable records and digital care without overstating unproven benefits.

Evidence at a glance

6

consensus principles guide WHO's governance of AI for health

The principles cover autonomy, safety, transparency, accountability, equity and responsive sustainability. They are governance requirements, not evidence that a particular AI product improves outcomes. [1]

Ethics and governance of artificial intelligence for health

Convergence should solve a defined care problem

AI can classify images, estimate risk, summarize records or support workflow. Sensors can collect longitudinal signals. Interoperability standards can move structured information between authorized systems. None of these capabilities is a clinical outcome by itself. A defensible programme begins with a specific unmet need, an intended user and a decision that the technology is meant to inform. WHO warns that health AI must be governed around human autonomy, safety, transparency, accountability, inclusion and sustainability. [1]

The most credible opportunities in women's health are often practical: making pregnancy warning signs visible across care settings, reducing repetitive history-taking, helping clinicians find trends in menstrual or symptom data, improving access to evidence-based mental health support, and monitoring treatment response. Each use case has different consequences if an output is wrong, late, missing or misunderstood. Risk classification should determine evidence, oversight and human review.

  • State whether the system informs, recommends, prioritizes or acts.
  • Identify the person accountable for the final clinical decision.
  • Define the harm of false reassurance as carefully as the harm of a false alert.

Authorization is use-case specific

The FDA maintains a list of AI-enabled medical devices authorized for marketing in the United States. The agency explicitly notes that the list is not comprehensive and that authorization relates to the submitted intended use and evidence, not to AI as a category. A model cleared for one population, input device or clinical workflow should not be assumed to work in another. [2]

Consumer wellness tools, clinical decision support and software as a medical device can fall under different regulatory pathways across jurisdictions. Product teams should determine classification before deployment, document claims narrowly and maintain a total product lifecycle plan. FDA guidance emphasizes transparency about purpose, target population, inputs, outputs, performance, failure modes and the role of professional judgment. [3]

  • Verify the exact authorized indication and version.
  • Check whether local data, devices and workflow match the validation setting.
  • Do not convert a wellness correlation into a diagnostic claim.

Connected data needs semantic fidelity

A technically connected record can still be clinically misleading if terms, units, time points or context are inconsistent. WHO SMART Guidelines translate evidence-based recommendations into software-neutral workflows, data dictionaries, decision logic and indicators. This approach helps preserve clinical meaning as recommendations become digital pathways. [4]

Women's health records often cross primary care, fertility, obstetrics, mental health, oncology and specialist services. Interoperability should therefore include provenance, consent status, medication and allergy context, pregnancy timing, laboratory units and the ability to correct data. The IMDRF framework asks whether software has a valid clinical association, correctly processes its inputs, and achieves clinically meaningful performance in the intended setting. [5]

  • Record where a data point came from and when it was measured.
  • Preserve uncertainty rather than forcing ambiguous information into a definitive label.
  • Test exchange with the actual clinical systems and terminology sets used locally.

Bias is a lifecycle property

A representative development dataset does not permanently eliminate bias. Performance can change when disease prevalence, devices, clinical practice or user behaviour changes. Subgroups with smaller samples may have wider uncertainty even when overall accuracy appears strong. WHO's inclusiveness principle and FDA transparency guidance both support subgroup reporting, documented limitations and ongoing performance monitoring. [1][3]

For women's health, evaluation should consider age, pregnancy status, menstrual or menopausal stage when relevant, race and ethnicity, skin tone for optical sensors, language, disability, comorbidity and access to technology. These variables should be chosen because they can affect inputs, care or outcomes, not added as decorative demographics. Monitoring needs predefined thresholds, an escalation route and the ability to suspend a model.

  • Report sensitivity, specificity and calibration with confidence intervals.
  • Assess missingness and device failure by relevant group.
  • Monitor automation bias, override patterns and delayed care, not only algorithm accuracy.

Procurement should demand evidence and control

A high-performing model can still fail if it interrupts workflow, creates alert burden or lacks a safe route for uncertain cases. Procurement should examine clinical association, analytical performance, prospective clinical performance, human factors, cybersecurity, privacy, interoperability and economic impact. The NIST AI Risk Management Framework provides a voluntary structure to govern, map, measure and manage AI risk across the lifecycle. [6]

Convergence becomes clinically valuable when each component has a necessary role and the combined pathway is tested end to end. Adding blockchain, a large language model or a quantum label does not strengthen evidence. The appropriate question is whether the complete system improves a patient-important or service outcome compared with current care, for whom, under what conditions, and with what residual risk.

  • Require a model card, data sheet, intended-use statement and change log.
  • Contract for incident reporting, post-deployment monitoring and safe rollback.
  • Use prospective evaluation before scaling beyond the validated context.
  • Publish known limitations in language patients and clinicians can understand.

What the evidence cannot yet answer

  • Regulatory authorization in one jurisdiction does not establish approval elsewhere.
  • Retrospective accuracy does not prove benefit in routine clinical care.
  • Subgroup estimates can be unstable when sample sizes or event counts are small.
  • AI product performance may change after software, hardware, population or workflow changes.

Questions worth taking into care

  1. What exact clinical decision does the technology support?
  2. Has it been prospectively evaluated in a population like ours?
  3. What happens when the system is uncertain, unavailable or wrong?
  4. Can patients and clinicians see provenance, limitations and change history?
  5. Who monitors performance and has authority to pause deployment?

Source record

Evidence used in this review

Sources were selected for clinical authority, methodological relevance and traceability. Links open the original guidance, public-health record or research publication.

  1. [1]
  2. [2]
    Artificial Intelligence-Enabled Medical Devices

    US Food and Drug Administration · 2026

  3. [3]
    Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles

    US Food and Drug Administration, Health Canada and MHRA · 2024

  4. [4]
    SMART Guidelines

    World Health Organization · 2025

  5. [5]
    Software as a Medical Device

    U.S. Food and Drug Administration · 2025

  6. [6]
    Artificial Intelligence Risk Management Framework

    National Institute of Standards and Technology · 2023

Editorial standard

This evidence synthesis is for general information. It does not diagnose a condition or replace care from a qualified health professional. Treatment choices depend on individual history, examination, local guidance and informed preference. Emergency or rapidly worsening symptoms need urgent local medical assessment.