AI is entering a high-stakes laboratory
AI in reproductive medicine is being studied for embryo image assessment, oocyte grading, sperm analysis, laboratory witnessing, cryostorage management and treatment prediction. These applications differ in both evidence and consequence. An administrative tool that reduces manual logging is not equivalent to a model that ranks embryos for transfer. The American Society for Reproductive Medicine describes implementation in the IVF laboratory as early-stage and recommends caution when interpreting largely retrospective data. [1]
Embryo selection is especially sensitive because the clinically meaningful outcome is not an attractive score or improved agreement between embryologists. It is a healthy live birth while minimizing treatment burden, multiple pregnancy and avoidable harm. Any model used in selection should be assessed within the complete treatment pathway, including patient age, laboratory practice, transfer strategy and the possibility that several embryos may be suitable.
- Separate workflow automation from clinical prediction.
- Define whether the model supports or replaces embryologist judgment.
- Prioritize cumulative live birth and safety over surrogate image scores.
Time-lapse imaging is not the same as proven AI benefit
Time-lapse incubators repeatedly photograph embryos without routine removal from the incubator. Software may then apply morphological rules or machine learning to those images. A Cochrane review of nine randomized trials involving 2,955 couples found uncertainty about differences in live birth or ongoing pregnancy for several comparisons. For one comparison, a conventional live-birth or ongoing-pregnancy rate of 35% corresponded to an estimated range of 27% to 40% with time-lapse imaging and conventional assessment, reflecting low-certainty evidence. [2]
These findings do not prove that every newer algorithm is ineffective. They show why device configuration, algorithm version and prospective trial design matter. A retrospective model can learn patterns associated with implantation in historical data yet fail to improve patient outcomes when used to choose an embryo. Changing laboratory conditions, image systems or patient mix can also reduce transportability.
- Ask whether evaluation was prospective and whether the algorithm influenced care.
- Check whether the comparator reflects current expert embryology practice.
- Demand version-specific performance and external validation.
Clinical validation must follow the intended claim
If a system claims to predict blastocyst development, that endpoint can test the claim but cannot establish improved live birth. If it claims to prioritize euploid embryos, validation requires an appropriate reference and an explanation of how testing error is handled. If it claims to increase live birth, a prospective comparative study should measure live birth and relevant harms. ASRM notes that a large randomized trial did not demonstrate noninferiority of algorithmic embryo selection to standard morphology for clinical pregnancy and calls for more prospective evidence. [1]
Reproductive AI trials should follow established reporting standards. CONSORT-AI asks investigators to describe the AI intervention, user skills, integration setting, input and output handling, human interaction and error cases. These details are essential because a model can change behaviour even when its numeric output is unchanged. [5]
- Prespecify the primary outcome and subgroup analyses.
- Report failed outputs, overrides and disagreements with embryologists.
- Use confidence intervals and calibration, not accuracy alone.
Fairness requires more than demographic balance
Embryo datasets may reflect a narrow group of clinics, incubators, protocols and patients. A model may inadvertently use image features linked to laboratory conditions rather than embryo biology. Performance should be tested across sites, equipment, age groups, ovarian response patterns and clinically relevant populations. WHO's AI guidance calls for autonomy, safety, transparency, accountability, inclusiveness and sustainability. [6]
Consent and communication also matter. Patients should know when an investigational algorithm materially influences ranking or treatment, what data it uses, whether images contribute to future model development, and what alternatives exist. A proprietary score should not prevent a clinic from explaining the basis, limitations and uncertainty of a recommendation. FDA transparency principles similarly emphasize intended use, training and testing data, known failure modes and lifecycle monitoring. [4]
- Evaluate data provenance and permission for secondary model development.
- Test whether performance differs by clinic, device or patient group.
- Preserve a meaningful route for human review and challenge.
A careful adoption standard
A clinic considering reproductive AI should first determine regulatory status and whether the proposed use matches the authorized or validated indication. It should then conduct local acceptance testing, train users, document responsibility, monitor outcomes and compare performance with the pre-implementation baseline. ESHRE's good-practice recommendations on IVF add-ons emphasize transparent information and evidence-based evaluation before routine adoption of adjunct technologies. [3]
The appropriate position is neither automatic rejection nor premature enthusiasm. AI may improve consistency, laboratory quality and decision support, but evidence should rise with the stakes of the claim. Patients should not pay for an implied improvement in live birth when only retrospective association or laboratory efficiency has been demonstrated.
- What patient-important outcome has improved in a prospective comparison?
- Is the evidence independent and specific to this version?
- How are model updates validated before use?
- Can the clinic identify and respond to subgroup performance drift?
What the evidence cannot yet answer
- Much embryo-selection research is retrospective and vulnerable to selection bias and site-specific confounding.
- Time-lapse incubation, time-lapse morphology and AI selection are distinct interventions.
- Implantation and clinical pregnancy are not substitutes for cumulative live birth and long-term safety.
- Rapid model updates can make published evidence obsolete for the deployed version.
Questions worth taking into care
- Is AI being used for administration, quality control or a clinical selection decision?
- Was the exact version evaluated prospectively against current care?
- Does evidence include live birth, safety and cumulative treatment outcomes?
- How are patients informed and how can clinicians override the output?
- What happens to embryo images and related data after treatment?
Source record
Evidence used in this review
Sources were selected for clinical authority, methodological relevance and traceability. Links open the original guidance, public-health record or research publication.
- [1]Artificial intelligence in the in vitro fertilization laboratory: a committee opinion
American Society for Reproductive Medicine · 2026
- [2]
- [3]Good practice recommendations on add-ons in reproductive medicine
European Society of Human Reproduction and Embryology · 2023
- [4]Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles
US Food and Drug Administration, Health Canada and MHRA · 2024
- [5]CONSORT-AI extension
Nature Medicine · 2020
- [6]Ethics and governance of artificial intelligence for health
World Health Organization · 2021
This evidence synthesis is for general information. It does not diagnose a condition or replace care from a qualified health professional. Treatment choices depend on individual history, examination, local guidance and informed preference. Emergency or rapidly worsening symptoms need urgent local medical assessment.



