Artificial Intelligence for Organic Synthesis: From Retrosynthetic Planning to Reaction Prediction and Experimental Decision-Making
S. Farook Basha *
PG and Research Department of Chemistry, Jamal Mohamed College (Autonomous), Affiliated to Bharathidasan University, Tiruchirappalli – 620 020, Tamil Nadu, India.
S. Peer Basha
PG and Research Department of Computer Science, Jamal Mohamed College (Autonomous), Affiliated to Bharathidasan University, Tiruchirappalli – 620 020, Tamil Nadu, India.
H. Vajiha Banu
PG and Research Department of Microbiology, Jamal Mohamed College (Autonomous), Affiliated to Bharathidasan University, Tiruchirappalli – 620 020, Tamil Nadu, India.
R. Arulnangai
PG and Research Department of Chemistry, Jamal Mohamed College (Autonomous), Affiliated to Bharathidasan University, Tiruchirappalli – 620 020, Tamil Nadu, India.
H. Asia Thabassoom
PG and Research Department of Chemistry, Jamal Mohamed College (Autonomous), Affiliated to Bharathidasan University, Tiruchirappalli – 620 020, Tamil Nadu, India.
*Author to whom correspondence should be addressed.
Abstract
Artificial intelligence (AI) is increasingly embedded across organic synthesis, from choosing disconnections to predicting products, conditions and yields, and from selecting experiments to controlling automated platforms. Yet performance at one computational layer does not necessarily translate into reliable synthetic practice. This critical narrative review evaluates AI for organic synthesis as an integrated decision stack rather than as a collection of isolated benchmark tasks. Literature published from 1 January 2010 to 10 July 2026 was searched across multidisciplinary, biomedical, engineering and scholarly indexing sources, with seminal earlier work retained when needed for context. The evidence indicates that AI has matured most convincingly in bounded inference problems where large reaction corpora provide close precedents: single-step retrosynthesis, major-product prediction, reaction classification and local optimisation. Multi-step planning has also progressed through combinations of learned policies and graph or tree search. However, benchmark success remains only an imperfect proxy for chemical utility because commonly used datasets over-represent successful reactions, incompletely encode conditions and work-up, contain duplicated or closely related chemistry, and often favour random splits that reward interpolation. Yield and selectivity prediction are particularly sensitive to dataset design and out-of-distribution shifts. Closed-loop optimisation and robotic synthesis provide stronger evidence of practical value because predictions are tested experimentally, but published demonstrations remain narrower in reaction scope, hardware compatibility and scale than the breadth implied by general AI benchmarks. Large language models and tool-using agents improve orchestration and natural-language access, while introducing additional risks of overconfidence, hidden tool failures and poorly calibrated reasoning. The most defensible near-term model is therefore augmented synthesis: machine-generated options, uncertainty-aware prioritisation, structured experimental execution and expert oversight. Progress towards trustworthy autonomy will depend less on marginal benchmark accuracy than on provenance-rich reaction data, realistic prospective evaluation, uncertainty propagation across the full synthesis stack, machine-readable procedures, and explicit validation of route feasibility, robustness, safety and reproducibility.
Keywords: Computer-assisted synthesis planning, retrosynthesis, reaction prediction, reaction optimisation, machine learning, autonomous laboratories, large language models, chemical robotics