OpenAI and Molecule.one reported on June 17 that a GPT-5.4 workflow helped identify conditions that improved a difficult Chan-Lam coupling across 10,080 microscale reactions. An automated laboratory supplied the testing capacity, while chemists selected proposals, corrected plans and repeated key experiments by hand.
The result is evidence for a tightly coupled research workflow, not an autonomous scientist. Its strongest contribution was the ability to screen a large matrix of conditions and feed measurements back into another experimental cycle. The public record does not show that the model outperformed chemists on the same task or that the method works beyond this reaction.
A model proposal became two rounds of physical testing
The project began with an open-ended goal of improving a reaction class used in process chemistry. GPT-5.4 proposed focusing on the Chan-Lam coupling of primary sulfonamides with boronic acids and suggested mild oxidants such as TEMPO, a stable aminoxyl radical, as possible additives.
OpenAI says scientists wrote the steering and grading prompts, reviewed the highest-ranked proposals and selected four for laboratory testing. Molecule.one's Maria system converted the selected plans into detailed instructions, ran the experiments and returned structured results to the model. Human chemists made limited corrections, including removing a solvent that they thought could react with stronger oxidants used for comparison.
TEMPO improved yields across a defined substrate matrix
Chan-Lam coupling uses copper to form bonds between organoboron compounds and molecules containing nitrogen, oxygen or sulfur. The primary-sulfonamide version is difficult because these compounds are weak nitrogen nucleophiles and the boronic-acid partner can degrade before forming the intended carbon-nitrogen bond.
The researchers tested 12 primary sulfonamides and eight boronic acids in two high-throughput campaigns totaling 10,080 reactions. The reactions ran in 96-well plates at a volume of 28.6 microliters, with the team varying the oxidant, copper source and loading, base, solvent, temperature and substrate structure.
The optimized TEMPO condition raised the mean estimated product yield from 16.6 percent to 25.2 percent. The share of reactions exceeding 30 percent yield rose from 15.6 percent to 37.5 percent. TEMPO also reduced the estimated formation of products associated with oxidative degradation of the boronic acid, although the authors said further work was needed to establish the mechanism.
Bench tests narrowed the claim
The screening figures were estimates from a high-throughput analytical method, so the team repeated selected reactions at a larger scale. Fourteen representative substrate pairs were tested manually, and 11 produced more product with TEMPO than without it. Eight showed increases of more than twofold.
The paper reports detailed quantitative nuclear magnetic resonance yields for four of those pairs. TEMPO raised the yield in three, from 10 percent to 33 percent, 27 percent to 59 percent and 36 percent to 99 percent. It reduced the fourth result from 49 percent to 42 percent, illustrating that the effect was not universal.
Another comparison held the copper loading constant across 10 boronic acids. TEMPO helped seven electron-poor boronic acids and one electron-rich example, while two other substrates showed no benefit or produced less product. The method therefore describes a substrate-dependent improvement rather than a general solution to low-yield Chan-Lam chemistry.
The advance came from throughput and feedback
The project turned literature review and hypothesis generation into measurements at a scale that would be slow for a chemist working through reactions one at a time. The system could compare 10 oxidants and many combinations of operating conditions, then use the first campaign to focus the second. That loop made the search tractable and exposed where the leading condition failed.
Calling the result an intelligence breakthrough would go beyond the evidence. The experiment did not compare the model with expert teams given equal laboratory access, time and budgets. It also did not separate how much of the result came from GPT-5.4, Maria's chemistry software, the experimental design, the automated equipment or human selection.
The paper notes that TEMPO had appeared in early and later Chan-Lam research, but had received little systematic attention for this use. The reported novelty is a broad evaluation of TEMPO in primary-sulfonamide coupling, not the invention of TEMPO or its first appearance anywhere in Chan-Lam chemistry.
Independent replication and scale remain open
The authors are affiliated with Molecule.one and OpenAI, the organizations that developed the system and announced the result. OpenAI says four outside chemistry experts reviewed the preprint, but expert review is different from an independent laboratory repeating the experiments. APPI News could not find a published independent replication at the time of writing.
The experiments also stop short of manufacturing evidence. Most of the dataset came from microliter-scale screening, while the manual validation used milligram-scale reactions and a limited set of substrates. Work on reaction mechanism, broader substrate coverage, reproducibility under other laboratory conditions and process-scale performance remains unfinished.
For research organizations, the case supports investment in the full testing loop: structured experimental data, automated equipment, analytical methods and chemists with authority to reject or revise a proposal. A model can expand the number of hypotheses considered, but the finding becomes credible only when physical measurements expose both the gains and the failures.