By the mid-18th century, chemistry was still based around the traditional four elements – earth, water, air, fire – until it was revolutionised following Antoine-Laurent Lavoisier’s discovery of the role of oxygen in combustion. This developed into experimental chemistry; now, chemistry is experiencing another revolution instigated by the rise of machine learning.

The question is, can machine learning become so accurate, successful and efficient that we can abolish experimental chemistry entirely and replace the human chemist?

What are the key principles of experimental chemistry?

At its core, experimental chemistry involves conducting laboratory experiments to analyse chemical properties and reactions, involving the treatment of experimental errors and the application of statistical methods to interpret data. Conversely, machine learning (ML) is a type of artificial intelligence (AI) focused on algorithms, able to learn the patterns of training data to make accurate inferences about new data. Thus, it does not require constant human input or hard-coded instructions.

The question at hand is whether experimental chemistry will be made redundant, and there are two key types of machine learning that we need to understand:

• Reinforcement machine learning (RL): where the machine learns to make decisions based off its environment, such as the machine learning via doing experiments and adjusting the method based on the outcomes

• Unsupervised machine learning (UML): where the machine utilises algorithms to analyse and cluster data in order to discover hidden patterns or data groups.

The human chemist could only be replaced if machine learning can perform the three key stages of the experimental process without human input: coming up with hypotheses and designing the experiment; executing and refining the experiment; and analysing and interpreting the data collected.

If a machine can accurately, successfully and efficiently do all three to make new discoveries, rendering the human chemist useless, then the era of experimental chemistry could come to a close and the era of autonomous chemistry begin.

Undoubtedly, machine learning is rapidly enhancing and optimising experimental chemistry at a rate never before seen. Whilst chemists used to spend weeks finding just one reaction pathway to synthesise a target molecule, they now spend mere hours choosing the best reaction pathway out of the many that the ML has provided them.

Experiments can easily be refined through the use of reinforcement learning techniques such as Monte-Carlo Tree Search (MCTS). MCTS has four key phases: selection, expansion, simulation and back-propagation.

In selection, the ML utilises the rule of Upper Confidence Bounds in order to balance the atom economy of a reaction with its precedent reliability (its likelihood of success) but also trying new pathways that a chemist hasn’t tried before. Once the selection reaches a “leaf” that isn’t terminal (i.e. the first step in the synthesis could be successful), it expands to attempt to complete a full reaction pathway and simulates the experiment. Once the experiment reaches a terminal stage – a success or a failure – it backpropagates up the tree.

If the trial was a failure, it updates all the nodes of the pathway it just took during the selection phase and refines the experiment to ensure a greater chance of success on the next trial. If the trial was a success, it relays the data to either a chemist or a robot who can then carry out the practical experiment.

However, how RL really enhances and optimises the process is by still continuing to refine the experiment after a success so that the true “best” reaction pathway for synthesis is chosen.

Whereas a human may settle after weeks of trying for a pathway that has a less than optimal atom economy, the ML can more efficiently discover a superior pathway, saving saves both time and money. All in all, the machine’s reinforcement learning with MCTS can trial roughly 40 pathways a minute so multiple highly successful synthesis pathways can be discovered in just hours.

So, is the human chemist really needed for refining experiments in the modern day?

In my view, reinforcement learning has reached a standard of accuracy and efficiency that renders the human obsolete.

Nowhere is this more evident than Berkeley’s Autonomous Lab which looks to identify materials that could help enhance research areas such as fuel cells and clean energy technologies. The A-Lab works as a closed-loop lab where there is no human interference whatsoever at any point: the AI selects a target molecule; robots synthesise them via adapted MCTS algorithms; x-ray diffractometers and automated electron microscopes analyse the molecule; and the data is then fed back into the

machine so it can learn and refine the experiment or alert the researchers that the target molecule has the desired chemical properties.

This closed-loop system without human intervention highlights how reinforcement learning is the future when it comes to experimental refinement and execution. The AI can process a staggering fifty to a hundred times more samples in a day than an equivalent human research team. With that in mind, it would certainly seem as if, with respect to the stage of executing and refining the experiment, we can remove humans from the process, thus replacing experimental chemistry and moving to an age of autonomous chemistry.

But can machine learning independently design experiments?

In my view, no matter how powerful the machine, the human always has to design the experiment. Even in MCTS synthesis, the target molecule that the machine is refining the experiment to synthesise has to be chosen by a human. However, it could be argued that the materials project (developed by Dr Kristin Persson at Berkeley Lab) displays evidence that, in fact, a machine can design an experiment and the target molecule for synthesis.

The materials project utilises density functional theory – DFT0, a form of computer modelling, in order to design and predict the properties of a molecule before inputting them into a large database. This then allows labs such as the A-Lab to practically test the computed properties from the library of over 570,000 molecules in order to fast-track materials for several research areas. Since the A-Lab is closed-loop and does not require any human intervention, it appears on the surface that machine learning can both design and execute experiments relating to molecular synthesis without the need for humans.

However, in reality, this is nowhere near the case.

While DFT computer modelling has greatly enhanced the field of material chemistry, it still is prone to designing materials with erroneous geometries or molecules that would be so costly to synthesise that, despite MCTS refinements, they’re virtually impossible.

As a result of this being inputted into the materials, before project database, the molecules’ properties are validated by scientists against benchmark sets to ensure structural, electronic and vibrational properties are accurate. Since computer modelling isn’t actually a machine learning program, humans then have to update the program manually rather than the machine reinforcing its learning. Furthermore, in order for closed-loop labs to test a certain molecule, a human must input instructions for the machine to test materials with a certain theorised property.

This evidently is still miles away from removing humans from the process and abolishing experimental chemistry.

The materials project and the A-Lab demonstrates that without humans in the process our experimental procedures would be flawed, costly and, in some cases, impossible. Computer modelling and machine learning, as it stands, are tools to aid the experimental chemist, not a replacement. Overall, when it comes to designing experiments and coming up with chemical hypotheses, machine learning does not yet have the required accuracy to facilitate a removal of human intervention and, thus, experimental chemistry cannot be abolished.

With previous data analysis techniques (such as least-fitting squares) having high computing costs, ML has become very prevalent in the field to minimise costs and speed up analysis. Already it can analyse the structure, the spectra and the structural parameters of a molecule. However the large majority of labs utilise supervised learning for data analysis which uses labelled data sets to train the machine so that they can analyse real-world data.

The accuracy and reliability of supervised ML data analysis almost entirely depends on the range, diversity and quality that the human researcher chooses to train it on; it cannot train itself. While it is often very reliable when confronted with labelled data, supervised ML falls apart entirely when experimental data was not in the training. This is particularly problematic when we look to make fundamentally new discoveries in chemistry which the machine will not have encountered before. Thus, if supervised learning is to continue as the primary tool for data analysis, the human must remain in the loop and we cannot abolish experimental chemistry.

However, the future of ML data analysis does seem to indicate a shift to unsupervised learning analysis which does not require human input. Techniques such as k-means clustering have become prominent in spectra analysis and are often very reliable, without needing human supervision like DFT and supervised ML. K-means clustering works by grouping all the data into clusters where all data points in a cluster similar and all clusters are distinct from each other. For example, k-means would be able to easily group similar IR or NMR spectra whereas a human may find it more challenging to distinguish between functional groups in the 500-1500 fingerprint region for IR spectra.

In other areas of data analysis, however, k-means can sometimes encounter problems with outliers by trying to overfit the clusters to include them. This can be resolved by utilising gaussian mixture models but current k-means clustering technology does still mean supervised ML is still more reliable.

In the near future we may well see a combination method whereby unsupervised ML clusters complex data sets in order for supervised ML systems to analyse them, thus solving the problem of supervised ML struggling to handle data sets originating from a mix of chemical phases. However, in its current form, machine learning does not have

the accuracy and reliability in order to remove humans from the loop so, unequivocally, we cannot abolish experimental chemistry.

To conclude, machine learning presents immense utility for an experimental chemist, particularly with regards to refining experimental procedures, but fundamentally lacks the accuracy for experimental design and data analysis to trigger a removal of humans from the experimental loop.

As the A-Lab’s Professor Gerbrand Ceder put it, Lab research has been the same for the last 70 years: the equipment may have gotten better, but ultimately a person is needed to take measurements, analyse results, and decide what to do next. Therefore, despite machine learning, we simply cannot expect experimental chemistry to be abolished.

Leave a Reply

Trending

Discover more from The 1509

Subscribe now to keep reading and get access to the full archive.

Continue reading

Website by Theo Damaskos