Webinar behind the scenes, making sure we don't have any technical difficulties. In this webinar, Dr. Jeremy Jones and Dr. David Miller will walk you through some of the newest enhancements of ADMET Predictor, which enables researchers like you to design therapies for complex and challenging targets. They'll describe several new predictive models, updated existing models that better cover the beyond rule of five space, and key improvements to the High-Throughput Pharmacokinetics simulation model, command line and REST API, and a new command line version of ADMET Modeler.
They will also provide a first look at our new AI-enabled platform, which combines a natural language prompt interface with the ADMET Predictor engine. We'll also be joined by Rafał Bachorz, who will give a short demo of Composer, our new operating system for AI-enabled MIDD. Our Scientific Product Specialist, Gül Güryel, will serve as our moderator today. A few housekeeping items before we begine We take your privacy rights seriously. By registering for and attending this event or participating in the Q&A session, you are allowing us to contact you for follow-up.
You may ask questions via the Q&A panel on your dashboard at any time. If you need assistance, please use the hand raise icon. We'll address all questions during the Q&A after the presentation. Before we begin, we'd like to conduct a short poll. This first poll has two questions. What models should we prioritize for update? I do see some answers coming in. Couple more seconds. I'll go ahead and end the poll there. Thank you everyone for your responses. Now I'll hand it over to Gül to introduce our speakers.
Thank you, Jasmin. Welcome everyone. I'm pleased to introduce our presenters today, Dr. David Miller, Dr. Jeremy Jones, and Dr. Rafał Bachorz. Our first speaker today is Jeremy, who received his PhD in microbiology and immunology from Stanford University and completed his postdoctoral fellowship at the University of California, San Francisco. Jeremy joined Simulations Plus in 2022 as a Principal Scientist focused on early-stage drug development and AIDD. Prior to joining the company, Jeremy was the Director of Oncology at Trethera Therapeutics and led his own consulting company. Before that, he was a professor at City of Hope Comprehensive Cancer Center for nearly a decade, where his lab focused on drug development and translational research in urologic oncology. David is Vice President of Cheminformatics Solutions at Simulations Plus. David started his career studying physics at Stanford University, followed by his PhD in biophysics at the University of California, San Francisco.
After his studies, he worked at CombiChem, then at Sage Informatics, a start-up cheminformatics software company that was acquired by Simulations Plus in 2005, where he has been since. David is one of the lead software engineers for the ADMET Predictor software platform. Our last speaker today is Rafał. He received his PhD from the Poznań University of Technology, where he also completed his academic training. He joined Simulations Plus as a Senior Principal Applied Scientist, where his work spans cheminformatics, data science, and machine learning for pharmaceutical R&D. Prior to this role, he spent over a decade at PSI Software leading advanced analytics and a competence center for high-availability energy systems. Alongside his industry work, Rafał serves as a lecturer at Poznań University of Technology. Without further ado, I will let Jeremy, David, and Rafał take you through what is new today. Thank you.
Jasmin, we don't have the audio.
Okay. Sorry about that. Let me try one more time. Okay. Hopefully, this works.
Hi, I'm Jeremy Jones, Principal Scientist at Simulations Plus. First apologies for pre-recording my portion of today's webinar. I live in Auckland, New Zealand, so it's the middle of the night for me. But rest assured, the rest of the presenters will be here live and have the ability to answer any questions you might have. That said, I am excited to start off the webinar to tell you a little bit about what's new in ADMET Predictor 14. I'll be telling you about our work with the beyond rule of five molecules, developing new descriptors and some models that encompass that chemical space. David Miller will be along shortly to tell you about new features in our High-Throughput PK module, and some new capabilities in command line and REST API, and some other cool stuff.
You're in for a treat because at the end, you'll get to see a demonstration of our new agentic front for our software called S+ Composer. I'm going to dive right into the rule of five stuff. Just to set the framework here, I think most of us are familiar with Lipinski's rule of five, but just as a refresher, this is a set of rules of thumb, not hard and fast rules, that when looking back at drugs that had progressed successfully through the clinic, some features of those compounds, that made them amenable to development. Those rules are a molecular weight of 500 or less, a logP of five or less, and these limits for the number of hydrogen bond donors and acceptors. Those rules do work pretty well.
But over time, we have advanced our chemistries and advanced our different modalities, and have begun to develop things like PROTACs and larger cyclic peptides, and we have evidence that some of these drugs can also show oral bioavailability and decent pharmacokinetics, and can make effective drugs. Here you see, I love this kind of diagram. In red, you can see the boundaries of the classic rule of five in a small molecule like ibuprofen fits wholly within those parameters, and you start to get slightly larger molecules, atorvastatin, one of which I am certainly familiar with and probably many of you too, mostly fits in within those parameters. But it started to expand.
And as we get bigger and bigger in here, we see an advanced PROTAC and even an older molecule, a big macrocycle, rifampicin, which are clearly well outside these Lipinski's rules, but still have proven to have decent PK and have proven to be effective drugs. So we have gotten really good at making predictions about the in vitro ADMET properties of small molecules. I think the field has seen in recent years that perhaps our predictions are not quite as good for these novel larger small molecules, the ones that are outside these traditional rule of five parameters. So one of the reasons that our predictions are difficult, especially around solubility and permeability and the things that those influence, is because of something that has been labeled chameleonicity of these compounds. And what that really is the ability for compounds to adopt different conformations in polar and apolar environments.
So for example, a cyclic peptide like this in a polar aqueous extracellular environment, it can adopt a conformation where the electrons on the oxygens are exposed, and can form hydrogen bonds with the solvent molecules. And so you would predict relatively good solubility, but because there is a lot of hydrogen bond donors, acceptors, you would predict poor permeability for such a molecule using our traditional descriptors and understanding of this behavior. However, this molecule can adopt a different conformation as well, one in which those oxygens can fold in and form intramolecular hydrogen bonds, masking those electrons, and really have observed permeability that is much higher than we might predict. Once it passes through that apolar environment back into a more aqueous polar intracellular environment, it can unfold again and form those bonds with the solvent molecules. So this is true of PROTACs as well.
Here you can see two examples of potential long-range intermolecular hydrogen bonds that can mask those electrons and really confound our ability to predict solubility and permeability well. A quick digression here to remind you of the importance of descriptors to machine learning ADMET models. I think many of you are familiar that we offer the ADMET Modeler as a module, as part of ADMET Predictor, which you can build your own neural networks, and various types of machine learning models. When we are thinking about building a model, I think there are certain things that really influence the quality of that model. The first and foremost is data quality. If you have garbage in, you have garbage out. This will be my little plea. We have adopted many AI type tools to help us curate data, but there really is no substitution for manual curation.
Data quality is definitely the most important. Choice of algorithm and how you build the model, how you separate your training and test set, how you protect against over-training, all very important too. But I think one underappreciated aspect are the descriptors, the inputs you use for your machine learning models. I like to show this publication from folks at J&J about two years ago now, where they were actually looking to see what factors most influenced the quality of the models they built for their own intrinsic solubility data. The too long, didn't read was number one was data quality, but the second most important thing was the set of descriptors they used in their model building. Here you can see just a chart, where overall they're comparing using RDKit or Mordred, some open source descriptors, versus the ADMET Predictor descriptors.
You can see all the same data sets, all the same model builds, but different descriptors. ADMET Predictor, and to quote right from the paper, "Models trained with ADMET Predictor descriptors performed best." So we do have this set of advanced proprietary descriptors that consistently outperform open source descriptors, and that's something I think we're really proud of. That being said, that chameleonicity that I just talked about, the way it influences solubility and permeability, really makes it difficult to capture the behavior with traditional descriptors like logP and one that people love to use, topological polar surface area. We realized this as we were collecting data on beyond Rule of Five molecules and trying to incorporate them into some of our models that our predictions weren't as good as we had hoped.
We set out to actually create several new descriptors that could potentially capture this chameleonic potential. Here you can see a list of some of the descriptors that we've added to ADMET Predictor 14. Most of them have to do with the potential to form these longer range intramolecular hydrogen bonds, right? The distance and the strength of those. We've also thrown in a macrocycle descriptor, just one we didn't have. Of course, these aren't molecular dynamic simulations. These are more static descriptors that can be calculated very quickly. As you'll see, the proof is in the pudding, really do capture the potential to form that behavior. As we build models out and using descriptor sensitivity analyses, we've seen that these new descriptors are very influential and among the most important as we're building these new models. So, we have these set of new descriptors.
We also built several models to be able to predict the outcomes of novel assays that have been developed to measure certain properties of these beyond Rule of Five molecules. The first of these is called EPSA, sometimes called exposed polar surface area or experimental polar surface area. What this is this chromatography technique that quantifies this experimentally accessible polarity, basically any of the hydrogen bonds or polar residues that are on the outside of the molecule. Interestingly, it has a very poor correlation with calculated TPSA. Here you can see just a graphic I pulled up from Enamine. Enamine and other CROs run this all the time. Most companies that are developing PROTACs or big cyclic peptides run this in-house and measure the EPSA. Just so you can see the data in our hands, you're looking at the calculated TPSA versus the measured EPSA.
The compounds in blue are Rule of Five compliant. The compounds in black are beyond Rule of Five. You can see the poor correlation between the two. The second assay that is now commonly employed is called the ChromLogD assay, and this is a chromatography technique that replaces the traditional shake flask logP measurements just because those measurements are only accurate up to a value of about 4.5, and a lot of these molecules have even higher logP. These two techniques are employed so often in the field now that the folks at AstraZeneca published their Rule of Oral PROTACs that is obviously in reflection of the changes to that old Lipinski's Rule of Five. You can see they use these two assays very heavily to establish the parameters for the oral PROTACs.
A ChromLogD value, an EPSA value about 140, and then related to the EPSA, the exposed rotatable bonds and hydrogen bond donors and acceptors as well. Yes, these two assays are now pretty standard in the field. There's a third assay, which is, I think, gaining more attention too, that was developed by an academic group, you can see the paper there if you're interested, and it's called ChameLogK. Basically, this is trying to mimic what happens in a biological setting for these compounds, where they start off in a more polar environment and go to a more apolar environment as they pass through a cell membrane. This is another chromatography technique, where the column actually starts off very polar and has a gradient where it becomes very apolar.
You're measuring the difference from what you would extrapolate in the polar section to what you actually see in the apolar section. If you're interested, please read the paper. It's a cool new technique that I know that several CROs are now running this assay as well. As I said, these assays are now run routinely, but like any in vitro assay, we'd like to not have to run them all the time. We'd like to at least be able to potentially predict the outcome of those assays, so we could prioritize compounds for testing. Right? Of course, we set out to build models to predict the outcomes of these assays, and here you see the outcomes of those. In the top, you can see a distribution of the sizes of the molecules versus the measured EPSA or ChromLogD or ChameLogK.
Again, blue compounds are Rule of Five compliant, black are beyond Rule of Five. You can see the model performances. Just like any model we make, if you ever look in the ADMET Predictor, you'll see the data set sizes, the test and train, and the performance statistics of all of ours. I think we have a decent training and test set sizes for the EPSA and ChromLogD and models that perform really well. The ChameLogK, currently we only have the 56 data points from that academic paper. I would say this model, while performing well with a small size, is a work in progress and we'll keep updating it as more data is published, but still a valuable tool to have. We created two other models, a PROTAC filter and a cyclic peptide filter.
Now, I realize that recognizing what a PROTAC is and a cyclic peptide is pretty easy, but when we often sort through very large data sets and would like to be able to pick these out very quickly and sort by them, we developed these classification models, and you can see perform almost perfectly as you might expect from something that's structurally so obvious. The other point here was sort of a proof of principle. When we were building these models, we looked at the descriptors that were most influential in these models, and they almost always, those new descriptors that we developed. We also used the predictions from the EPSA, the ChromLogD, and the ChameLogK as descriptor inputs as well. All of these were always among the most influential. The top 10 were mostly populated by these.
So really showing the importance of these new descriptors and being able to capture what it is that makes these beyond Rule of Five molecules different. All right. We took those descriptors and the predictions from those models, and we tried to gather as much data from PROTACs and cyclic peptides and big macrocycles as we could to add to our data sets and rebuilt many of the models of our ADMET models. Our first pass here was really focusing on the ones that would feed into our High-Throughput PK. That includes things like rebuilding the liver microsome intrinsic clearance, as well as the fraction unbound in microsomes, the fraction unbound in plasma, the blood to plasma ratio, five different species for those.
We rebuilt the intrinsic water solubility models, the S+Peff jejunal permeability model that we use as standard in HTPK, as well as several different other cell-based and PAMPA permeability models. Some of these are brand new, and I'll tell you a little bit about those as well. David will tell you more about this in the HTPK section, but bringing in the ability to scale some of these to our jejunal permeability so you can use them in HTPK. We also updated our ADMET risk score, which I want to tell you about in a second, too. As I said, full performance characteristics of all of our models can be found in the manual. I just wanted to show you, highlight some of the improvements from the models we had in AP13 to these new beyond Rule of Five inclusive models in AP14.
These are some molecules cherry-picked that weren't in the training set of either the models for AP13, the old S+Sw or this permeability model, or in the new AP14 models. In blue, you're looking at the measured solubility. In orange, you're looking at the predicted solubility from AP13. In green is the predicted solubility from the AP14 model. Traditional small, Rule of Five compliant molecules, our predictions are about the same, equally good, and that's generally true throughout the data sets. But there were definitely several beyond Rule of Five molecules where we had missed severely in AP13, a big macrocycle here, a PROTAC here. You can see with our new models, still not perfect, but certainly way better. This one got obviously really good.
Our predictions got significantly better for these beyond Rule of Five molecules, and the same pattern holds for MDCK-LE permeability and PAMPA permeability as well. You can check out the statistics in the manuals as well. I did want to mention our ADMET risks model. We have an ADMET risk score. A lot of folks have a drug ability score, a drug likeness score. I think of our ADMET risk score as sort of our version of that. It incorporates many of our different models, such as our logP, our permeability, our solubility model, our Vd for distribution, several most important tox and metabolism models. Closer to zero means a lower risk, so better for developability. What you will notice is that at the top here, some of the rules are the classic Lipinski type rules.
As my colleague Robert noticed, these rules are what he called hostile to these beyond Rule of Five molecules. What you can see here is the weight distribution versus the ADMET risk score, higher score, higher risk, so worse molecules. In green, you see highlighted the PROTACs in this data set with all high scores. We went back and modified these rules. They included call-outs for PROTACs and macrocycles. If they were, we brought in some of the predictions, ChameLogK and EPSA, so that PROTACs and other beyond Rule of Five molecules were not penalized as much. The updated rules are much more friendly to PROTACs than these other beyond Rule of Fives.
Just in general, if you have been using a classic drug ability or drug-like score, and you are trying to work with these bigger beyond Rule of Five molecules, it might behoove you to check out the rules and update them like we have done here as well. I will finish with just talking a little bit more about that HTPK and how the new models lead to better PK predictions for the beyond Rule of Five molecules. I think many of you are familiar with our HTPK, but in case you are not, this is our modified version of the full advanced compartmental absorption and transit model that is available in GastroPlus. It is slimmed down a little bit, so it runs a lot faster, and you can simulate the in vivo PK of thousands of compounds in minutes.
You can produce full CP time curve profiles along with the important endpoints, oral bioavailability, AUC, et cetera. You can calculate the optimal dose and schedule. It is available for three different routes of administration in five different species. This is I think the best way to do this is this hybrid approach, where you use machine learning models, or if you have the measured data, of course, as inputs into this mechanistic modeling and simulation of the in vivo PK. You saw we rebuilt the models that are important here, the solubility, the permeability, the fraction unbound in plasma, et cetera. We set aside several beyond Rule of Five molecules that were never touched by any of our machine learning models, just to see if our PK simulations improved. I will give you just one example here.
This is human data from our great friends at Nurix. I always thank them for actually publishing their human PK data on their PROTACs. What you can see here on the left is a single dose, either 50, 100, or 200 mg in humans. The dashed lines here are the actual observed values. You can see the Cmax, not really that high, not bad, but much lower than our models. We use the AP13 models as inputs for HTPK. We vastly overpredicted the Cmax so that the AUC was much too high. When we then went back and used the new AP14 models as inputs for HTPK, what you see is something much closer to what was observed.
The Cmax values are not perfect, but much closer, and certainly for an early stage development where you are trying to bring some of this downstream PK information back up into the selection and prioritization process. This degree of error is certainly acceptable. We get the Cmax as much more right. You get the slower elimination looks much more similar. We still miss Tmax a little bit, but the overall curve is much more similar to what was observed in reality. With that, I will conclude my part. Just reminding you here that the traditional descriptors, while great for modeling of traditional Lipinski's rule of five compliant molecules, really are insufficient to describe that chameleonic behavior of beyond Lipinski's rule of five compounds.
Expanding that descriptor space and retraining the models with the explicit inclusion of that beyond Lipinski's rule of five chemistry led to real meaningful improvements in not only the in vitro ADMET models, but the PK simulations as well. To me, what is most important here is that we can bring this sort of PK simulation information earlier into drug design and optimization, even at the point of design. For instance, we have our AI driven drug design module that is generative chemistry plus multi-parameter optimization. You can bring these PK, HTPK simulations as one of the parameters so you can optimize for oral bioavailability in the design of the compound. I think that is the dream. With that, I will stop and turn it over to David. Thank you.
All right. We will pass it over to David.
Okay. Thank you. I take it we're not doing another poll, Jasmin. Let's-
Go ahead.
I was going to say, we can save the poll for after.
Okay, great. We are running a little tight on time here, so I am going to just quickly go over some of the other new features that we've added to version 14. I am going to start with HTPK, the simulation module that Jeremy just introduced. The first enhancement here is what we call an advanced logging feature that's in the graphical user interface. The idea here is that, we hear from users that it becomes difficult to keep track of all the parameter changes that they have made over time in HTPK, particularly when they go back to a project after a couple of months. This new feature allows for the automatic save of settings, inputs, and outputs, for each simulation. You run a simulation, you'll get a set of saved files. There's a timestamp in the filename.
There's a text file that stores all the chemical structures, any in vitro inputs that you used in the simulation, and then all the simulation results. There are two settings files that are saved. One, the .inp file with species or physiology parameters, and then an HIA file. Sorry, let me quickly do my pointer. An .hia file with compound specific parameters. This set of files serves both as a permanent record of what you've done, but also as a means of ensuring reproducibility. Later on, you can open the text file directly in ADMET Predictor, so you have all your chemical structures and the in vitro inputs. Then for the settings files, we've added a new user interface feature to load those settings files directly into a running session. This was something that was rather cumbersome to do in previous versions.
With the settings loaded, you just click Run, and you'll be guaranteed to reproduce the previous results that you had. The text file also saves the source of each HTPK input at the compound level. Using a parameter like solubility as an example, if you've set the simulation to use an experimental solubility preferentially, and then only to fall back to the ADMET Predictor model, called S+SW, in the case of a missing value, that information is retained. So you have not just the solubility values, but where they came from. This is particularly important because this can be a list that's longer than two items. You can have multiple experimental solubilities and multiple models. To enable that logging, there's a Simulation tab in the main settings page.
You enable the archiving, is what it's called, and there's a second option to be prompted each time you run a simulation, so you can add some custom notes, and then those get saved to the log file. The second enhancement is Papp to Peff conversion. I think Jeremy alluded to this. This is a feature that we have taken from our flagship GastroPlus software. It allows you to create a correlation between your in vitro permeabilities and the human jejunal permeabilities that are used in HTPK. Once you've created and saved that correlation, it will be available for use in your future simulations. There are various modes of operation, and you can use the correlation in all of them. This is what the graphical interface looks like. You would select your experimental permeability, and then there's a new dropdown to choose any conversion that you've created.
To create these correlations, you actually do need to use the user interface for the time being. It looks very similar to what you would see in GastroPlus. You are presented with the list of so-called Lennernäs standards. These are the 30-odd compounds for which the human jejunal permeability has been measured. You would pick a subset of these compounds to measure in your assay, then enter the data here, and then using the tools, you'd find the best linear correlation, and then save it with a particular name that you can refer back to later. We've had parameter sensitivity analysis in HTPK for some time. This is what the chart display looks like in the graphical interface. We're looking at one PK endpoint here, this is bioavailability, and then four parameters.
Solubility, permeability, clearance, and fup and looking at how bioavailability is affected as you vary these four parameters independently. What is new in AP14 is that we have enabled PSA to be run on multiple compounds simultaneously, so up to thousands of compounds. Instead of seeing all these charts, all the data gets summarized in a series of plots in the spreadsheet. There is one column that is added for each PK endpoint. Here is an FA column, an FB column, and there are some other hidden ones. Looking at just an individual compound in one plot, which we call a star plot, the wedges represent the parameters. The length of the wedge represents the absolute value of the slope of this chart at the midpoint. We use a solid texture for a positive slope and a kind of a hatched one for negative.
You can quickly see that solubility has the largest positive slope and clearance the largest negative. You can run this on an entire data set and then just quickly scroll through and have all this information be summarized. The command line and the REST API have supported multiple compound PSA for some time now, but this is just bringing that same functionality into the graphical user interface. We have added a few new HTPK parameters. Two of the ACAT variety. Again, that is the sort of physiology species parameters. Then two of the HIA, which are the compound specific. In the first category, we have added amount of microsomal protein per gram liver tissue, and the number of hepatocytes per gram liver tissue. These were previously hard-coded values, but we had some more advanced users who wanted more control to set those, so we have exposed them.
We have added a parameter related to Peff conversion, which I already discussed, and then one related to clearance scaling. It is a multiplier for scaling the intrinsic clearance. The default for this is meant to reproduce the scaling that is used in GastroPlus compartmental simulations. That is exactly how AP13 in earlier versions worked, so you do not need to set this parameter. Again, it is a case of some users who have their own internal scaling that they use, so we have exposed the parameter for that purpose. It is a bit of a complex issue here, and there is a discussion of it, a new section that we added to the user manual that you can consult. Two new HTPK result values. The number and list of HTPK inputs that are outside their model applicability domain.
This applies to the situation that you are using one or more ADMET Predictor models as inputs to HTPK, something like our blood to plasma model or permeability model. This just serves as an alert that for a particular compound, one or more of your inputs is out of scope, so to speak. Okay. Some changes in the command line. Jeremy kind of previewed this, but we are now releasing a command line version of our ADMET Modeler tool. Modeler is the tool that we use for building models, and as Jeremy said, customers can use the tool to build models using their data. In the past, you could only use Modeler by launching it from the ADMET Predictor GUI. Now it is a separate command line tool. It runs on both Windows and Linux, so the Linux is new.
This does make it quite a bit easier to explore model parameter space just by writing a small script around the command line executable. We've introduced a few new, what we call script file workflows. These are workflows other than just predicting ADMET properties or running pharmacokinetic simulations. There's a text file that you send to Predictor, and it has the instructions for the workflow you want to run. There's a new one for creating structure image files, one for images related to this structure sensitivity tool that we have in the GUI. That's shown here. That is for sort of flagging regions of a molecule that are important for a particular property, like hERG blocking in this case. Then one workflow to rank compounds using something called Multiple Criteria Decision Analysis, or MCDA. We've added some new endpoints to the REST API.
One to retrieve detailed pKa microstate information, find the dominant ionization state at a user-specified pH, to generate 3D conformers, perform arbitrary compound transformations like standardization, generate structure image files, and then some utility endpoints. One for retrieving current license usage and one for retrieving the version of the underlying ADMET Predictor engine, so 14.0 in this case. Just one quick slide on the pKa microstate endpoint. This is showing what the pKa microstate display looks like in the GUI. This is looking at one compound, the predicted pKa values. It enumerates the ionization microstates and shows their prevalence values and charge states, and so on. With this new endpoint, you get all the same information back. There's one block for each of the ionization microstates.
The endpoint also has the option to generate an image file that reproduces what you would see from the GUI. Okay, this is my last slide. Just a couple other enhancements in the API. The two endpoints related to HTPK, the basic workflow, and the dose optimization. They now let you create a new summary file that lists the sources of all the HTPK inputs per compound. That's just like this new logging feature that I described earlier. The same two endpoints now support specifying ACAT parameters. Things like liver-to-body ratio, the liver blood flow, GFR, and so on. In the past, you were limited to the compound-specific parameters, like solubility and permeability. So we've expanded that. The run command endpoint lets you emulate any ADMET Predictor command line run through the API. This has been extended to support the new script file workflows.
We've simplified the metabolite prediction endpoint, so now you get the metabolites directly. In the past, they were stored on the server and files, and you had to retrieve them with an additional call. Then finally, the ADMET Predictor installer has a folder with lots of request and response examples for all the REST API endpoints. We're getting lots of companies now trying to integrate the API into their existing cheminformatics infrastructure. This is just meant to help with that effort. Okay. I apologize for the breakneck speed there, but I'm excited to hand things back over to Gül and to Rafał to talk about what is coming in the next couple of months.
Okay. Thank you, David. I cannot take over the screen if you are sharing, I think.
Let me see. Stop share. Okay.
Okay. Thank you. Let me quickly share. Yep. I'm hoping you can see my screen and you can hear me well.
Yep.
Thank you. Good morning, everybody, or good afternoon, depending on where you are. What I want to show you today is something we've been building recently, which is S+ Composer, our agentic front-end and ADMET agent. You can think of it as an AI drug discovery assistant that sits right inside the tool you're already using, meaning ADMET Predictor. The idea is very simple. Instead of clicking through modules, exporting the files and re-importing them or somehow stitching the results together by hand, you describe what you want in plain natural language, and the agent orchestrates the whole pipeline for you. Over the next couple of minutes, I'm going to walk you through how it is put together and how this can be used for early drug discovery tasks. Yeah, stay with me till the end.
First, I like to show how it actually work under the hood. At the center of each agent, we have a large language model. It can be Claude, it can be something from OpenAI company or Gemini or whatever. But the point is that the model itself does not know anything about computing ADMET properties. What make this useful is everything which is around. The model is connected to a set of specialized tools. The ADMET Predictor does what it's always done well, which means it predicts ADMET properties. The S+ Composer, this is our agentic front-end. This is simply the editor, the graphical user interface that connects you to this ecosystem and lets you interact with it. It can reach into the vector store, where we have ADMET Predictor manual.
It can search the ChEMBL databases, and it can also operate in the file system. The key point is that the LLM is the reasoning layer that decides which tool to call and when. But every number it reports come from our validated scientific engines, I think this is important, not from the model guessing. That's a very important remark, I think. What is also important is that it is prone to the extension. It's very flexible. That's why on the left-hand side, we have this future MCP server. We can add here the additional capabilities, additional context, which is later on provided to the LLM. Today, I'd like to guide you through a quite simple but not trivial prompt. This is a prompt that a discovery scientist would, in principle, type.
You might notice that it reads like instructions you could give to your junior colleague. This is not a code. It simply says that the agent should take this more or less 400 compounds, calculate the PBPK properties, or ADMET properties, and give the histograms of solubility and permeability. Then to define the ADMET profile by combining solubility, permeability, and ADMET risk, and then rank the top 10 of them. Then for the top three of them, it is supposed to predict the key in vivo endpoints by using HTPK simulations and show the plasma concentration curve for human at the 10 mg dose. And finally, it's supposed to write everything up and create a nice markdown report. We have here four distinct tasks, several hundred compounds, multiple tools, and normally that would take some hours to go through it.
Let me quickly show you how this can be achieved by using Composer. Here we have the Composer. I'm just pasting these four tasks here, and I'm just asking the Composer to go through it. You can look what is happening here. The agent is reasoning about which step comes first. Here we have some skill activation. It is going to call the file system tools to read the SDF files. At some point, it is going to run, which is taking place right now, the ADMET Predictor endpoint to calculate the ADMET Predictor properties. It is actually operating here in this ecosystem where in the center of that we have the large language model, we have S+ Composer, which is this agentic front end, and all the other tools that can be used.
Not all of them have to be used, but if the agent decides that something can be important, for example, searching through the vector store to get some additional hints taken from the ADMET Predictor documentation, it will go through. It has gone through the ADMET prediction. You can see here that it is inspecting the files, looking what has been predicted. All the files you have here. All the 400 compounds, the predictions related to them are visible here. Everything is recorded. All the session details are stored on a dedicated file. You can always go back to it, look through, and check what was happening at that time. It is pretty autonomous. The only information that was given to the agent was these couple of sentences. It is autonomously going through all the predictions, all the reasoning behind.
Everything is happening within the so-called agentic loop. This is really the loop that is formed by various activities of the LLM. LLM is conscious about the presence of all the MCP servers, all the capabilities that are around, meaning this ecosystem. It is autonomously deciding what should be called when. If there are some problems on the way, then it is just redoing some sort of calculations or looking at the results from a different angle. You have quite a lot of autonomy here. You do not have to know anything about the graphical user interface. Under the hood, you still have reliable scientific engine ADMET Predictor. You can simply avoid all the clicking, all the details associated with the graphical user interface. Everything is done here autonomously by the tools that are delivered to a large language model.
I think now the LLM is at the level, or the agent is at the level of report creation. It has gone through all these steps, like ADMET predictions, HTPK predictions, MCDA ranking mentioned by David, this multi-criteria decision analysis ranking. Now we are just waiting to get the final report, which is supposed to be placed somewhere here in this reports directory. As you can see here, these plots, they have been created on the way. We have here one of the tools that we have here is the plot generator. You can really get publication-ready plots directly from this agentic approach. Here we have the summary. The workflow has been completed. All the tasks have been fulfilled. The task one, PCB properties calculation and histograms ranking, then in vivo pharmacokinetic predictions, and then report creation at the end.
Here we have the report. This is a markdown file. We can render that nicely to see a human-friendly document, which can be easily exported to PDF. At the beginning, we have the executive summary, and then all the tasks that have been accomplished by the agent. All the plots that were created on the way. All the results, the ranking of the compounds, the properties that are important from the perspective of this ranking and broadly speaking, ADMET profile. Then you have the task three, which is predicted in vivo pharmacokinetic endpoints. You have both the static endpoints as well as the full plasma concentration curves. Then at the end, you have the visualization of all the structures that have been ranked as the most promising drug candidates. I think this is it.
I'm hoping that I gave you some overview about what is going to happen in the near future. We are on the way to release our ADMET agent, and I'm pretty happy to share further details in the near future. Thank you very much for your attention.
Thanks.
Thank you.
Oh, go on.
Sorry, Gül. We just have one last poll to launch before we jump into the Q&A, and the last question is, if you're interested in having someone from Simulations Plus follow up with you about ADMET Predictor, and it's a simple yes or no here. We'll give everyone a couple seconds before jumping into the Q&A. I do see some answers coming in. All right, and I'll go ahead and end it there. Go ahead, Gül.
Yeah. Thank you. Thank you to all our speakers, Jeremy, David, Rafał. It is now time for questions from the audience. I will check. I think we have a couple of questions, and the first one I can see is how permeability can be better predicted by ADMET Predictor for such molecules. I believe it's, yeah, talking of during this question concerning Jeremy's talk, might be PROTACs, macrocycles, and cyclic peptides. If not, just correct me, please. David, what do you think?
Could you repeat that, Gül? I didn't quite get the question.
Yeah. It's how permeability can be better predicted by ADMET for such molecules.
You said how permeability can be better predicted?
Yeah.
Let's see. I think Jeremy did discuss the issue of the new descriptors that we've added in ADMET Predictor 14.
Yeah
That are meant to sort of capture the intermolecular hydrogen bonding and whatnot. I think he showed a couple of examples. Maybe the ones he showed were more related to solubility, but I think the same applies to permeability as well. We did find that overall on some data that was withheld from model building, that the performance did improve on beyond rule of five compounds.
Yeah. Okay. Thank you. If you haven't seen it, there is a Q&A box. Please add your questions while I'm asking other questions. Are there any publication that assess the accuracy of HTPK predictions?
Yes, there are. I'm not sure I have any on the tip of my tongue here, but I think on our website, we make an effort to collect all the publications that mention our software. I know with Predictor, there are at least a couple from Roche in particular, big HTPK users, and they've looked at the accuracy of predictions and discussed that in the context of using certain in vitro inputs to HTPK as opposed to model predictions. So those publications are certainly listed on our website. And then we have a publication of our own. I think it's at the preprint stage right now. We'll release more information about that publication in the near future.
Okay, great. Thank you. I think a couple of questions, and then I think we're done. Another one, is AP14 available for download?
My understanding is that it is now up on the S+ Cloud. I think we've changed the mechanism for distributing the software, and I think there's an automated way now where each organization that's a user of ADMET Predictor should have been notified now of how to go up to S+ Cloud and download the software. If not, please send an email to info@simulationsplus.com. My understanding is right now, both the Windows and the Linux versions are available for download.
Yeah. I'm also thinking about the time. The last question. The other question, if we couldn't reply, definitely we will reply after the webinar. One of them is like, does HTPK alone can help researcher to publish the data in a good journals?
You're going to have to ask that one one more time again, Gül. I'm sorry.
Yeah. Does HTPK alone can help researchers to publish the data in a good journals? I mean, it's quite depends, isn't it? Like depends like what you are getting, what you are researching, in my opinion, but I don't know.
I'm sorry. I should have had the panel up on my screen so I could read the question.
Yeah. We can reply later if-
Yeah. Well, why don't you send that question? I'm just not sure I quite got the gist of it.
Yeah. No problem. Okay, I think, yeah, we are coming to end. Thanks to speakers again, and thank you everyone to join and post a question. I will hand it to Jasmin to complete. Thank you.
Yeah. Just want to say thank you to everyone. For any questions we were unable to answer live today, someone from our BD team will reach out to you with answers following the webinar. We invite you to visit our website to learn more about how you can leverage our early-stage discovery program. I put a couple links in the chat to our website. You can also follow us on social media. Lastly, this webinar has been recorded for playback and will be available on our website and YouTube channel. This concludes our webinar for today. Thank you, everyone.
Thank you