Hello, welcome everybody out there to a next edition of our IR virtual tutorial from Deutsche Post DHL, in our effort to make an experience for you what we mean by "Excellence. Simply delivered.," our claim. For today's round, as you've seen from the invite, I'm glad to have with me today Katrin König. Thanks very much for taking the time and making yourself available. I'm pretty sure it's going to be an interesting, I don't know, 60 minutes, 70 minutes, something like that. We have introduced this format not too long ago to have a chance to still educate you on specific topics in the group which we think are of interest for you. I'm pretty sure you're going to agree with that also after today's session. Invites are already out for the next one with John Pearson and team on why e-commerce works for DHL Express.
Looking forward to that one, and there's more to come. Let's now focus on today's topic. You see here our strategy, pictorial group strategy as announced last fall. Needless to say that over the past couple of months, we had plenty of opportunity to see that the strategic approach of our group is actually working out. One element we want to focus on today is, of course, digitalization as one of the key enablers of the group's future success. What do we mean by digitalization? Obviously, in line with our three bottom lines, we want to come to a better customer experience with the help of digital tools. At the same time, there are plenty of opportunities to make work easier and more efficient for our employees and overall increase operational efficiency.
That's where you guys come into play because that's going to typically spell some sort of financial benefit. We will take a look today at how we are dealing with the theme of digitalization. We have been identifying five areas of digital improvements and initiatives where we, on a group level, run so-called centers of excellence, which are then sort of developing the right know-how and the tools and everything that's necessary and then make sure that this is scaled up through all the divisions, which is clearly a big advantage over smaller players who only have one division and have to take the investment, and we can much better scale that. You see the areas that we are covering with this approach.
Today, quite obviously, we want to talk about data analytics, and I think what Katrin is going to show you is that indeed this approach is one of the key enablers really to digitalize our operations and to bring it into life and fruition throughout the group. There are, as we will see, a multitude of solutions, whether that's benefiting the operational experience from an employee point of view or leading to a better customer experience or simply saving money. We will share with you where we stand in rolling this out throughout the group. One more technical remark. For those of you who followed earlier tutorials, you know the procedure. You see on your screen a Q&A field. Whenever you have questions, just punch it in. It will be ending up here on my screen, and we're going to deal with it in the Q&A session.
With that lengthy introduction, finally, over to you, Katrin, please.
Yeah. Thanks, Martin. Welcome also from my side. I'm really glad I have the chance to talk about that exciting topic today. Let me start the whole topic with a short teaser. This teaser is an example for one of the data science solutions that we developed, and I think it nicely shows what data science can do for us as a company and what benefit it brings. What is it about? Well, it's about a problem that you probably all know from your personal lives as well. Sometimes you need to go to a place you have never been before, it's an address, and you need to find this address on the map. Could be the hotel of your next vacation destination or something like that.
I think all of us to do this use external map services like Google, Bing, or HERE to solve that. When we look at our company, we actually face this problem millions of times every day because we get the shipments from our customers with some address information which can be more or less clean or structured, and we need to locate that as well, somewhere on the map on the earth. In order to find those addresses, we actually came up with our own internal approach because, as you will see in a minute, this works for us much better than the existing services. I will show you now how we leverage our unique internal data assets together with machine learning algorithms to solve this problem of geocoding.
Let me quickly just share our geocoding engine with you. This is our geocoding engine and the front end that we developed, and I think that shows nicely what we can do here. What you see on the left-hand side here is a delivery address from the Netherlands, as we typically get them from our customers. Now, and this looks similar to the typical search engines you know, we want to know where this address is. Let's search for it, and we see it's somewhere in the area of Tilburg. If we zoom in further, we see that in this case it's a hospital, and we get the geocode on a very specific point in that hospital area.
I want to explain you quickly first how we do this, and then secondly, why this works for us much better than the existing services that are out there. I mentioned already we have millions of shipments that we deliver every day, and whenever we do that, we actually capture the geocode of those locations. This is what you see here now on the map. All those blue pins are historical deliveries in that area. This is the base for our solution. We use that data and we apply machine learning algorithms to it to structure the data, to match it, and then to determine the final geocode. Why is this internal approach with our own data so much stronger? The advantage is that the information that we use is much more relevant and specific for our purposes.
This is a hospital, as I just said. Very often, we need to deliver to large buildings like hospitals or malls or so. You can see here that our data knows exactly the location where the courier needs to go to. It wouldn't point us to the main entrance or the center of the hospital, but really to the courier entrance, so to say. An external service, as you see it here, with a red and yellow flag, just cannot know this information. This is just much more powerful, and you can imagine in countries where we don't have a structured address system in place, this becomes even more beneficial for us. This was a quick example and a teaser, so let's get back and start with the actual content and look a bit more into the topic.
We want to do that based on deep dives in four areas today. We want to start to look into the big opportunities that we see and that we have in the area of data analytics. Second, we then want to share with you our understanding and our approach to analytics with a typical project cycle. We want to share further use cases and especially talk about the big initiatives that we are running currently in the area of analytics and the benefits that we get from those. Finally, we will share a bit more about our approach on how to scale a topic like analytics into a large organization like ours. Starting with the big opportunities. What you see on the next page, I think is not new to you.
Those are our three bottom lines, and I think the geocoding has already shown you that analytics can bring benefits in many different ways. When we look at those bottom lines, analytics can clearly support all of them. Let me just go through them based on the example of the geocoding that we just saw. Clearly, if we have more precise geocodes, we can do a better route planning, and we end up with shorter routes. This leads to clear cost savings, but of course, also reduces our CO2 emissions that we produce. Also when we look at our couriers, they have a better tool support. They get more precise routes. They have to do less detours, and they can deliver better service.
This finally, when we look at our customer, pays off for the customers because they experience this better service quality that we can deliver to them. We will go through further examples along the presentation, and you will see that all of them will address at least one of our bottom lines. Before we do that, I want to have a bit of a closer look into the financial aspect and the investment of choice bottom line. What you see on that page are our main two cost blocks that we have. The staff cost and the transport costs. When we had a look at all the analytics initiatives and projects that we ran, we saw that almost all of them address at least one of those buckets. I can maybe quickly highlight that on one very concrete example.
This example is about volume prediction. If we have a data-driven forecast of the volumes that we expect to handle in our sites and the volumes that we want to transport on our routes, we can optimize the resource planning and also the transport routes based on that. We've seen in the projects that we ran that that can lead to significant reductions in both areas. This is also why, and I think Martin said it in the beginning, analytics is one of the substantial pillars in our digitalization strategy and will have a significant contribution to it.
You do see big numbers on that chart, but today it will not be about sharing with you concrete numbers on savings use case by use case, but it's much more about understanding how we really enable an analytics culture, how we create a mindset shift. This will really enable our organization to manage the overall company better in many different ways and aspects. All right. To illustrate that to you, I think I will now talk a bit about our view on analytics and really on our approach to the topic. To do that, let me maybe start by sharing with you our understanding of analytics, because I think that's quite important.
Because what I often see is that there is an impression that analytics is purely about taking huge amounts of data, throwing that into some magic tool, running some fancy algorithms, and then automatically you get exciting insights that are, of course, business relevant and value-adding. Well, our understanding and our view of the reality is a bit different. Actually, in our understanding, we start the other way around, and you see that illustrated on the left-hand side. For us, analytics always starts with a very concrete business problem and a question you want to solve. Taking the volume prediction, I want to know how much volume I need to sort in my sorting center tomorrow. Only then you start collecting data, data which can help you address this question, and that can come from different sources, can be in different formats.
Then you start the whole process of cleaning, processing the data, and you run your mathematical models. You always have the business question in mind, and you always combine it with the business knowledge and the expertise that you have, because only then you can really come up with valuable insights in the end. We will go through an example a bit more in depth in a minute, but I would also quickly like to show you a bit what is inside this machine that you see on the left here. On the right-hand side of the screen, you see one way of classifying data analytics. This is a quite common framework, so you might have seen that before. Let me quickly go through it. It starts at the bottom left with descriptive analytics.
This is really about understanding the past and understanding what has happened. The typical business intelligence or dashboard type of activities. When we then start looking into the future, trying to make guesses about what will happen, we are then in the predictive area. You can make all sorts of predictions. You can predict volumes, you can predict customer behavior, demand, or you can even predict the geocodes as we've seen it in the beginning. The next step after predictive is then the prescriptive analytics. This is about optimizing your processes based on those predictions. This might sound complex, but there's a simple example. We have a prediction of our volumes in our transport network. Well, based on that, we optimize the transport routes for the next day. The descriptive area is an equally important part for us.
In the following, when I will talk about data analytics and data science, we will refer to the predictive and really the prescriptive analytics applications. All right. Now, we have a common understanding of analytics. What I would like to share with you now a bit more in detail is our approach to a typical analytics project, because I think it has some quite interesting insights that we see along that process. Those general six steps that you see here, they are applicable to more or less any analytics project that we run. I want to make this again based on a concrete example, which is the operational volume prediction at Express.
Which is actually referring to the daily challenge that our guys in the Global Express network have to deploy the right capacity on a given flight from a hub to somewhere in the regional network to have an idea what's about the volume we're expecting on the next day or on a given day, and what type of capacity planning do we need to come to, right?
Exactly. It's for the next couple of days, and it helps us to steer this capacity in two ways. On the one hand side, we can make sure we have enough capacity in the network. On the other hand, we can also sell excess capacity. It's really an important input for all planning that follows then.
Selling excess capacity, not the usual problem these days, but in normal times.
Yeah.
Perfect. Okay, let's get through it.
Good. We start with the first step, the problem definition. I mentioned already earlier, that we always start with the business problem, and this is what this first step is all about, really getting a very good understanding. Martin already shared a bit what it is about. It's really about predicting the volume on each flight in our flight network for the near future with the two clear objectives that we have, steering our capacity for quality and for selling excess capacity. This is so really the very first step, having an understanding of what it is about. Second step is the process understanding.
Here we really want to understand how the model or the algorithm that we build will fit into the processes, because only then, if it's fully integrated in the end, we will see benefits from the work that we are doing. Here, the data scientists, typically, they join the operations, they talk to users of the tool and so on. It's a quite important step. Here, for example, we learned that there was an existing tool and a process already in place, so a tool with a rather simple forecast. That was really good news because we knew that there was not a big change in process or also change management with the people that we had to do. That was quite helpful in that case. Then we can get to the third step, which is data collection.
Here, you can really get creative and think about all potential data sources that could help you answer this question. To clarify that already up front, there's always a lot of ideas about fancy data sources, and it's good to think about them. What we've seen in this case, but also in many other projects, the most important data is really your own historical data. You want to predict your volumes? Well, you need to have your historical volumes as granular and as far back as you can have. Of course, there are other important data sources, like information on holidays or the flight plan that you collect in that phase. That brings us to the fourth step, the data cleansing and preprocessing. This is quite an important one because here the data scientist really sits very closely with the business experts.
Here, it's about trying to understand patterns in the data. For example, you see an outlier in your data series. You want to understand the business reason behind so that you know whether you take out this outlier, whether you average it out, whether you leave it in. This is really very important and very interactive between, well, data science and business. Just to give you one concrete example, again for the Express case, heavy weights were a big topic in that part, right? Understanding whether they are part of the historical data, whether we want to take them out, whether we want to forecast them maybe separately to avoid distortions. This is what this is all about in that phase.
All right, the modeling step comes, and I think this is quite interesting because when you think of data science, you think it's all about modeling and all about what this guy here on the picture is doing. Actually it's only step number five. There are a lot of important steps, yeah, before that. Why? Well, in this step then, the data scientist really combines his methodological knowledge with all the business information and all the business knowledge that he or she gained so far. A quite creative process starts because data science, well, it's called science, sometimes it's more an art because there's not a textbook solution to it. There are a lot of algorithms that you can try out, play around with, it's all about integrating this business knowledge that you found out in a smart way into your model.
Again, I can make a quick example here. I talked about holidays, and I think it's quite obvious that holidays do have quite an impact on our volumes. What we found out, for example, is that holidays in Australia do not only impact the volumes in and out of Sydney, but they also have an impact on our flights from Hong Kong to Leipzig, for example. Those complex interdependencies need to be somehow taken into account in the model. This is a little bit of a glimpse into what happens there. Once we are done, well, we are never done with modeling, but once we found a model we are happy with that has promising results, we then get into the last step. That's the implementation and really the business process integration that takes place here.
I already mentioned that we were quite lucky in this case because we only had to integrate our algorithm into the existing tool. What we saw then was, well, on the positive side, that the improvements that we saw on the theoretical side when we looked into the modeling also turned out to realize in the practical or in the live environment. We saw improvements of 3 percentage points in the forecast accuracy, which is huge in that area. Of course there we also found in the initial phase some smaller pain points, which we then of course took back to the model, fine-tuned the model to get rid of those. Yeah, now model is up and running and integrated into the business.
I think this hopefully gave you a bit of an understanding how we see and how we approach data analytics, which I think is quite important for the following part. I already shared some use case examples along the way, but I want to spend a bit more time on further use cases and especially on really the big initiatives that we're running in that area. Yes, what you see on the next page, those are really the six key analytic solutions that we are developing and working on at the moment. Of course, there are a lot of other very specific projects that are driven from different teams. Those topics that you see here are really the ones that have the highest relevance and the highest potential across our divisions.
This is why we develop those really into group-wide solutions with a very high degree of scalability. Given they are really important for us, I would quickly like to go through them. You can split them into two groups. The first three topics that you see here, they are really addressing our core processes and optimizing them. The first one, the operational volume prediction. We've just seen the example of the Express sector forecast, which is exactly this. This is just key for our business to plan our resources in our sites, to plan our transport assets. Every percentage point of improvement that we see here just has a huge benefit for all the subsequent planning processes.
The vision that we have here is really that each planner in all sites and all facilities has such a data-driven forecast at hand that supports them in their planning steps. The second topic, the staff scheduling. This is the natural next step, I would say, because we do know our expected volumes from the prediction, and we also have a rough shift plan in place. Now it's about the short-term planning, the exact shift scheduling, determining who works where, when, considering sicknesses and breaks and all those things. This is a super complex optimization problem. Our staff dispatchers in our sites, they spend hours with that task every day. What we did here is we developed an interactive optimization tool that supports the dispatchers in that process.
That gives them more optimal solutions, but also helps them to run this process much faster than they could before. Again, we want to have all our dispatchers using such a tool support. That brings me to the third topic, which is routing optimization. Routing is, of course, not new. That's the core of our business, and we have been doing that for years and decades. What we see here is that there are really new trends coming up because somehow, our world is changing faster, and we see that the requirements of our customers, but also from our operations, are changing faster, and we want to change our routes more dynamically. What we've seen here is that existing software that we used and tools that were in place were often not able to cater for those requirements and flexibility needs.
On the other hand, we saw that our internal solutions, they have that flexibility and that they can be really tailored to those needs. Let me go a bit more into the details of routing. The page that you see here, we could actually draw for any of our solutions. Typically, there are a lot of application areas and a lot of projects below each topic. It's not one project, but a multitude of application areas. The picture for routing, you see here. We have application areas and many we use. A lot of them really focusing on last mile, of course, because this has a huge cost lever, it's huge cost driver for us. Our vision for routing is really that we leverage our business expertise that we have and our method knowledge to develop truly customized tools for our specific problems.
Through that, we are able to really address this flexibility needs that we see. Now, this might now still sound a bit abstract, which is why I brought a very complete project from the space of routing, and this is the RaptOR tool that we developed for freight. The RaptOR tool supports our dispatchers in our freight terminals in their daily work. Those dispatchers, every day get a hundred or hundreds of orders, which they need to assign to tours and delivery vehicles. When they do this, and you see that in some of those yellow boxes, there are a lot of constraints and a lot of things they need to take into account. Vehicle sizes, capacities, stackability rules, driving regulations. The option space for them simply explodes, and it's impossible to solve that for a human being in an optimal way.
This is what RaptOR now does for them. The algorithm takes all those constraints into account and comes up with an optimal delivery schedule that the dispatchers can then use. RaptOR does that in seconds. That brings me back to the bottom lines we talked about earlier. You see, first of all, that this is a huge support for our dispatchers. They save a lot of time that they can spend on other tasks, like finding even better deals for the routes that we came up with.
This solution is being derived in seconds. However, using the process that you described earlier, all the input that you had from previous observations, what a role does a local holiday play? How are Wednesdays compared to Fridays? What's the current traffic situation? All these sort of things, right?
All these sort of things are taken into account in the algorithms, and of course, the development was not done within seconds, but now the algorithm really runs in seconds. Whenever you give a new input, you get a new solution out for that.
Okay. There's probably also something like a learning curve and further fine-tuning of the algorithm in the course.
Yes, definitely. I think one of the strengths of the algorithm is that we can really integrate or include new constraints rather quickly. I can give you one example for that. For one of our customers, we typically roll it out customer by customer. We had one customer with a transport network that had a lot of shipments from Eastern Europe into Germany. What would RaptOR do? Calculate the shortest routes with all the constraints, and it would travel, of course, also through non-EU countries. We learned from the customer that they absolutely want to avoid that because if customs processes, that's effort, that's time. This for us was a new constraint. The team took that back and could include that into the algorithm within a few days.
This is what I meant when I talked about this flexibility that we need because you won't find a software that has all and everything included. With the internal capabilities, we can just quickly adapt it, and that's then truly customized for our customers.
Okay. Your team is then also in constant exchange with the actual users of the tools, yeah?
Yeah. The team spent a lot of time on-site, really with the dispatchers. In the initial development phase, it was more or less like a constant line that they had. Dispatchers run something, they said, "This doesn't make sense." We looked into it and iterated through it. I think this also helped the dispatchers to really gain trust into the tool and into the work that we are doing.
That's important, yeah.
Yeah. All right. Now a quick look at the other three solutions, which are our so-called advanced data services. Here we really leverage our internal data to generate new benefits or new insights. The first topic, we've seen in the very beginning, the geocoding. Here we leverage our internal delivery data, to come up with geocodes for new addresses. The second topic, the product classification for customs, that's based on a pretty similar idea. Yeah, what is it about? Whenever we send a shipment across border, and you can imagine we have a lot of those shipments, we need to determine the customs code for this shipment. This is a code that specifies exactly what's inside the shipment. That could be a cotton t-shirt for men in red or something like that.
This process today literally is still done either manually by the customs agents or externally via brokers. As you can imagine, both is quite some effort in terms of time or money. Again, here we do have data assets that we can leverage because we have a huge database of historical shipments. Those shipments, they all have information like a description and value and other information. Of course, for the historical shipments, we also do have the customs code attached. We can learn from this data, again, apply our machine learning algorithms, and then predict for each new shipment what is the most likely customs code. If you remember the process as it has been in the past, this has a huge potential for automation, of course. Now, the last topic in that area is the invoice overdue risk prediction.
That's a topic from the finance area, and again, I want to go a bit more into detail here. Again, we do leverage here data that we have. In this case, it's our internal accounts receivable data, and we use that really on an invoice level. We come up with algorithms to predict then for any new invoice, the risk or the likelihood that it will be paid late, on time, or even early. Now, this prioritization in the past has been done either manually by a collections expert or based on a set of fixed rules. What we now have is a machine learning algorithm that takes into account hundreds of statistical features, and it's constantly learning and adapting. We've seen that this is a really valuable input for the collectors and the collections processes.
Since the IOR score, as we call it, is in place, we've also seen huge benefits for our cash inflow processes and steering, and of course, also during the corona situation, this has been a very valuable input.
You're filtering out those invoices that we have sent out, which are running a relatively higher risk for not being paid on time, but not simply by only taking past payment patterns, but also taking into account state of the customer industry and other factors. Yeah? It's not only about how well this specific account is doing, right?
Yes, exactly. As it says here, it's hundreds of really variables and characteristics about the invoice, the customer, the industry, past behavior, not only looking back a few weeks, but really looking at trends and doing extrapolations. This is all what happens in this quite sophisticated algorithm. What it spits out is really for an invoice risk flag. Is it high? Is it medium? Or is it low?
That means all the sales guys out there who are taking care of their customers have a clear list, focus list, which customers to contact and call and make sure that they are aware of that and don't waste their time on more secure cases.
Exactly. It's a very crucial input now for the collectors and for prioritizing their day. Maybe that's a good moment to also show you how that looks in practice, right? Because I talk about machine learning models, statistical features. That sounds abstract, but also fancy. I can quickly show you how that really looks like in our daily work. To do that, let me quickly share my screen again. What you see here is a dashboard that a collector at DGF is using to plan and prioritize their day. As you can imagine, the data that we show here is no true and real life data, but it's test data that we set up for training purposes of our collectors. What you see here is a list of customers for one of those collectors with all sorts of information.
This is what they had already in the past. What is now new is this colored column here, the IOR column, that assigns this risk flag, so medium, high, or low risk of paying late. You see it's really straightforward. It has the color coding. It has three categories. At the first glance, it already helps you to prioritize where to look at first. If we look at one of those customers a bit closer, let's take the second one here, the customer with a high-risk flag. You get even more details. You see here all the invoices that are attached to their customers. Again, all with a score here. This is really a high-risk customer because each single invoice gets a high-risk flag. You can also see here then the actions that are triggered based on that.
Some of them automatically, some of them, of course, are derived by the collectors. You might now be a bit, I don't know, disappointed, but you might have expected to see something more fancy because the algorithms behind are quite sophisticated. What you see here is really doing exactly what it is supposed to do. Because we have those algorithms, and the results of those algorithms are really fully integrated into our business processes.
Well, I think what our specific audience does find fancy is, what is the actual real-life impact? Have we seen any improvement or any positive change to the collection success?
Yes. We have. For DGF, in the cash inflow, we've seen a positive impact of a double-digit million euro figure. On the EBIT side, this led so far to a single-digit million improvement that we've seen. That on the financial side, but of course, this comes with a lot more of automation potential. We, for example, or DGF, they now installed bots that are the so-called digital twins that are treating some of those customers with a very clear pattern automatically. This is a huge efficiency gain on top. We could reduce our bad debts because we really identify problematic customers earlier. There are a lot of benefits to that.
Right. Very tangible and relevant.
Yes.
Very good.
Yes, maybe one comment also to make on that. Of course, it had to come with quite some change management and training we had to do. Of course, we are now, well, providing a, at first glance, black box to our collectors, and we had to build up trust with them because they should rely on something they didn't know what it was about. We did a lot of measures in that area. One nice one was that we took a set of new invoices. We asked the collections experts to classify them into high, medium, low, and we let the algorithm do the same. In the end, the algorithm was better, and this was one way that really helped to overcome the skepticism. Now the collectors really see that as a support, which it should be, right? It helps them really focus on the key things.
Often you have to overcome this first wall of resistance because these guys, these collectors, they've been doing that for years, and obviously, their starting point is, "Well, I know my customers best," right?
Exactly. This is often, and I think that's a valid skepticism that you can have in the beginning. If the machine cannot beat them, well, then it also shouldn't. In this case, it was then clear, but very often it's also not purely about replacing what they have been doing in the past, but providing input to them so that they can really focus on the things that the machine doesn't know, on the local expertise that they have. Very often it's more valuable input that they can use to then do the last 10%, 15% even better.
All right. Good.
Okay. Those were, I think, quite some examples that showed you how we leverage analytics in our daily work and how we integrate it into the business. Let me now finish by talking a bit about our approach to really scale this topic of analytics into our organization, which is not as trivial as it might seem. Starting with the organizational setup. What we have established here is a hub-and-spoke approach. Because for us, it really combines the advantages of a central and a decentral approach. We have our hub, which is our central data analytics center of excellence. In this team, one of the main things we do is we are developing those group-wide solutions. The six topics I just talked about that are applied in every division. Besides that, we also act as a talent pool for the organization.
We really bring in the latest methods, best practices. Of course, we try to connect the data scientist community across the company. The central team is then complemented by divisional or functional spokes. Those are smaller data scientist teams that sit within the business. That could be an ops analytics team within P&P, or it could be a dedicated finance team. Those teams, of course, they are much closer to the business, and they bring in the subject matter knowledge. Just looking at the invoice overdue risk topic, that was so valuable to have a dedicated finance data scientists there who knew the finance processes, the finance terms, rather than just a data scientist who has never dealt with finance topics before.
This is really the benefit that we see, and very often we combine also projects with data scientists from the hub and the spoke to really get the best of the two worlds out of that.
That's also addressing one of the questions we got in already from Alex Irving from AB. This center of excellence, and you're part of more of the corporate center.
Yes.
Having counterparts in the divisions. It's basically an offer to the divisions. The tiny but still cost of this central group is not borne by the divisions, but they don't have to have any fear of, "If I try this, it's only gonna cost me." Yeah.
Yeah.
How do we decide on where to deploy now these tool and techniques? Who is making that selection? Is that coming from the divisional side?
You mean more or less the prioritization of topics we're dealing with?
Yeah.
Well, this is. Maybe we can jump to the next page because this explains very nicely how this is done. It comes mainly from the divisions, because we have a very clear setup now in terms of prioritizing that. You were right, we are offering our service. Of course, we are also making an estimate of whether this is valuable or impactful, beneficial topic, and that has a lot of aspects and criteria to it. Of course, also the divisions are prioritizing the topics they want to work on and their key focus areas, and then asking hence for our support on those. I think you can see that in the governance that we set up for steering the whole scaling of analytics, because we believe that analytics has to be fully embedded into the business. You cannot purely drive that from a central perspective.
You have to have it embedded in the core functions. You have to have people there who know the topic and who embrace the topic. What have we set up on that? It's three main pillars. The first one, that's the analytics executive sponsor within each division. This is a designated top manager, typically a divisional board member, who is driving the topic into their organization. They are on the one hand, an ambassador for the topic, but they are also really responsible for defining an analytics roadmap for executing and implementing that. Roadmap means very clear use cases. What are your top five cases you want to run this year? Which ones can you do with your own teams? Which ones do you want our central team to support with?
This is more or less the process for prioritizing that within each division.
The second element of the steering that we have is the analytics coalition that we establish in each division. Analytics is a cross-functional topic, it has to be driven from three parties, more or less. Clearly from business, because if business doesn't want it, doesn't commit to it, you will never achieve anything. It's business, it's IT, and it's the analytics function. In this coalition, the senior leaders from those three functions get together to really lift this shared accountability. They pragmatically decide on prioritization, roadblocks, how to handle them. To really accelerate the whole execution for that. Of course, we have the third element, which is the group-wide element, our analytics steering board. In that board, which is chaired by Frank Appel, the divisional sponsors get together.
They track and share their roadmaps, so the divisional roadmaps that everyone defined, but they are also quite open in sharing learnings and failures. We really trying to help each other progress on that journey. What we also do there, So the second part of the answer to your question, there we then prioritize those cross use cases. The solutions, so which should be solution number seven, this is a topic that we then discuss in the board with all the divisional sponsors, of course. Lastly, what we do in the board is, we talk about topics that we should drive from a group-wide perspective. There's one quite nice example for that, and also very important one, which is about education.
We discussed in the last board the challenge of capability building and upskilling, which is quite still a challenge in our huge organization. We decided on a set of trainings that we want to develop. Our COE is now developing trainings for different target audiences and our first e-learning. That's a basic awareness training, more or less for every employee that will go live in the next months. Those are the topic that we could drive there.
Okay, good. What would you think if you were to think of the whole digitization process that we want to go through now until 2025 with the EUR 2 billion spend and the benefit? I think I have to think of that as a sort of an S- curve.
Are we basically done by now with all the preparatory work and sorting out the details and now we're really aiming to increase the actual number of use cases out there?
Yes, I think that's a fair assessment. I think we are now really at the point where we can do the scaling and really transfer it largely into the organization. We have been dealing with the topic for now five to six years, so the foundations are in place. We have a significant amount of people there. Of course, we are not done yet, but this is really, I think, the point to make it large, and this is also why we have established this governance now because it helps us accelerate and really translate it into the organization. Yeah.
All right. Good.
Okay. I think with that, we are at the last page and a short summary that shows our holistic approach to the topic. Some of the elements you see here, we've talked about a lot, or I've talked about a lot. The use cases, the data scientists, the people, our approach, and the scaling. There are also other very important elements which I haven't talked about yet. Of course, to really scale it holistically, they are very important enablers we need to have, and we need to get in place. It's the data infrastructure and data governance, which we are also giving a lot of importance to, but also the education part, which I just briefly touched upon. Really teaching everyone what data analytics is and what it can do.
Together with all these elements, I think we have shown you that we have a setup in place that can really help us to generate significant benefits on the financial side, but really also for our employees and customers.
Great. Thanks, Katrin, at this stage for taking us through the slides. I think a very thorough introduction of how we're approaching the whole theme of data analytics and make it deployable and usable within this group.
A couple of questions we have come in the meantime. One again from Alex, but also from Muneeba, from BofA. Is this something that we are developing internally with your help, that you have to do internally, or is there any sort of off-the-shelf third-party software or solution available that you can buy in the market? If so, the question, how do you think the two approaches compare?
Yeah, I think there's no one-fits-all answer to that. I think there are certain processes that are just so much core to our business and key to us, like the routing I talked about, like the geocoding, where we definitely get better results. The better approach here clearly is to develop that internally. It's our core competencies. We can constantly adapt it. We can bring in our logistics model. We have seen that in a couple of examples for routing, that whenever we try to go externally with a vendor or with an existing tool, at some point we fail because there was this one requirement that the tool couldn't cater for. The things that are really close to our core, we develop internally. For some of the cases, we also went out to benchmark ourselves.
We ran pitches against experts providing those solutions.
In all the cases, we were better or at par with those. I think it shows that it's really also not just the belief that it should be close to us, but we are also really better there. There are other topics, where there are good tools out there, and it's maybe a topic that is not so close to our core. If we think about topics like natural language processing, so it's converting voice or text into a structured format and then doing algorithms with it. On the tools that translate the voice or the text into a, well, digestible format and all of that, there are really good providers out there. There, I think we don't have to reinvent the wheel, but rather pick on what is out there and then maybe use this as an input for then the next step or the next level.
Okay. Wherever one of the core input elements is our own collected data, probably the more unique and useful it gets.
Our own data and also our own knowledge. This holds true also for this volume prediction. There are a lot of companies who offer like i t's time series forecasting, more or less. We understand our processes, we understand the interdependencies of the logistics flows and all of that, and this allows us to tweak the models in a way that they, in the end, result in that output.
Which brings me to one question, which was also raised by Muneeba, and probably comes to many's mind. You talk about predicting and taking past data. We've just been through a period of massive disruption where on all levels, a lot of former rules somehow no longer apply. How much of a shock has the whole corona situation sent to the workability of your models?
Mm-hmm. I think it's a bit different depending on the time horizon you look at. In this invoice open risk case, for example, we saw, because the model typically looks back only, as a very short term, and those algorithms, it's not the machine that is learning, but they are constantly well adapted. The parameters are recalculated every day based on the recent history. Here, we've seen that the models could really quickly pick up those changes and the trends. We haven't seen problems, but we've really seen that it worked pretty well. When we think of long-term predictions, right? We want to predict annual volumes. Of course, here we have to see how we handle this year, right? This is exactly then this outlier question, right?
Right.
What do we now do? Do we average out 2020 by the last three years' average volume? This is then also something that you have to discuss also closely with the industry experts and see how you handle that. This is really more for the long-term prediction with the problem.
Okay. Yes, we're the largest player in the world, but there are other large players as well. Are you aware to what extent our peers are following a similar approach? Is there any exchange even?
I wouldn't call it exchange. All of our competitors are looking into those topics. Taking the topic of routing, for example, this is an optimization topic, and operations research has one large conference every year. Of course, all our competitors are on that conference, and we are fighting for the best talent there. It's definitely everyone looks into that. What is special about our approach without knowing how they do it, but for us it's really this close interlinkage between, well, the data science part, but not from an academic perspective, but really closely linking it to the business from the very beginning and combining with our logistics expertise, which is our asset, right, with those new technologies.
It's probably the advantage is in the size and the multitude of the divisions. We can use data input not only from forwarding but also from Express supply chain, what have you. I would think also the other way around. You said you're running 100+ data analysts. I don't think that there's a big number of forwarders out there having access to that number of resource, right?
Yeah, exactly. Not only that we can take the data of our divisions, we can also, once we've found a good approach, we can easily roll it out. Not only we develop something for Express, we can transfer the algorithm and the intelligence behind to the other divisions.
All right. Good. Just a quick reminder before we come to the last question that I have received so far. If you want to place any of your questions, simply punch it in there. I think one of the questions is the EUR 2 billion digitization spend. I think this is something that I'm happy to clarify. This is obviously something that is baked into our guidance short and midterm and part of the overall performance. It's not something that comes on top of or as a one-off expense. That's all baked in there, but I think it's a very important signal to the organization that we're serious about it, to spend a total amount of EUR 2 billion only on digitization.
From listening what you said, I think there shouldn't be much doubt that we are able, with a multitude of projects, that we're going to harvest that and get to a very decent run rate of financial benefits over time. Great. Well, that's fantastic. It's bringing us to the scheduled full hour.
Perfect.
I think we covered all the ground. I hope we covered all the questions, so I definitely took care of everything that you sent in. Thank you very much. Thanks, Katrin. Hugely interesting and gives you an idea that one could talk about this probably for many days if you really dig into.
I definitely could, yes.
Thanks for sharing that valuable hour with us. I hope it was useful also for you guys and to give you a better understanding on how we're steering the group through the next number of years. Okay. Very good. Thank you very much. Looking forward to seeing you on October 5th when we're running our next tutorial. Until then, bye-bye. Have a good rest of the day. Thanks, Katrin.
Thank you. Bye.