00:00:00 Introduction
00:02:49 Can Claude really forecast?
00:07:49 Do LLMs actually process numbers?
00:09:12 How will AI affect forecasting vendors?
00:14:38 What is the difference between forecasting and planning?
00:18:03 Can LLMs support supply chain decisions?
00:22:31 From revenue forecasts to daily operational decisions
00:27:11 Why are naked forecasts unusable?
00:30:12 Forecasting granular supply chain decisions
00:34:57 Can an AI independently generate purchase orders?
00:42:24 Which industries will adopt AI fastest?
00:44:25 What becomes more or less valuable?
00:52:48 How should companies evaluate AI vendors?
00:55:51 Using ChatGPT for vendor research
00:56:19 Why adversarial prompting matters
01:00:36 Can Lokad pass the same stress test?
Summary
Conor Doherty interviews Joannes Vermorel (CEO and Founder of Lokad) about Claude Fable 5.1 and its advertised forecasting capabilities. Joannes argues that the real breakthrough lies in agents’ ability to automate analytical and coding tasks. They examine why time-series forecasts are not supply chain plans, how LLMs can support decision automation, and which industries may adopt them fastest. The discussion concludes with a practical framework for challenging AI vendors, assessing their technology, demanding measurable productivity gains, and using adversarial prompts to verify marketing claims.
Full Transcript
Conor Doherty: So Joannes and I just wrapped up our discussion on Anthropic’s release of Claude Fable 5.1, specifically its at least publicized forecasting capabilities. Now what produced the conversation was a video that we received from a friend of the channel. They wanted our opinion, our response to the implications of this potential forecasting breakthrough, which is of course an LLM model. So Joannes and I discussed that and it sort of led to a broader discussion on how you really should evaluate vendors in the space including Lokad and by the end Joannes issued a I think a very practical guide or diagnostic test if you happen to be confronted or presented with a proposal from a vendor who says “Yeah, man, my AI can completely autonomously forecast, plan, and execute your supply chain end to end.” As I said Joannes had some very concrete and useful insights on how you should stress test those claims. So, if you take nothing else from this episode, stick around to the end or even just skip to the end, I guess, if you if you prefer, and take away that practical information. And with that, I give you today’s conversation. So, Joannes, I brought you here to discuss a video that was sent to us by a friend of the channel. I sent it to you. We discussed it privately, and now we will sit here and record a podcast on it. And that is the video that was released by Anthropic that described or let’s say advertised the forecasting capabilities of Claude Fable 5.1. Now, for the rest of the video, I’m just going to say Fable because I don’t want to say Claude Fable 5.1 every single sentence, but henceforth, we’re referring to Claude Fable 5.1. In this demo, it is shown taking in raw B2B data and then running revenue forecasts unattended overnight, producing reports for analysts. To then review. It can perform its own back tests through a simple like ChatGPTot interface. All of that’s fantastic. Now, the Lokad follower sent it to me and said it would be great to get your guys’ take on that for obvious reasons because if you have commercially available AI that essentially runs forecasting unattended overnight producing the language was producing results. Now we’ll get into what result like forecast versus decision and what that means but commercially available AI producing forecasts overnight. They wanted to get our opinion particularly yours obviously on what that means for the marketplace. So join us you’ve seen the video you’ve got the context. Were you blown away?
Joannes Vermorel: Effectively a little bit yes but it is impressive to be fair. It is very impressive and but what is much more impressive than the that this use case let it on is that forecasting is just an example. They could I mean what has happened since essentially the last September I would say so one year ago is those coding agents have are now I would say have gained a degree of autonomy for relatively straightforward-ish tasks that is absolutely super impressive and so I would say for like one year now it and it was before even the Claude Fable 5.1. It’s again the real leap for what happened I would say around last summer. Not this one, the one before. You can actually ask relatively straightforward-ish tasks that you would ask to a business analyst. One of them would be “Build me a forecasting model and then just make sure it fits the data with that we have historical data with a back test and then just revisit this thing every day just to make sure it is still like relevant and yes, it will work.”. It will work and what it will do is in fact it will build the connector to connect from the data sources assuming that it that it can you know write a few SQL queries. It will dump the data in a few flat text file. It will compose a Python script. It will run the Python script. It will then compose a mini web user interface and it will just present you the results in a nicely packaged fashion. And on the on the side it will also build the back test utility so that to make sure that the model the fit is reasonably good on the data that you have. And then with something separate that is not exactly clarify in the video, but you can just have you can have your LLM being called on according to a schedule daily to revisit the whole contraption and just refresh the forecast. The type of forecast but let’s make no mistake the type of forecast that is going to be built is just going to be a regular you know old school forecasting model there is nothing in this video that hints that Fable was doing anything particular here it will just compose you know some kind of regressive models with a few parameters to take into account the cyclicities business as usual they mention a Monte Carlo process 1000 Yeah, that’s what he said. Which is probably like the looping the constant that they that they used in Python just to run a few iterations. Again, nothing too fancy. But what is so what is really impressive is not the forecasting results which I think is extremely you know run-of-the-mill but the fact that you can with those with those agents you can throw pretty much anything you would throw at a business analyst and in quasi autonomy you will get a fairly decent result for those tasks that are relatively straightforward assuming that you have already an environment that is already set up. So you see because the setup environment is most likely going to be the most complicated thing such as do you have like a read-only access to your database. You don’t want to give your agent you know a read-write access to production. So does it have like a read-only access? Do you have Python installed? Do you have node installed? Do you have u this and that and that and that and all the bells and whistles and if you want to share the thing with your colleagues then you need to have a way to deploy this web app otherwise it will be just on your desktop and if you want to share then your colleagues will have to install I don’t know git checkout have the exact same permission so the main problem will be in the plumbing the infrastructure the setup but yes, with modern agents and modern LLMs like the state-of-the-art LLMs, those things are literally one-shotted.
Conor Doherty: Yeah. So, that actually you mentioned the phrase straightforward tasks. And I just want to reference something that we talked about a few years ago. I recall when ChatGPT 3.5 came out. So, this would have been late 2022, early 2023. We were discussing LLMs. Again LLMs, large language models, and you made the point that well, they excel at text-based tasks, and now we’re talking about LLMs essentially crunching numbers, more or less. Is that the claim that you’re making?
Joannes Vermorel: No, it doesn’t touch any number. Almost none. It touch Python.
Conor Doherty: Okay.
Joannes Vermorel: You see, it’s if they the approach the model will almost never touch any number. It will compose a SQL query that is text. It will compose a Python script. This is text. It will run it will compose a Python script for the back test. This is this is again text and then it will get like a mini summary of the fit that is just going to be like a handful few numbers and that it can process you know an LLM even back when it was GPT 3.5 it could process like half a dozen of numbers in a conversation what it could not do is process thousands of numbers but here this is not what is being done and Fable doesn’t even try to process thousand of numbers it’s Python script with classic execution that does that.
Conor Doherty: Okay. Well, then that sort of comes back then to the framing question at the start which was when the video was sent to us the well explicit question was how does this breakthrough even if it is still quite nascent but how does this general breakthrough of LLMs into forecasting how does that impact vendors in the space for example people who sell forecasting services like why would anyone continue to pay a subscription to a vendor if they can just get a subscription to Anthropic
Joannes Vermorel: Yeah I mean I mean First we need to pose those vendors were already obsolete 20 years ago you know it’s I mean 20 years ago way before the LLM. So the fact that you know I mean this is this is where it might be strange but imagine you have people who are using typewriters and now Apple say you have the next iPhone. I mean, is it because you have like an even better iPhone that it will discourage people who are still using typewriters from using the typewriter? No. You see, it’s don’t I mean, the vast majority of softwares in I would say around planning in the realm of supply chain are conceptually completely outdated and they have they have been outdated for literally two decades. It’s very interesting because there is even new companies that are still emerging in this area and day one they are completely obsolete by 20 years. This is mindboggling. You we still have competitors that literally started 2025 2026 and they start with concepts that would have been obsolete 20 years ago. So that’s the first thing is a reality check is that those vendors I would say are completely unimpacted because if their client had any unfortunately any sense they would have realized that they were completely obsolete way be beyond before the LLMs. So you see that’s where I say yeah it’s super nice but it doesn’t really change much. Then we have another thing which is and it’s something that we have discussed many times is naked forecast is an anti-pattern you know it’s going for the forecast with a time series perspective is broken we have discussed that at length on this channel it just doesn’t work and that’s the problem is that for many companies and many people for them it’s if you tell them the problem is the time series as a concept is what you should get rid of. So that the problem is that can an LLM craft the very best time series forecasting model? Oh, absolutely. Absolutely. And I would not be surprised if we give, you know, Fable time to compete, you know, on the Kaggle competition that it would score quite high. No, maybe not state-of-the-art, not maybe not the top, but it I would not be surprised that in a you know in a classic forecasting competition that something that would be completely mashing off would score quite high. But the problem is that the time series perspective is defective. So if you if you ask and that’s by the way that’s a problem that LLMs don’t fix is that the LLMs are very compliant. If you ask them to say “Find me a faster horse,”, the thing will find you a faster horse. It may suggest that horse are not super appropriate, but ultimately those LLMs and that’s a good property are extremely docile. They do what they’re being told which is it’s part of the alignment you know that thing and it’s great that makes them extremely useful but also that means that if you ask for something that is essentially a technological dead end the LLM will be absolutely happy to deliver on that and so that’s the problem is now we can ask the question is the LLM able to so a state of the art model able to do the alternative approaches. Absolutely. Absolutely. So I’m not saying that you know Fable 5.1 is not able to do the alternative. It is absolutely capable to do the alternative. Will it do it? Not unless you ask it to. So that you see that would be the first the problem is that it’s it very becomes like the prompting skills but not in the sense of fancy prompting techniques. Is more like the high level perspective on what you’re asking. So naked forecast if we’re back is an anti-pattern. You don’t want to generate time series forecast. I mean it’s completely useless. And yes the LLM can completely automate that but that just that’s just the wrong thing to automate.
Conor Doherty: The way you framed that there were a couple of sentences you used and I think it actually sets up the next question quite well which is what are you asking this tool to do and that actually comes back to something you said right at the start which was and I’m paraphrasing you said let’s say it can forecast that’s not the same as saying it can plan and I think the difference between forecasting and planning is something that is worth spending a bit of time on so for example if you said to well actually no why would I’m going to let you outline if you hand please the functional difference between having a tool that produces a forecast and having a tool that produces a plan which presumably is a series of decisions to be taken.
Joannes Vermorel: So unfortunately for the market the paradigm is the forecast is a plan. You see that’s and that’s we are back to time series are defective and most vendors operate on obsolete paradigm and the problem is that the problem is so widespread that even companies like Anthropic when they want to advertise their capacity of their model what do they do they go back to a time series model which is completely broken and defective because that’s
Conor Doherty: Well it’s what the market wants or it’s what the market is familiar with or used to or many
Joannes Vermorel: Exactly. It’s what the exactly that’s what the market is f is familiar with. But this is this is just a defective paradigm. So that’s the you see and I think that’s part of the problem is that we I would even reject you know entirely the classic notion of planning because the classic notion of planning is a two-stage process where step one you know the future step two you just orchestrate the resource allocation to match and ensure compliance with this future. You know that’s planning is just essentially forecasting plus orchestration and what I say is that this perspective is broken. So yes planning in the classical sense is more than just forecasting but the agent can very easily do the orchestration. It’s super easy and it will be just as broken because again the problem is that the paradigm is broken and if you ask an agent to operate in a broken paradigm it will happily comply but again the paradigm is broken. So it doesn’t matter how smart and capable is the agent. It will just not deliver the return on investment that you expect. Conor Doherty: Well, I’m going to get you to speculate a little bit here because granted the video that actually prompted this discussion dealt very explicitly with revenue forecast. So, it was not it was not a demand forecast per se.
Joannes Vermorel: Yes.
Conor Doherty: And obviously when we get into decision-making in a supply chain context, this is a supply chain company. It’s a supply chain channel. Your like demand would be a very key part of that forecasting process. Now when we get into the difference between forecasting and planning or what we would call making decisions, I’m curious how you see and you can speculate how you would see Fable or similar products invol excuse me my brain medicine how you would see similar products involving themselves in the decision-making process in supply chain. So, for example, if you’re director of supply chain, yeah, you want to know if the company’s making good revenue, if it’s going to be a profitable year, sure. But on any given day, if you have 10,000 products and 10 locations, that’s a lot of possible SKU location combinations that you have to consider and then you have to make decisions with that information. And so in terms of that lived experience and those decision-making problems, do you see tools like Fable, I’m not saying Fable specifically, but like Fable, that sort of LLM product helping in those situations?
Joannes Vermorel: I mean, yes, absolutely. I mean even Lokad we have we have started more than a year at using those LLM tools as productivity tools you see because it can it can now pretty much replace an entry-level data analyst that’s completely clear those tools also those large models when it comes to I would say advanced mathematics advanced statistics they are just completely superhuman already So on specific aspect of the work they are already completely superhuman and they have been for quite a while. Again when I say quite a while again it’s like one year ago.
Conor Doherty: Okay. Not that long ago in terms of companies don’t move in one year cycles.
Joannes Vermorel: So I mean yes but it’s not like super fresh and it’s not it’s not related to the latest release of Fable.
Conor Doherty: Okay. It was
Joannes Vermorel: It was it has been the case for more than a year.
Conor Doherty: Okay.
Joannes Vermorel: And I think you know what it changed is that all the vanity metrics which used to employ I mean in companies to get vanity metrics the CFO wants tons of vanity metrics. The CEO wants tons of vanity metrics and every single directors had their pet metrics that they want to keep track of. And in classic companies, you end up with a business intelligence division that employs again potentially dozens of people. When I say I mention a company, you know, the company I picture is like a $1 billion North American company. So if we $1 billion of annual revenues, bam, you have like a 10 to 20 people who are just essentially doing business intelligence and building reports that are requested by the various directors and usually those reports they are very transient you know it’s just like they just want it once or maybe they will request it again next year. This is the sort of things where okay agents are just going to completely automate that. When it comes to supply chain agents can deliver a massive boost in productivity in the right hands. So if you give to a very smart capable supply chain scientist at Lokad an agent then those people can do great things with those agents. Because supply chain scientists are already in the mindset of we are crafting and that’s that has been the mindset for more than a decade now a decade and a half actually we are we were already in the mindset of the way we approach supply chain is by crafting a numerical recipe that address the problem end to end and now with a coding agent I can just do that much faster and yeah it’s awesome It’s a it’s a massive win. It works. If you want to have auto the if you want to apply the same sort of idea in a organization where the supply chain decisions are done manually then LLM barely helps at all. You know that’s that the problem is that we are back to things that we discussed in previous episode. LLMs are quite slow when it comes to revisit tons of things. So the LLM can be super human if it comes to write the script that does your forecast economic forecast does your craft your numerical recipe that will generate all the decisions unattended. Yes, absolutely. But if you want the LLMs to accompany you in your traditional way of doing the for of doing supply chain which is go line by line on the long spreadsheet to adjust the min and the max of inventory positions then this is going nowhere. This is going nowhere. You really need if you want the LLM to be of any use, you need to embrace the paradigm of an attended decision making where decisions are completely are steered with a numerical recipe which is not an LLM which is typically just classic software except that this classic software is now written in parts with a coding agent.
Conor Doherty: Okay. So I want to tighten the question because I think in that you’ve answered a lot of the things I asked but I wanted to be very clear for people who I know are going to ask essentially this version of the question. So I’ll take the and I’ll use very round numbers to make this super straightforward. So you gave the example of let’s say a North American company making a hundred million revenue each year. A billion a billion sorry a billion a billion with the big boys. We work with the big boys. So, a billion, right? And let’s say even let’s say that company, I’ll make it even simpler. Let’s say that company uses Fable one day and it gets they’re told, yeah, you know, next year you’ll also do you’ll do 1.1 billion in revenue. That’s what we that’s what we forecast for your revenue. Okay, let’s say that company has 10,000 SKUs and has 10 locations. That’s 100,000 SKU location combinations. Knowing that I will do 1 point that I might do 1.1 billion of revenue next year is a piece of information. It’s certainly nice. It’s 10% profit over what I had last year. Okay, cool. However, what does where do I send the units that I have today? Do I send this one to
Joannes Vermorel: Yeah, the store in New Orleans. That’s what I was saying about the vanity metrics is that and that’s the problem with naked forecasts is that naked forecasts are useless because again they cannot be used. So that’s you see If you if you approach the problem incorrectly, so that’s the classic pro planning paradigm where you say I know the future and then I orchestrate. Yes, then it can be used. But this paradigm is just broken. It’s broken because your number when you say you assume it is not just approximate. You say I know the future and I have quasi 0% error on that. But you know that you have error. So it’s it doesn’t work. You see that’s the thing is that if you project 10% growth okay but it is not a relevant information you see it’s it do not cascade decision I can’t do something it cannot do and it’s not even at the right aggregation level because your decisions leaves time wise the scale of decision is like daily decisions and the granularity is granular to the SKU. So your macro forecast is irrelevant. So that’s why I say you have the vanity metrics. They are for the directors. They are for high level human understanding and it’s fine. It’s completely fine and agents can do that. It is fine. Boom. And then you have really the automation of the supply chain decisions. The myriad potentially thousands or tens of thousands potentially millions of daily allocations of resources. And this requires another piece of software that is going what we call at Lokad the numerical recipes just because it’s a mix we it’s a mix of machine learning algorithms optimization algorithms data crunching for preparation preparing the data etc and the LLM can help you to compose this those numerical recipe no problem and then those numerical recipes will run independently from the LLM. No problem. But again it requires for that to work you need to be in to approach the problem from the perspective where supply the supply chain decisions the allocations of resources are steered by unattended decisions that are the consequence of a written numerical recipe you know and that’s all that I’m saying. So I’m not sure if it clarify if you do not so and this paradigm that I just described with numerical recipe is incompatible with the classic paradigm which is I know the future and I can just orchestrate the resource allocation th those two paradigms are just not compatible. So you can only pick one.
Conor Doherty: I think where confusion might arise is your you is your use of the term used the verb used. You said like a forecast like a naked forecast you can’t use and someone hearing that might say might think well you also like we use probabilistic forecasting for example obviously we use that information in service of something else. So clarify what you mean by you can’t use this information
Joannes Vermorel: Every single time you people mention forecasting it’s implicitly time series forecast relatively aggregated okay always you know it’s so I’m not even specific you see there are so many assumption it’s going to be equispaced time series so forecast per day per week per month I mean in theory you could consider non-equispaced time for time series forecast. I’ve never seen anybody in practice use that. So, so you see when we use the term forecast, it comes with extremely narrow assump mathematical assumption just like when you say safety stock it means normal distribution. It’s like in Yes. Yes. Yes. Mathematician would say oh yeah in theory in principle you can have a safety stock that doesn’t follow a normal distribution. Yes, except that 99% of the software of the market when they say safety stock it’s a normal distribution. So, okay. So, you see it’s this sort of things where in theory you could do it differently but I’m just using the term in the way that is generally understood. So, when I say forecast and I say forecast is the naked forecast on pattern we are talking of the equispaced time series forecast. So it’s and it’s going to be at a at an aggregation level where you’re looking at time series that are not super sparse. So whenever people think a time series they picture a time series they picture something that looks like a curve. But if you are at the SKU level what you have is mostly zeros and occasionally one. So if you look at the time series for a SKU that is it’s pass data over the course of one year it’s like 90% zeros and occasionally ones that represent the fact that you’ve served one unit or sold one unit and that’s absolutely not what people have in mind when they think time series. And for example, in this video of code that they that they put for their forecasting, they had like a super nice smooth hyper-aggregated time series. And again, that’s exactly what people have in mind, not the super disaggregated situation that you face in supply chain when you are at the granularities that matter for the decision for the allocation of resource. So that’s what I’m saying and when you operate pred you want to make predictions I’m using this term on purpose different from forecast predictions at the at a granularity that matters for your decisions then this thing this mathematical instrument is so radically different from the time series forecast from the from what people call actually forecast that It’s better to even call it from another name because those things have almost nothing in common.
Conor Doherty: Okay. When we talk about supply chain decisions, I’ve just written a short list of the ones that I think would speak to the greatest preponderance of people in supply chains. We’re talking about what to purchase, which supplier to purchase from, when to order, where to put the inventory, what to manufacture, what to expedite, what to discount. We can go on, but those are those are just
Joannes Vermorel: Yeah, we could go on.
Conor Doherty: Those are just seven, right? So the question really is
Joannes Vermorel: And at which price point
Conor Doherty: And at which price point exactly and should I keep things today or should I switch suppliers like there are all these there’s different classes and there’s second order and third and fourth order we can get very meta here but just to take seven really concrete ones that people can build in their heads. So, with those examples in mind, telling me that my revenue next year is 1.1 billion, that’s information, but does that get me any closer to being able to granularly answer those questions on a day-to-day basis? And if the answer is no, why not? Because if I watch that video, I might get that impression. And I’m not picking on Anthropic. I’m just taking as an example, it was presented to me, what does this say about companies like, for example, Lokad? Well, we’re not selling isolated information like here’s your revenue for next year. It’s we’re selling answers to those questions essentially.
Joannes Vermorel: Yeah. I mean again here the point is that the intuition of many people when it comes to higher dimensional statistics is just plain wrong. So let’s take an image. You’re thinking a movie is a series of image. Imagine that you want you have one frame and you want to forecast anticipate what will be what will be the pixels for the next frame. Okay. The revenue next year is just the average color of your image. So your average color is just going to be something like a shade of gray. That’s average over the entire picture.
Conor Doherty: Yeah.
Joannes Vermorel: And now you forecast the shade of gray which is going to be the average color for the entire next frame. That’s your next your revenue forecast. Does that help you to forecast the pixels? You know when you are one pixel at a time for the image to forecast what comes next? Not at all. Why? Because if you aggregate everything, so you just re you just, you know, collapse your entire image into one average color.
Conor Doherty: Yes.
Joannes Vermorel: You lost the image. You don’t you see nothing. And so when it comes to project what it will look like, what the next frame will look like, the answer is should you restart from the average color? Certainly not. You want to restart from the most granular information, which are the pixels themselves. And you see it’s and the idea that can you have like a very accurate average forecast for the next frame when it comes to the average color? Yes. Why? Because the average color from one frame to the next is pretty much the same. You know, as long as you don’t change what you’re looking at, it’s going to be pretty much the exact same average color. So from even if from image from frame to frame a lot of things are changing if you just average out the whole image it’s almost the same frame to frame unless you know you cut and you have a different perspective. So what I’m saying here is that when people think the forecasting the revenue yes you can do that can you have a relatively accurate forecast possibly because it’s hyper-aggregated will this hyper-aggregated forecast be of any use to when you when you are actually back into the super disaggregated level which would be the equivalent of the pixels in our image. The answer is no. You see and again if you want to understand why just think of you want to go from one frame to the next in terms of an image and you think should I should my logic to predict the next image just go through the idea of just collapsing everything to the average color. The answer is no. So if I were I want to I want to frame this question very concretely. So imagine I have a commercial subscription to any commercially available AI tool. I’m not picking on paper. I’m just saying for example, yes, I’m at that company, the billion dollar turnover and I have a subscription to yeah, a readily available commercially available AI tool and I say “Here’s access.” I create an API, give it access to my ERP all my historical transactional data and I say “Okay, I want to know where to send all these units tomorrow.. So like I need or sorry excuse me I need to place a PO tomorrow. Tell me what I should buy, in what quantities.. You can use whatever forecasting model you want. I’m really simplifying the question, but you can use whatever forecasting paradigm you wish. Just give me an optimized purchase order.”. Would that work? And I realize that’s a simple question, but that is that is the way a lot of people consider this, which is not to say it’s wrong, but it’s just they want to know what benefit I get from this technology. So, I mean right now I suspect it would not work unless you are steering the model in sensible decisions. Again there is too many those LLMs are extremely good but they are not patient. They don’t there is so much they don’t know. Yes. You need to steer the model. You know the for example the model the LLM cannot guess that you have a CRM that has been in production for seven years until you say so. You see if you say all my data is in the ERP but in fact a big portion of the data is in the CRM and you never told the LLM that it should also look at the CRM then it doesn’t know then you can be a very smart prompting you can be very smart at prompting and you can go into reverse mode where you say oh I want to do that “Ask me all the questions and you can ask the LLM to a “Ask me all the questions that you want to be answered. I will do my best,” and then let the agent take over the process. It can work to some extent. It can work to some extent. My own experience is that even with the latest release like Astra from OpenAI which is almost a class of its own you still need high level steering but it is it is good and I think you know the if you wanted to get good results what you should ask a tool like Astra would be u help us replicate what look at does and I suspect for small setup it might actually get you quite far. The again the problem is maintainability and auditability. Yeah. Yeah. You see that’s a key concern and here you so if you ask today a very smart agent just “Replicate Lokad for me,” to some extent it might actually succeed. Maybe not to deal with terabyte of data because that’s that creates there are so many shenanigans with that. But assuming that your data can fit comfortably on a powerful workstation that you have, you know, where the coding agent is running, you could pro probably go quite far. The challenge will be to deal with the dark intelligence. You see, that’s a challenge at Lokad. We have been working with this challenge for again a decade and a half. That problem was there even before we had LLMs. It’s how do you make sure that the numerical recipe remains you know under control that you understand what’s going on that and by the way you can have this problem with human you have a colleague that do something extremely fancy and complicated and then this person leaves and then nobody has any clue on why it works how it works etc so that’s a very real problem that you have with I would say with dark intelligence when you just vibe code something that is very fancy and is that you may you may actually very much struggle into maintaining that over time and the LLM itself might when it revisit this code because you have to think LLMs are kind of stateless so the agent is going to every time it’s going to revisit the codebase it’s going to be as if it was the first time seeing the codebase and the next time the agent revisit this codebase, the agent might actually be lost in okay, why was it done like this? I don’t know. That’s that I with the level of agentic intelligence that we have right now I believe it is not reasonable to not have high level human supervision for things where you for processes that are mission critical that would be my take you can you can just go yolo you know and say “You know what? I’m just going to trust the agent.”. It’s going to vibe code the whole thing but I suspect it’s a recipe for disaster just because we are not quite there yet in term of agentic intelligence and also people may not realize but part of the problems when you’re using very smart agents is that I would say at least half of the foolish things that they are doing is on you. It’s because you are suggesting something that is dumb and again the agent is very compliant. It will do the stupid things and that’s a big problem. You know the agent they are still lacking the capacity to push back on stupid human request. You know they don’t have a mode for that. Maybe the future models will have a mode that is like push back but right now it’s still very weak and that’s a problem that means that you are one bad request away from doing something that is tragically bad for your codebase.
Conor Doherty: Speaking of high level essentially high level supervision reminds me of that the fact that so the commercial director at Lokad, Fabian Hoehner and I will be doing a live demo of Lokad’s aerospace account. Now in preparation for that he and I were discussing it and he mentioned that obviously much like in any other vertical in which Lokad operates the decisions are automatically generated but because of the value of every individual decision in aerospace those still get checked. So, for example, if it’s buy one unit and that unit costs $100,000, that’s probably going to be checked, not because they don’t trust the system, but because the cost of being wrong there is so high. But there are other verticals where buying one unit might only cost a handful of cents again depending on price breaks. So when we’re talking about AI essentially commoditizing a lot of these tasks, do you see certain verticals as being more easily subsumed by this kind of technology than others?
Joannes Vermorel: I mean obviously there are verticals that are very I would say engineering minded you know oil and gas, aviation, consumer electronics. So I suspect e-commerce. So I suspect those companies will be on the front line of automating everything. Then will come the industries where it’s just super tedious because there is just too much stuff like fast fashion.
Conor Doherty: Well, that’s what I was example.
Joannes Vermorel: Yeah. That will that will come. Also people are maybe less I would say oriented toward technologies but the complexity makes those technologies extremely appealing because otherwise it is extremely tedious. And last we will probably end up with as it is very frequently the case FMCG companies that are just you know lag behind because for them they are dealing with few products comparatively very high volumes. So even if the productivity is not good it is not as impacting as it is for other verticals.
Conor Doherty: Well, I asked you earlier in the framing of the discussion was the potential consequence of this kind of technology on vendors in the space and the question is like let’s just assume for the sake of discussion that the that excuse me that the direction of traffic is real. So that this technology is only going to get better and better and better which it presumably will. So that means if you take the Fable video as an example that data sorting in and of it data sorting and cleaning gets quicker and faster and cheaper, forecasting gets quicker and faster and cheaper. Generating reports, back testing is done in a handful of seconds and it’s all just automatically done at a click of a button. Okay, cool. Let’s say all of that’s true and it gets really good. What becomes less valuable in this space in terms of vendor provision and what become and what becomes more valuable?
Joannes Vermorel: So my take is that 90% of those vendors are going to go bankrupt.
Conor Doherty: Okay.
Joannes Vermorel: But not because of because of exactly they are going to go bankrupt because their own clients will go bankrupt.
Conor Doherty: Very different.
Joannes Vermorel: You know it’s so the way I see it is that and again that’s something I was discussing for a few years prediction. You know, it’s my take is that especially with the those agents over 90% of the white collar work is going to evaporate. No, it’s and it might even be more than that. Right now, I mean it’s a 90% was a prediction I was making a few years ago. Now we are beyond that. I mean frankly if we look at what people on average are doing I mean it’s probably over 90% of the white-collar stuff that can be automated completely. Which means that some companies will go will get on board and many will not many will not. It’s if you look at the history of innovation, you might think that electricity is obvious, but when you look at what happened at the end of the 19th century, the answer was most companies went bankrupt and did not adopt electricity. And that was the same with the next revolution that was automotives. Most companies did not adopt automotive and went bankrupt etc. So it’s it has been the cycle has been repeating over and over and with like major tech changes the reality is that most companies never manage to make the leap and just disappear and I think in our lifetime in my lifetime I suspect AI will be the biggest one. I mean I’m I can imagine a technological jump that would be even more significant but frankly it looks like it’s absolutely massive because in modern in modern economies like France or the US you know white collar workforce is 80% of the workforce and this is we’re talking of having this automated away at 90%. So the impact will be absolutely immense and the and the way I see it is that a few just a few companies will actually adopt that embrace this sort of things and they and they will just push to bankruptcy all the other companies and again that’s Schumpeterian revolution it’s it will just happen it’s not like again don’t I’m not predicting mass unemployment because other companies will appear so it’s just fine you know it’s the market will resolve that very cleanly as it does which is lot of companies go bankrupt. People are you know get unemployed for a few months and then they go back to another company because the other winning companies will be hiring like crazy right and left for other jobs. So back to the vendors, my take is that again if we go back to the beginning of this discussion, my take is that most of the vendors in the realm of enterprise supply chain software were already obsolete 20 years ago. So you see that’s the thing is that okay one more tech revolution does it change anything? No. No. I mean, if you’re obso you’re already obsolete 20 you’re 20 years old obsolete, the fact that you know you’re like 30 years old obsolete is well as long as your clients are willing to pay for obsolete stuff we can still read in the I’m still reading in the news that apparently for system of records that mean things like SAP which is the thing that is like the easiest to vibe code. It works beautifully. It has been working beautifully for years now. I mean even I say there was like an agentic breakthrough one year ago but even before that CRUD software for systems of record was already made trivial like at least two or three years ago. So the bottom line is there if you have a company that is still spending more than $1 million annual no matter your size on a system of record. This is just foolish spending. You know this is like pure loss. You’re just spending your money on frivolous thing. So I mean that’s a baseline. No system of record is worth more than like $1 million per year. I mean again if you have really a lot of data we can discuss the fact that you have hardware cost but even you know $1 million a year for records you can I mean you could you could you get close to be able to manage all the transactional records of Walmart so for this budget. So that really begs the question of what’s what sort of records are we talking about? You know, if you’re saying that you need a budget that is bigger than the one that would be needed to actually keep the data of Walmart again represented in a in a smart way, not super bloated and extra but again coding agents are very good at optimizing the code. So that they could they could do that. So that’s if we go back to the pictures is my take is that vendors selling obsolete products will keep selling their absolute product as they did for the last 20 years and the and the and this charade will only end when their own clients go bankrupt because you see at the end of the day the market is not a great educator it’s just a filter and so the companies who operate on obsolete paradigms And if you operate on obsolete paradigms, the consequence will be that instead of having a 90% workforce reduction for most of the jobs, you will have like a 5% reduction in workforce. So that you will just have like a cosmetic improvement and but that’s not what you should be looking for. Again, I think nowadays with the tools, a typical AI project should when you when you go for a function in your company, establishing like at least I don’t know slashing the headcount by two should be like a very non I would say it should be like a baseline. You see if you can if you carry a project where you are deploying AI and you say when we are done the headcount of whatever function we want to augment with AI has not been divided by two it’s a it’s an abject failure that’s two is really not ambitious nowadays for this sort of automation that you can get if a more reasonable would be like 75% and the long-term goal let’s say three to five years from now would be 90% But head count slashing. Yes. Again we are talking of selectively choosing. So here if it’s a function where AI you there is a way to use it and for many things there is then slashing that count by two is not even really ambitious. And I think that the problem right now is that many companies on that are absolutely not you know seeing those results because the way it the way they approach AI and the way they approach automation is extremely I would say obsolete in term of paradigm.
Conor Doherty: Well, we were talking about vendors and this conversation actually came about as a result of a friend of a channel sending us that video to get our response. And there was one question that he wanted me to ask you and I’ll close with it. And it was, imagine tomorrow a vendor walks into my boss’s office and says, “Our AI,” quote, “Our AI autonomously plans your supply chain.” What should I ask them to demonstrate?
Joannes Vermorel: The question is you should ask how and you need to understand how what is going on under the hood that’s wholesale company I should say for context a wholesale company yes I mean for you need to understand what goes under the hood you see because that’s the key is that and what goes on under the hood. “Walk me through it. I want to understand.” And then we’ll Can we go on a path where you are committed to actually deliver the headcount savings that we discussed together? You see and conversely there are red flags if the vendor is charging per seat that means that their it is in the best interest of the vendor to just minimize productivity to make sure there is as many people as possible and retain headcount and maintain account.
Conor Doherty: Yes.
Joannes Vermorel: Exactly. Because you see if for example another red flag would be if the vendor say we are going to augment what your people are doing. Okay absolute red flag. This is not the way it would work. You know you cannot achieve 90% headcount reduction by saying we augment the productivity of people. You have to rethink the work. So that’s why I say okay walk me through how do you achieve this ideally we want to go the perspective five years from now is 90% and that would be a nice baseline 90% headcount reduction so we are 20 people five years from now we will be four tell me exactly how it works what your system is doing and why you’re confident that with four people sorry that would be sorry two people yes out of 20 we will be able to operate and that’s that will be one of them.
Conor Doherty: So it’s one other person in your example. Yes.
Joannes Vermorel: Yes. And again I think usually enterprise software vendors are pushing technologies that are so obsolete that when you actually ask what is going on under the hood very quickly you realize that it’s extremely obsolete and you can even nowadays it’s a magic of ChatGPT is that even if there are like technical terms that you don’t understand just record the conversation and ask ChatGPT to for an assessment and ChatGPT will give you a reasonably solid assessment you know just ask ChatGPT to take an adversarial stance otherwise it’s just way too kind with in term of criticism but just take okay just set the premise on Dear ChatGPT, this vendor is probably you know inflating his claims and embellishing the situation “Be merciless and criticize what you present to me,” and you will get a decent criticism, but you need to have a look at what goes under the hood.
Conor Doherty: To be fair, just to echo that thought, it has never actually been simpler to get pretty solid market research on just about any vendor.
Joannes Vermorel: Yes.
Conor Doherty: And you don’t you don’t even need the most powerful model. We’re not sponsored by ChatGPT or excuse me, by OpenAI, but you don’t even have to use the most sophisticated model of ChatGPT. Here is a vendor, here’s the website, evaluate, give me a market report, compare it to peers,
Joannes Vermorel: But beware, you need to ask to have an adversarial stance because otherwise ChatGPT is way too trusting. So you have to say do not trust case studies, do not trust any claim made by the vendor. So do not substantive.
Conor Doherty: Exactly.
Joannes Vermorel: Make sure that you can ground your assessment into facts. What does the technology look like? Assume you need to assume the worst. Whenever a piece of technology is not described, assume it’s a moving average. Whenever something
Conor Doherty: No, it’s enterprise software. This is functional.
Joannes Vermorel: This is functional. Yes. Enterprise software works like that. If the vendor does not give you the detail, there is no magic recipe. There is no magic ingredient like an hyper advanced technology. No, no, no. This is a special AI. Yeah. No, no. Assume the worst. For example, if there is no technical documentation, you need to assume that it’s a nightmare and that’s why they don’t publicize their technical documentation. It’s because the product is a big pile of mud and they are and they are the vendor is ashamed is so ashamed that they don’t want to publicize. You know, it’s so you just need to tell ChatGPT to adopt this adversarial mindset. You know, ground your assessment into what you can observe. Make sure it’s factual. Just any claim, be extremely skeptical, especially with case studies. Assume that all the case studies are forged because usually they will be and or the very least and reason from first principles. So reason from first principles means okay the vendor is telling that is saying that they are doing everything in a SQL database. Is it is it something that is really compatible with advanced machine learning technologies? You know, just ask ChatGPT reason from first principles. Just look at what they’re doing and from that extrapolate with what whatever they’re saying is really credible or not and whether it will, you know, be it will achieve this 90% headcount reduction, you know, five years from now, which should be on the horizon.
Conor Doherty: What I’ll do is when we post this recording, I will have a ChatGPT made summary of that diagnostic and we’ll post that with this clip to promote this podcast because I think that is actually very functional but also very attainable. It’s not like you’re saying you need to pour through petabytes let’s even say of data and you need to become an expert in forecasting like that period. It was possible to say at one point, yeah, you really need to have a greater degree of mechanical sympathy, which I know we will say is always good, and it is always good, but realistically, most people are not going to do that. But there’s very little excuse for saying, I literally don’t have 60 seconds to review what I’m about to spend maybe more than a million dollars on. That’s really a bit rich.
Joannes Vermorel: And you just have to keep in mind that by default that’s the reason why I was saying that you need to prompt all the models like that is because all those LLMs are essentially fine-tuned to be nice and docile and gentle and not antagonistic. Yes. And as a consequence of this default behavior they will because maybe they are asking something about your colleague and your colleague is there to see. So you see they don’t try to be to present anything in a bad light by default. But if you if you press them with very specific prompt then no problem. The model can actually be deliver a harsh criticism of anything. But you need to help and steer a little bit the model because again the default fine-tuning and that’s the same for pretty much all the models as far I know is to be relatively gentle and kind and positive looking on things Conor Doherty: Generally speaking you issue that challenge with presumably a tremendous degree of confidence that Lokad will not run afoul of that standard. Joannes Vermorel: Yeah I mean which is fine it’s a good position to be in. I’m just pointing that out that again I say comparatively I mean obviously the I did I did I did stress test Lokad in the past on that I haven’t tested the latest model Astra from OpenAI or the latest from Anthropic but Lokad was faring decently on this stress test. Conor Doherty: Okay. Well, I don’t have any further questions, Joannes, but if anyone sends me their feedback that they got from ChatGPT, I will forward it directly to you for comment and we’ll address that at a later date. As always, a pleasure and to you for watching. Thank you very much. If you want to get in touch with Joannes and me, if you want to share with us your LLM feedback